Reference
- Upgrading S3 deployments — S3 storage now always writes a local disk cache, so an S3 container that ran v0.0.8 or earlier can fail or use more resources after the upgrade:
- The cache directory,
/var/cache/helixunlessHELIX_DISK_CACHE_DIRnames another, must be writable. Mount a volume there (required with a read-only root filesystem), and setHELIX_DISK_CACHE_DIRwhen running as a user other than65532. - The hard open-file limit must be at least 26,600 at the default budget (about 8,200 or fewer at 1 GiB or less). Raise it with
--ulimit nofile=65536:65536. - RSS grows by the block cache’s index, about 20–70 MiB at the default budget and twice that briefly after a restart.
- Without
HELIX_DISK_CACHE_BYTES, the 8 GiB default budget must fit the cache’s filesystem, or startup fails.
- The cache directory,
- S3 always uses a local disk cache in
/var/cache/helixby default, with an 8 GiB default budget. Memory-only S3 is gone: once a vector index outgrew the memory cache, every graph hop waited on its own S3 request.HELIX_DATA_DIRand in-memory storage are unchanged. - Faster first vector searches on S3. Cold batch reads overlap their block fetches, expansions read parents concurrently, and projections prefetch the records they need. After open, the server also warms vector search rows into the disk cache in the background; readiness never waits for it.
HELIX_DISK_CACHE_WARM=off(defaulton) turns startup warming off. On 2 CPUs, 4 GiB of memory and an 8 GiB cache over S3, the first global vector top-50 after open went from 30.8 s on v0.0.7 to 714 ms once the ~26 s warm finished. - Index sets are kept across
ForEachiterations and unrelated writes in the same request, which removes the v0.0.8 known limitation. A set is read again only when an iteration binds a parameter it depends on, or a write creates, deletes or changes indexed properties of nodes with its label. A batch that expands, filters on an index and adds an edge per iteration runs about twice as fast. - Source filters and counts always use indexes.
N<L>.where(...),E<L>.where(...)and counts on indexed properties no longer scan the label, including partly indexed ORs, longINlists, parameter lists,== null, and counts over.as()/.store(). A label scan only happens when no condition has an index. On 50k rows, counts after.as()went from up to 1.1 s to under 5 ms, and 200-valueINlists from up to 1.5 s to about 0.2 s. - Upgrading — the stored data format is unchanged. Databases written by v0.0.8 open directly with v0.0.9.
- Known issue — updating a node’s vector (re-embedding it) and later deleting that node can leave vector index (HNSW) links pointing at the deleted node, and searches that reach them fail with
missing simhash. Earlier releases are affected too, and a fix is in progress. To repair an affected index, drop and recreate the vector index. - CLI 3.4.2 defaults to image v0.0.9. With
--diskor S3 storage,helix startmounts a per-instancehelix-<project>-<instance>-cachevolume at/var/cache/helixfor the disk cache, with a 64 MiB budget in disk mode and 1 GiB for an S3 bucket.helix stopkeeps it andhelix pruneremoves it.
- Filters after a traversal always use indexes. A
.where(...)on indexed properties after.out(),.in()or another expansion is answered from the index, including in counts, limits,existschecks and branch or union bodies, instead of reading records one by one. A filter that combines several indexes with AND or OR reads them concurrently, and label-only filters use label bitmaps. On 312k nodes, an indexed count after an expansion went from 41 ms to 33 ms (13 ms to 3 ms on 31k nodes), and the first cold query from about 280 ms to 33 ms. - Known limitation — inside
ForEachbodies (forEachParam/for_each_param), and after a write in the same request, index sets are re-read once per iteration or write. An improvement is in progress. - Upgrading — the stored data format is unchanged. Databases written by v0.0.7 open directly with v0.0.8.
- CLI 3.4.1 defaults to image v0.0.8.
- Traversal-scoped search is much faster. A
.where(...)after a traversal is answered from equality-index bitmaps instead of reading every record. Scopes of up to 8,192 candidates (at 768 dimensions) are scored exactly, and larger scopes use a bounded graph walk. On 312k 768-d nodes in S3 with a warm disk cache, scoped vector queries that took 2–6 s, and a combined filter and scope that did not finish in 30 s, now return in about 60–110 ms. See prefiltering. - Local disk cache — set
HELIX_DISK_CACHE_DIRto keep recently read data on local disk, ideally NVMe, across restarts. - Planner — a conjunction of four or more indexed equalities no longer falls back to a full label scan.
- Vector index builds keep their planning cache across backfill steps, so large backfills run several times faster.
- Upgrading — the stored data format is unchanged. Databases written by v0.0.6 open directly with v0.0.7.
- CLI 3.4.0 defaults to image v0.0.7. Cloud commands resolve workspaces, projects and databases without IDs, and a global
--jsonflag gives machine-readable output.helix --quiet querynow prints the result. - Local
--diskmode now uses SeaweedFS, because MinIO withdrew its community images. Existing MinIO disk data is not moved automatically; see Migrate MinIO disk data.
- TypeScript SDK 3.1.0 exposes
HelixError.code,statusCode,serverMessage,rawBody, and explicitretryablemetadata. npm 3.0.4 did not expose these fields. Legacy response formats and unknown future codes remain readable. - Transaction conflicts can be recognized with
error.statusCode === 409anderror.code === "transaction_conflict". Docker v0.0.5 already supplies this contract; no Docker or CLI release is required. A lost response remains distinct from a confirmed abort, and mutations are never automatically replayed. - This release updates the TypeScript SDK only. Embedded runtime packages and the Rust, Python, and Go SDK versions are unchanged.
- Go SDK v0.3.1 — operation-tree query construction, typed parameters,
/v2/queryexecution, index lifecycle requests, and traversal-scoped vector and BM25 search. Upgrade from v0.3.0: this patch keeps unreleased embedded adapters out of public module discovery. Embedded execution and native graph bindings are not distributed in this release. - Python async client —
helix-db0.3.3 adds reusable HTTPX connection pooling, concurrent requests, cancellation-safe response cleanup, and asynchronous embedded writer and reader clients.
- Open-source database — the engine moved out of
helix-hyperscaleinto the public, modularhelix-dbworkspace. - New query planner — a dedicated logical and physical planning layer optimizes the typed operation-tree AST before execution.
- Embedded database — run the same engine and queries in process with memory, disk, or object storage.
- Vector prefiltering — traverse and filter an exact candidate set before vector ranking, so results cannot escape graph or permission boundaries.
- Row bindings —
bind+projectBindings/projectDistinctBindingscorrelate values captured at different hops of a single traversal. - Python SDK — new zero-dependency synchronous SDK (
helix-db, imported ashelixdb). - CLI 3.0.6 —
--pathflag onhelix add.
- Go SDK
- Better cache metrics for better query insights
- Query plan improvements
helix chefcommand for local bootstrap- CLI DX improvements
- Local docker deployment
- Query plan improvements
- Query insights, metrics and suggestions
- Stability and reliability improvements
- Better memory usage and caching