Reference
- Strong text search keeps working during ingest. A
strongtext search now has its own bound, 512 MiB of unpublished text per partition by default (it was one publication’s 64 MiB), and reuses analyses of pending text across searches. Ingest bursts of 20K–40K documents that made v0.0.10’s strong text searches fail withindex_backpressurenow complete with no rejections, and a strong text search behind 5,000 pending documents dropped from about 200 ms to 30 ms. SetHELIX_STRONG_TEXT_SEARCH_MAX_ANALYSIS_BYTESto change the bound; past it the search fails with retryableindex_backpressure(pending_text_analysis_bytes). - Strong searches decode pending index work once per request. Several searches in one request share one decoded copy of the pending queue, decoding and exact scoring run off the async workers so they no longer stall other requests, and four strong vector searches in one request behind 160K pending 768-dimension changes went from about 5.0 s to 2.2 s.
- Strong vector searches are bounded. A
strongvector search decodes and scores at most the index’s retained-byte limit of unpublished work by default (1 GB, so writes reach backpressure first). SetHELIX_STRONG_VECTOR_SEARCH_MAX_PENDING_BYTESto lower it; past it the search fails with retryableindex_backpressure(pending_vector_bytes). - Queue admission tracks real memory. Each queued index operation now counts its encoded size plus 576 bytes against the 1 GB per-index limit, so a backlog of many small operations, such as deletes, can no longer use several times the limit in memory. Such backlogs reach
index_backpressuresooner; the 250,000 pending-entity limit is unchanged. - Writers reopen faster with a large backlog. Opening a writer loads its queued index work one queue at a time instead of decoding every queue at once: a 2.26 GB backlog opened in about 5 s instead of 25 s, with peak memory bounded by the largest queue instead of the whole backlog.
- Publication drains in linear time. Selecting work to publish no longer rescans the whole queue each time, which was 9.4 s of bookkeeping per 250K-operation drain and is now 0.16 s, and a queue retired by an index drop is discarded without being re-read.
- No storage changes since v0.0.10: upgrading from v0.0.10 needs no migration. Upgrading from v0.0.9 or earlier follows the v0.0.10 notes below.
- CLI 3.4.4 defaults to image v0.0.11. Projects that saved a
taginhelix.tomlkeep using it until you change it.
- Upgrading: index storage version 5. The first time a v0.0.10 writer opens a database, it upgrades the index storage version from 4 to 5. It rewrites only the version marker and rebuilds no index. After that, v0.0.9 and earlier refuse to open the database with
unsupported_index_storage_version:- Upgrade readers before the writer; v0.0.10 readers serve both versions.
- Never start a v0.0.9-or-earlier writer against an upgraded database.
- Take a backup before upgrading if you may need to roll back.
- Vector and text indexes publish asynchronously. A write commits its graph change together with a durable queued index operation, and a background index worker publishes it into the vector or text index, so writes no longer wait on index maintenance. Steady vector insert publication went from 6.2 to 239 per second. Index builds no longer livelock under concurrent writes: on a c7i.8xlarge with 100 writes/s during the build, backfilling 100K 768-dimension vectors on disk finished in about 2,500 s, where v0.0.9 never finished.
- Search consistency. Read requests set
search_consistencytostrong(the default) oreventual.strongsearches include every committed write, published or not.eventualsearches include a bounded amount of unpublished work and serve the rest as last published, so they can return a node or edge that has since changed tenant, label or indexed property. Write requests always searchstrong, and a write batch that searches an index fails with a retryabletransaction_conflictif the index worker publishes into that index first. The SDK options ship in the next SDK releases; until then, sendsearch_consistencyin a direct HTTP request. - Index backpressure. Each index retains at most 1,000,000,000 bytes or 250,000 pending entities of unpublished work. Three cases fail with
index_backpressure(HTTP 429,retryable: true) before the request takes effect: a write past either limit, a whole-indexstrongsearch behind more than 800 unpublished changes, and astrongtext search that would analyze more unpublished text than one publication may (64 MiB by default). Retry it unchanged with bounded backoff. A single write that stages more than 8 MiB of index work for one index fails withindex_operation_batch_too_large(HTTP 400); split it. - Text write limits are per document.
active_text_mutation_limit_exceedednow checks each indexed document against its own allowances, about 16,000 unique terms, instead of limiting a write to 512 text-indexed entities. - A failing index entry no longer stalls its index. An entry whose index update fails the same way every time, for example on damaged index data, is held back while the rest of the index keeps publishing.
GET /healthzcounts it underblocked_index_entity_count, the server logs an error naming it, and the worker retries it about once a minute. See troubleshooting. - Fixed:
missing simhashafter re-embedding and then deleting a node (the v0.0.9 known issue). Re-embedding a node no longer loses track of the vector index (HNSW) links that point at it, so deleting it later removes them all, and a node is never linked to itself. Indexes written by v0.0.9 and earlier can still hold such links; drop and recreate the vector index to repair them. - Fixed: counts over a range index with other indexed filters. A
.count()overN<L>.where(...)orE<L>.where(...)that combined a range condition with other indexed conditions counted the range alone, for example 149 rows instead of 49. Counts now apply every filter. - CLI 3.4.3 defaults to image v0.0.10. Projects that saved a
taginhelix.tomlkeep using it until you change it.
- Upgrading S3 deployments — S3 storage now always writes a local disk cache, so an S3 container that ran v0.0.8 or earlier can fail or use more resources after the upgrade:
- The cache directory,
/var/cache/helixunlessHELIX_DISK_CACHE_DIRnames another, must be writable. Mount a volume there (required with a read-only root filesystem), and setHELIX_DISK_CACHE_DIRwhen running as a user other than65532. - The hard open-file limit must be at least 26,600 at the default budget (about 8,200 or fewer at 1 GiB or less). Raise it with
--ulimit nofile=65536:65536. - RSS grows by the block cache’s index, about 20–70 MiB at the default budget and twice that briefly after a restart.
- Without
HELIX_DISK_CACHE_BYTES, the 8 GiB default budget must fit the cache’s filesystem, or startup fails.
- The cache directory,
- S3 always uses a local disk cache in
/var/cache/helixby default, with an 8 GiB default budget. Memory-only S3 is gone: once a vector index outgrew the memory cache, every graph hop waited on its own S3 request.HELIX_DATA_DIRand in-memory storage are unchanged. - Faster first vector searches on S3. Cold batch reads overlap their block fetches, expansions read parents concurrently, and projections prefetch the records they need. After open, the server also warms vector search rows into the disk cache in the background; readiness never waits for it.
HELIX_DISK_CACHE_WARM=off(defaulton) turns startup warming off. On 2 CPUs, 4 GiB of memory and an 8 GiB cache over S3, the first global vector top-50 after open went from 30.8 s on v0.0.7 to 714 ms once the ~26 s warm finished. - Index sets are kept across
ForEachiterations and unrelated writes in the same request, which removes the v0.0.8 known limitation. A set is read again only when an iteration binds a parameter it depends on, or a write creates, deletes or changes indexed properties of nodes with its label. A batch that expands, filters on an index and adds an edge per iteration runs about twice as fast. - Source filters and counts always use indexes.
N<L>.where(...),E<L>.where(...)and counts on indexed properties no longer scan the label, including partly indexed ORs, longINlists, parameter lists,== null, and counts over.as()/.store(). A label scan only happens when no condition has an index. On 50k rows, counts after.as()went from up to 1.1 s to under 5 ms, and 200-valueINlists from up to 1.5 s to about 0.2 s. - Upgrading — the stored data format is unchanged. Databases written by v0.0.8 open directly with v0.0.9.
- Known issue — updating a node’s vector (re-embedding it) and later deleting that node can leave vector index (HNSW) links pointing at the deleted node, and searches that reach them fail with
missing simhash. Earlier releases are affected too, and a fix is in progress. To repair an affected index, drop and recreate the vector index. - CLI 3.4.2 defaults to image v0.0.9. With
--diskor S3 storage,helix startmounts a per-instancehelix-<project>-<instance>-cachevolume at/var/cache/helixfor the disk cache, with a 64 MiB budget in disk mode and 1 GiB for an S3 bucket.helix stopkeeps it andhelix pruneremoves it.
- Filters after a traversal always use indexes. A
.where(...)on indexed properties after.out(),.in()or another expansion is answered from the index, including in counts, limits,existschecks and branch or union bodies, instead of reading records one by one. A filter that combines several indexes with AND or OR reads them concurrently, and label-only filters use label bitmaps. On 312k nodes, an indexed count after an expansion went from 41 ms to 33 ms (13 ms to 3 ms on 31k nodes), and the first cold query from about 280 ms to 33 ms. - Known limitation — inside
ForEachbodies (forEachParam/for_each_param), and after a write in the same request, index sets are re-read once per iteration or write. An improvement is in progress. - Upgrading — the stored data format is unchanged. Databases written by v0.0.7 open directly with v0.0.8.
- CLI 3.4.1 defaults to image v0.0.8.
- Traversal-scoped search is much faster. A
.where(...)after a traversal is answered from equality-index bitmaps instead of reading every record. Scopes of up to 8,192 candidates (at 768 dimensions) are scored exactly, and larger scopes use a bounded graph walk. On 312k 768-d nodes in S3 with a warm disk cache, scoped vector queries that took 2–6 s, and a combined filter and scope that did not finish in 30 s, now return in about 60–110 ms. See prefiltering. - Local disk cache — set
HELIX_DISK_CACHE_DIRto keep recently read data on local disk, ideally NVMe, across restarts. - Planner — a conjunction of four or more indexed equalities no longer falls back to a full label scan.
- Vector index builds keep their planning cache across backfill steps, so large backfills run several times faster.
- Upgrading — the stored data format is unchanged. Databases written by v0.0.6 open directly with v0.0.7.
- CLI 3.4.0 defaults to image v0.0.7. Cloud commands resolve workspaces, projects and databases without IDs, and a global
--jsonflag gives machine-readable output.helix --quiet querynow prints the result. - Local
--diskmode now uses SeaweedFS, because MinIO withdrew its community images. Existing MinIO disk data is not moved automatically; see Migrate MinIO disk data.
- TypeScript SDK 3.1.0 exposes
HelixError.code,statusCode,serverMessage,rawBody, and explicitretryablemetadata. npm 3.0.4 did not expose these fields. Legacy response formats and unknown future codes remain readable. - Transaction conflicts can be recognized with
error.statusCode === 409anderror.code === "transaction_conflict". Docker v0.0.5 already supplies this contract; no Docker or CLI release is required. A lost response remains distinct from a confirmed abort, and mutations are never automatically replayed. - This release updates the TypeScript SDK only. Embedded runtime packages and the Rust, Python, and Go SDK versions are unchanged.
- Go SDK v0.3.1 — operation-tree query construction, typed parameters,
/v2/queryexecution, index lifecycle requests, and traversal-scoped vector and BM25 search. Upgrade from v0.3.0: this patch keeps unreleased embedded adapters out of public module discovery. Embedded execution and native graph bindings are not distributed in this release. - Python async client —
helix-db0.3.3 adds reusable HTTPX connection pooling, concurrent requests, cancellation-safe response cleanup, and asynchronous embedded writer and reader clients.
- Open-source database — the engine moved out of
helix-hyperscaleinto the public, modularhelix-dbworkspace. - New query planner — a dedicated logical and physical planning layer optimizes the typed operation-tree AST before execution.
- Embedded database — run the same engine and queries in process with memory, disk, or object storage.
- Vector prefiltering — traverse and filter an exact candidate set before vector ranking, so results cannot escape graph or permission boundaries.
- Row bindings —
bind+projectBindings/projectDistinctBindingscorrelate values captured at different hops of a single traversal. - Python SDK — new zero-dependency synchronous SDK (
helix-db, imported ashelixdb). - CLI 3.0.6 —
--pathflag onhelix add.
- Go SDK
- Better cache metrics for better query insights
- Query plan improvements
helix chefcommand for local bootstrap- CLI DX improvements
- Local docker deployment
- Query plan improvements
- Query insights, metrics and suggestions
- Stability and reliability improvements
- Better memory usage and caching