Skip to main content
Reference
  • Strong text search keeps working during ingest. A strong text search now has its own bound, 512 MiB of unpublished text per partition by default (it was one publication’s 64 MiB), and reuses analyses of pending text across searches. Ingest bursts of 20K–40K documents that made v0.0.10’s strong text searches fail with index_backpressure now complete with no rejections, and a strong text search behind 5,000 pending documents dropped from about 200 ms to 30 ms. Set HELIX_STRONG_TEXT_SEARCH_MAX_ANALYSIS_BYTES to change the bound; past it the search fails with retryable index_backpressure (pending_text_analysis_bytes).
  • Strong searches decode pending index work once per request. Several searches in one request share one decoded copy of the pending queue, decoding and exact scoring run off the async workers so they no longer stall other requests, and four strong vector searches in one request behind 160K pending 768-dimension changes went from about 5.0 s to 2.2 s.
  • Strong vector searches are bounded. A strong vector search decodes and scores at most the index’s retained-byte limit of unpublished work by default (1 GB, so writes reach backpressure first). Set HELIX_STRONG_VECTOR_SEARCH_MAX_PENDING_BYTES to lower it; past it the search fails with retryable index_backpressure (pending_vector_bytes).
  • Queue admission tracks real memory. Each queued index operation now counts its encoded size plus 576 bytes against the 1 GB per-index limit, so a backlog of many small operations, such as deletes, can no longer use several times the limit in memory. Such backlogs reach index_backpressure sooner; the 250,000 pending-entity limit is unchanged.
  • Writers reopen faster with a large backlog. Opening a writer loads its queued index work one queue at a time instead of decoding every queue at once: a 2.26 GB backlog opened in about 5 s instead of 25 s, with peak memory bounded by the largest queue instead of the whole backlog.
  • Publication drains in linear time. Selecting work to publish no longer rescans the whole queue each time, which was 9.4 s of bookkeeping per 250K-operation drain and is now 0.16 s, and a queue retired by an index drop is discarded without being re-read.
  • No storage changes since v0.0.10: upgrading from v0.0.10 needs no migration. Upgrading from v0.0.9 or earlier follows the v0.0.10 notes below.
  • CLI 3.4.4 defaults to image v0.0.11. Projects that saved a tag in helix.toml keep using it until you change it.
  • Upgrading: index storage version 5. The first time a v0.0.10 writer opens a database, it upgrades the index storage version from 4 to 5. It rewrites only the version marker and rebuilds no index. After that, v0.0.9 and earlier refuse to open the database with unsupported_index_storage_version:
    • Upgrade readers before the writer; v0.0.10 readers serve both versions.
    • Never start a v0.0.9-or-earlier writer against an upgraded database.
    • Take a backup before upgrading if you may need to roll back.
  • Vector and text indexes publish asynchronously. A write commits its graph change together with a durable queued index operation, and a background index worker publishes it into the vector or text index, so writes no longer wait on index maintenance. Steady vector insert publication went from 6.2 to 239 per second. Index builds no longer livelock under concurrent writes: on a c7i.8xlarge with 100 writes/s during the build, backfilling 100K 768-dimension vectors on disk finished in about 2,500 s, where v0.0.9 never finished.
  • Search consistency. Read requests set search_consistency to strong (the default) or eventual. strong searches include every committed write, published or not. eventual searches include a bounded amount of unpublished work and serve the rest as last published, so they can return a node or edge that has since changed tenant, label or indexed property. Write requests always search strong, and a write batch that searches an index fails with a retryable transaction_conflict if the index worker publishes into that index first. The SDK options ship in the next SDK releases; until then, send search_consistency in a direct HTTP request.
  • Index backpressure. Each index retains at most 1,000,000,000 bytes or 250,000 pending entities of unpublished work. Three cases fail with index_backpressure (HTTP 429, retryable: true) before the request takes effect: a write past either limit, a whole-index strong search behind more than 800 unpublished changes, and a strong text search that would analyze more unpublished text than one publication may (64 MiB by default). Retry it unchanged with bounded backoff. A single write that stages more than 8 MiB of index work for one index fails with index_operation_batch_too_large (HTTP 400); split it.
  • Text write limits are per document. active_text_mutation_limit_exceeded now checks each indexed document against its own allowances, about 16,000 unique terms, instead of limiting a write to 512 text-indexed entities.
  • A failing index entry no longer stalls its index. An entry whose index update fails the same way every time, for example on damaged index data, is held back while the rest of the index keeps publishing. GET /healthz counts it under blocked_index_entity_count, the server logs an error naming it, and the worker retries it about once a minute. See troubleshooting.
  • Fixed: missing simhash after re-embedding and then deleting a node (the v0.0.9 known issue). Re-embedding a node no longer loses track of the vector index (HNSW) links that point at it, so deleting it later removes them all, and a node is never linked to itself. Indexes written by v0.0.9 and earlier can still hold such links; drop and recreate the vector index to repair them.
  • Fixed: counts over a range index with other indexed filters. A .count() over N<L>.where(...) or E<L>.where(...) that combined a range condition with other indexed conditions counted the range alone, for example 149 rows instead of 49. Counts now apply every filter.
  • CLI 3.4.3 defaults to image v0.0.10. Projects that saved a tag in helix.toml keep using it until you change it.
  • Upgrading S3 deployments — S3 storage now always writes a local disk cache, so an S3 container that ran v0.0.8 or earlier can fail or use more resources after the upgrade:
    • The cache directory, /var/cache/helix unless HELIX_DISK_CACHE_DIR names another, must be writable. Mount a volume there (required with a read-only root filesystem), and set HELIX_DISK_CACHE_DIR when running as a user other than 65532.
    • The hard open-file limit must be at least 26,600 at the default budget (about 8,200 or fewer at 1 GiB or less). Raise it with --ulimit nofile=65536:65536.
    • RSS grows by the block cache’s index, about 20–70 MiB at the default budget and twice that briefly after a restart.
    • Without HELIX_DISK_CACHE_BYTES, the 8 GiB default budget must fit the cache’s filesystem, or startup fails.
  • S3 always uses a local disk cache in /var/cache/helix by default, with an 8 GiB default budget. Memory-only S3 is gone: once a vector index outgrew the memory cache, every graph hop waited on its own S3 request. HELIX_DATA_DIR and in-memory storage are unchanged.
  • Faster first vector searches on S3. Cold batch reads overlap their block fetches, expansions read parents concurrently, and projections prefetch the records they need. After open, the server also warms vector search rows into the disk cache in the background; readiness never waits for it. HELIX_DISK_CACHE_WARM=off (default on) turns startup warming off. On 2 CPUs, 4 GiB of memory and an 8 GiB cache over S3, the first global vector top-50 after open went from 30.8 s on v0.0.7 to 714 ms once the ~26 s warm finished.
  • Index sets are kept across ForEach iterations and unrelated writes in the same request, which removes the v0.0.8 known limitation. A set is read again only when an iteration binds a parameter it depends on, or a write creates, deletes or changes indexed properties of nodes with its label. A batch that expands, filters on an index and adds an edge per iteration runs about twice as fast.
  • Source filters and counts always use indexes. N<L>.where(...), E<L>.where(...) and counts on indexed properties no longer scan the label, including partly indexed ORs, long IN lists, parameter lists, == null, and counts over .as() / .store(). A label scan only happens when no condition has an index. On 50k rows, counts after .as() went from up to 1.1 s to under 5 ms, and 200-value IN lists from up to 1.5 s to about 0.2 s.
  • Upgrading — the stored data format is unchanged. Databases written by v0.0.8 open directly with v0.0.9.
  • Known issue — updating a node’s vector (re-embedding it) and later deleting that node can leave vector index (HNSW) links pointing at the deleted node, and searches that reach them fail with missing simhash. Earlier releases are affected too, and a fix is in progress. To repair an affected index, drop and recreate the vector index.
  • CLI 3.4.2 defaults to image v0.0.9. With --disk or S3 storage, helix start mounts a per-instance helix-<project>-<instance>-cache volume at /var/cache/helix for the disk cache, with a 64 MiB budget in disk mode and 1 GiB for an S3 bucket. helix stop keeps it and helix prune removes it.
  • Filters after a traversal always use indexes. A .where(...) on indexed properties after .out(), .in() or another expansion is answered from the index, including in counts, limits, exists checks and branch or union bodies, instead of reading records one by one. A filter that combines several indexes with AND or OR reads them concurrently, and label-only filters use label bitmaps. On 312k nodes, an indexed count after an expansion went from 41 ms to 33 ms (13 ms to 3 ms on 31k nodes), and the first cold query from about 280 ms to 33 ms.
  • Known limitation — inside ForEach bodies (forEachParam / for_each_param), and after a write in the same request, index sets are re-read once per iteration or write. An improvement is in progress.
  • Upgrading — the stored data format is unchanged. Databases written by v0.0.7 open directly with v0.0.8.
  • CLI 3.4.1 defaults to image v0.0.8.
  • Traversal-scoped search is much faster. A .where(...) after a traversal is answered from equality-index bitmaps instead of reading every record. Scopes of up to 8,192 candidates (at 768 dimensions) are scored exactly, and larger scopes use a bounded graph walk. On 312k 768-d nodes in S3 with a warm disk cache, scoped vector queries that took 2–6 s, and a combined filter and scope that did not finish in 30 s, now return in about 60–110 ms. See prefiltering.
  • Local disk cache — set HELIX_DISK_CACHE_DIR to keep recently read data on local disk, ideally NVMe, across restarts.
  • Planner — a conjunction of four or more indexed equalities no longer falls back to a full label scan.
  • Vector index builds keep their planning cache across backfill steps, so large backfills run several times faster.
  • Upgrading — the stored data format is unchanged. Databases written by v0.0.6 open directly with v0.0.7.
  • CLI 3.4.0 defaults to image v0.0.7. Cloud commands resolve workspaces, projects and databases without IDs, and a global --json flag gives machine-readable output. helix --quiet query now prints the result.
  • Local --disk mode now uses SeaweedFS, because MinIO withdrew its community images. Existing MinIO disk data is not moved automatically; see Migrate MinIO disk data.
  • TypeScript SDK 3.1.0 exposes HelixError.code, statusCode, serverMessage, rawBody, and explicit retryable metadata. npm 3.0.4 did not expose these fields. Legacy response formats and unknown future codes remain readable.
  • Transaction conflicts can be recognized with error.statusCode === 409 and error.code === "transaction_conflict". Docker v0.0.5 already supplies this contract; no Docker or CLI release is required. A lost response remains distinct from a confirmed abort, and mutations are never automatically replayed.
  • This release updates the TypeScript SDK only. Embedded runtime packages and the Rust, Python, and Go SDK versions are unchanged.
  • Go SDK v0.3.1 — operation-tree query construction, typed parameters, /v2/query execution, index lifecycle requests, and traversal-scoped vector and BM25 search. Upgrade from v0.3.0: this patch keeps unreleased embedded adapters out of public module discovery. Embedded execution and native graph bindings are not distributed in this release.
  • Python async client — helix-db 0.3.3 adds reusable HTTPX connection pooling, concurrent requests, cancellation-safe response cleanup, and asynchronous embedded writer and reader clients.
  • Open-source database — the engine moved out of helix-hyperscale into the public, modular helix-db workspace.
  • New query planner — a dedicated logical and physical planning layer optimizes the typed operation-tree AST before execution.
  • Embedded database — run the same engine and queries in process with memory, disk, or object storage.
  • Vector prefiltering — traverse and filter an exact candidate set before vector ranking, so results cannot escape graph or permission boundaries.
  • Row bindings — bind + projectBindings / projectDistinctBindings correlate values captured at different hops of a single traversal.
  • Python SDK — new zero-dependency synchronous SDK (helix-db, imported as helixdb).
  • CLI 3.0.6 — --path flag on helix add.
  • Go SDK
  • Better cache metrics for better query insights
  • Query plan improvements
  • helix chef command for local bootstrap
  • CLI DX improvements
  • Local docker deployment
  • Query plan improvements
  • Query insights, metrics and suggestions
  • Stability and reliability improvements
  • Better memory usage and caching

Launch 🚀

  • Graph database with vector search and full-text search
  • Fully ACID
  • Queries via Rust DSL
  • Backed by object storage
  • Multi-tenancy included by default