Guide
The ghcr.io/helixdb/helixdb:v0.0.9 image runs the standalone HelixDB server.
It supports Linux on amd64 and arm64.
It exposes the same operation-tree request contract used by Helix Cloud at
POST /v2/query; gateway-only Cloud features are not included.
Choose storage
Start with the CLI
macOS and Linux:http://localhost:6969/v2/query.
For an existing project, set tag = "v0.0.9" in its [local.dev] block in
helix.toml before starting it. An explicitly saved tag overrides the CLI default.
Use an existing object store
Place credentials in a project-root.env file or export them:
.env
- S3-compatible
- AWS S3
--s3-allow-http only for a plain HTTP endpoint.--persist stores the resolved configuration in helix.toml. Helix does not delete
externally managed object-store data during stop, restart, or prune.
Run the image directly
Memory mode is selected by leaving bothHELIX_DATA_DIR and S3_BUCKET unset:
Run with native directory persistence
SetHELIX_DATA_DIR to a mounted directory inside the container:
docker stop helixdb. Starting a new container with the
same volume preserves its database. Removing the volume deletes the data.
This mode needs no object store; the CLI’s --disk mode uses SeaweedFS.
The image runs as user and group 65532:65532. Its /var/lib/helix directory
is prepared for that user, including when initializing a new Docker named volume.
If you use a bind mount or an existing volume, its directory must be writable by
that user.
Run with SeaweedFS persistence
This Compose configuration runs SeaweedFS as a single-node S3 store with a static identity, creates thehelix-db bucket on startup, and stores HelixDB
objects in the seaweedfs-data volume. HelixDB starts once a signed request
to the bucket succeeds:
docker-compose.yaml
configs content needs Docker Compose 2.23.1 or later. The SeaweedFS
container is not published to the host. Change the example credentials in all
three places if you expose it.
Run docker compose down to replace the HelixDB container without deleting
the SeaweedFS data. docker compose down -v deletes the persisted database.
For an existing object store, use the same HelixDB environment variables and
omit the SeaweedFS service. For AWS S3, omit the endpoint and HTTP override.
The endpoint must be reachable from inside the container; container
localhost is not
the host machine. DB_PATH is a logical prefix, not the storage directory. To use
native directory persistence, set HELIX_DATA_DIR and mount the directory there.
Add a local disk cache
WithS3_BUCKET, the server always keeps recently read data on local disk, ideally
NVMe, in /var/cache/helix unless HELIX_DISK_CACHE_DIR names another directory.
Vector indexes on S3 need it: once an index outgrows the memory cache, every graph
hop that misses it would otherwise wait on its own S3 request. Mount a volume at
/var/cache/helix so the cache survives restarts; without one, each new container
starts with an empty cache. With HELIX_DATA_DIR, set HELIX_DISK_CACHE_DIR to add
the disk cache. Memory storage rejects the cache variables.
The disk cache needs v0.0.7 or later, and the v0.0.6 image and earlier ignore the
HELIX_DISK_CACHE_* variables. v0.0.7 and v0.0.8 cache S3 on disk only when
HELIX_DISK_CACHE_DIR is set, so the example below sets it to the default./var/cache/helix to the helix service. A new
named volume mounted there is already writable; a bind mount must be owned by
65532:65532. When running as another user or outside the image, set
HELIX_DISK_CACHE_DIR to a directory that user can write. With a read-only root
filesystem, mount a writable volume (for example a tmpfs or emptyDir) at the cache
directory. Use one cache directory per running server and per database; changing
DB_PATH on the same directory leaves the old database’s full-text cache behind.
Half of the disk budget caches object-store files, 3/8 caches database blocks,
and the rest caches full-text search data. With
S3_BUCKET, the object-store cache
also keeps the files the server writes; with HELIX_DATA_DIR they are already on
local disk, so it keeps only the files the server reads. Changing the budget
usually changes the block size, which discards the database-block cache at startup
so it refills from the object store. The object-store cache keeps its files and
evicts down to a smaller budget as it admits new data. Full-text search data is
cached on demand, once searches have used it twice; the server downloads none of it
at startup. At startup and on every admission the full-text cache evicts down to
its share. Each of these trims spares data admitted or recorded as used in the
second before it (use is recorded at most once a minute) and data still open in
searches or the 64 MiB full-text memory cache, so the cache can stay over its share
by that data until the next admission. It never caches a split larger than its
whole share. After a restart, the first search to use each cached split reads and
checksums the whole split, up to 64 MiB, before it answers. Budget at least twice
the data the server reads often: the object-store cache fetches whole file parts
(4 MiB, or smaller under a 2 GiB budget), so once that data outgrows half the
budget, cold reads fetch more than memory-only caches would. The budget must also
fit on the cache’s filesystem: its free space plus what the cache already uses.
When HELIX_DISK_CACHE_BYTES is unset and the 8 GiB default does not fit, startup
fails naming the variable; without a volume the cache would otherwise share, and
could fill, the filesystem of the container runtime. A budget you set that does not
fit only logs a warning at startup.
Once storage opens, the server warms the disk cache in the background: it caches the
database blocks that index the newest files and, with S3_BUCKET, the rows vector
searches read, using at most half the object-store cache. The warm delays neither
startup nor readiness; queries that arrive first read from the object store as usual.
Set HELIX_DISK_CACHE_WARM=off to skip it, and the caches then fill only as queries
read.
The block cache also indexes its disk share in memory, outside
HELIX_DISK_CACHE_MEMORY_BYTES. Once full, the index takes roughly 2–9 MiB of RAM
per GiB of HELIX_DISK_CACHE_BYTES (about 20–70 MiB at the default, 2–9 GiB at
1 TiB), and a restart briefly needs about twice that while it rebuilds the index.
Size the container’s memory limit for this index as well as the memory tier.
The block cache keeps up to 32,768 files open, so the server needs at least 34,792
open files in the worst case (26,600 at the default budget), plus one per open
full-text split. It raises its soft open-file limit to the hard limit at startup.
If the container’s hard limit is below that minimum, startup fails naming
HELIX_DISK_CACHE_BYTES; raise the hard limit with --ulimit nofile=65536:65536.
Lowering the budget is not a reliable fix above about 3 GiB, where the file count
does not fall steadily with it; budgets of 1 GiB or less need about 8,200 or fewer.
Run natively on macOS, the limit is also capped by sysctl kern.maxfilesperproc.
Startup also fails with a message naming the variable when a size is invalid or
out of range, a size is set with HELIX_DATA_DIR but without HELIX_DISK_CACHE_DIR,
the default budget does not fit, the cache directory is not writable or cannot be
locked (some network and FUSE filesystems do not support locks), or another running
server already uses it. Stop the old container before starting its replacement on
the same cache directory.
Stop or inspect
Next steps
Get started
Create a project and run your first query.
Embedded database
Open HelixDB directly inside your process.
Helix Cloud
Move the same application requests to a managed cluster.
Local CLI workflow
Manage instances and raw requests.