
HydraDB - fast graph database on object storage
HydraDB is an object-store-native distributed graph database written in Rust. It combines durable graph storage on SlateDB with snapshot-consistent OpenCypher queries, GraphBLAS traversal, Neo4j-compatible Bolt connectivity, and an HTTPS query API.
Storage and compute are fully disaggregated. S3-compatible object storage is the
durable source of truth, and compute runs as two independent roles: data
nodes (graph-node) serve queries and canonical mutations, while indexers
(graph-indexer) build immutable traversal indexes in the background. Both keep
only disposable state in memory and on local SSD or NVMe, so they can be replaced
or scaled without moving the graph itself.
…
Each data node owns a private local SSD/NVMe cache; the object store is the shared layer beneath the whole tier and the only durable copy of the graph.
Data nodes serve reads and canonical graph mutations. Indexer workers build immutable CSC generations asynchronously and publish them through atomic object-store pointers. Readers remain correct when an index is absent or behind because the visible WAL tail is applied to the indexed base.
See architecture.md for the storage model, query pipeline, writer coordination, index lifecycle, and failure semantics.
There are two ways to bring up a single development node: the published Docker image, or a build from source. Either way, once the node is listening, use Verify a running node to confirm it works — a listening port is not proof; a round-tripped write is. TLS is required by default in deployed environments, so the local flows below enable plaintext explicitly.
Run with Docker — fastest, no local toolchainRelease images are published to
ghcr.io/hydra-db/hydradb.
Each v* release is tagged with its full version, compatible minor and major
versions, the commit SHA, and latest (for example 0.1.0, 0.1, 0,
latest, sha-7bf77ac):
docker pull ghcr.io/hydra-db/hydradb:latest
Images are published for linux/amd64 and linux/arm64, so Docker selects the
right one for the host and Apple Silicon needs no extra flags. Releases up to
and including 0.1.0 were linux/amd64 only, and pulling one of those on an
ARM host fails with:
no matching manifest for linux/arm64/v8 in the manifest list entries
That message means the tag predates multi-architecture publishing, not that the
pull is misconfigured. Move to a release after 0.1.0, or run the older tag
under emulation with --platform linux/amd64 — correct but slower, and it
requires Rosetta on Apple Silicon. To see which architectures a tag actually
carries before pulling it:
docker buildx imagetools inspect ghcr.io/hydra-db/hydradb:latest
This starts one plaintext node backed by a host directory mounted into the container:
…
The node runs in the foreground. LOCAL_PATH must point at a directory that
already exists, which is why hydradb-data/store is created before the mount.
--user "$(id -u):$(id -g)" is required: the image runs as UID/GID 10001,
but the bind-mounted hydradb-data is owned by the host user, so without it the
container cannot write its store or cache and fails on the first storage
operation. Running as the host user makes the mounted directories writable and
keeps the created files host-owned. The image entrypoint is graph-node; it
also ships graph-indexer. For production, pin an image digest rather than
latest — see the Helm chart guide.
HydraDB requires Rust 1.91 or newer, a C/C++ toolchain,
libcypher-parser, and SuiteSparse GraphBLAS.
Ubuntu or WSL:
sudo apt-get update
sudo apt-get install -y \
build-essential clang libclang-dev cmake pkg-config \
libcypher-parser-dev libgraphblas-dev \
curl git python3 python3-venv
The last line is not needed to build, but the steps below use it: curl for
the Rust installer and the readiness checks, git to clone, and python3-venv
for the Neo4j driver used by scripts/runtime_smoke.sh.
macOS with Homebrew:
xcode-select --install
brew install just cmake pkg-config llvm suite-sparse
brew install cleishm/neo4j/libcypher-parser
# Rust, only if `rustup toolchain list` does not already show a stable toolchain:
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
libcypher-parser is not in homebrew-core; the fully-qualified
cleishm/neo4j/... name adds the tap automatically. A plain
brew install libcypher-parser fails with No available formula.
Rust comes from the official installer rather than Homebrew because the
rustup formula is keg-only and no longer ships a rustup-init binary, so
brew install rustup leaves nothing named rustup on PATH.
rust-toolchain.toml pins channel = "stable", so any rustup-managed stable
toolchain works.
No PKG_CONFIG_PATH export is needed: libcypher-parser is not keg-only, so
Homebrew links cypher-parser.pc into the default pkg-config search path.
just is the supported command runner for the
repository. Install it with cargo install just --locked when your package
manager does not provide it. Docker is optional and is used only by MinIO,
Neo4j comparison, image-build, and Kubernetes harnesses.
git clone https://github.com/hydra-db/hydradb.git
cd hydradb
just native-check
just smoke
The smoke example creates a local graph, writes and deletes edges, runs a sparse
traversal, closes the database, reopens it, and verifies the durable result.
The recipe creates and removes an isolated local object-store directory. Use
just smoke-graphblas to pin the traversal kernel to SuiteSparse GraphBLAS.
To exercise the same flow against an ephemeral MinIO instance:
just minio-smoke
The following starts a single plaintext development node backed by a local directory.
…
The node runs in the foreground and does not return; that is it working, not hanging. Confirm it from a second shell with Verify a running node.
For a fully scripted Bolt and HTTP round trip against a source build, install the
Python Neo4j driver and run. Homebrew's and Debian's Python both refuse a bare
pip install under PEP 668, so use a virtualenv (apt-get install -y python3-venv on Debian/Ubuntu):
python3 -m venv /tmp/hydradb-venv && /tmp/hydradb-venv/bin/pip install neo4j
# macOS: this script calls cargo directly, so it does not inherit what the
# justfile exports. Without this it fails at bindgen with
# `'cypher-parser.h' file not found`. Linux needs neither.
if command -v brew >/dev/null; then
export BINDGEN_EXTRA_CLANG_ARGS="-I$(brew --prefix)/include"
export LIBRARY_PATH="$(brew --prefix)/lib"
fi
PYTHON=/tmp/hydradb-venv/bin/python bash scripts/runtime_smoke.sh
Prints runtime-smoke-ok. The node's log is at
/tmp/sgk-runtime-smoke/node.log; read it first if the script fails.
However you started it, the node listens on:
| Endpoint | Address | Purpose |
|---|---|---|
| Bolt | 127.0.0.1:7687 |
Neo4j-driver-compatible queries |
| HTTP | 127.0.0.1:8443 |
JSON and NDJSON query API |
| Admin | 127.0.0.1:9090 |
readiness and Prometheus metrics |
In another terminal, write and read a small graph through HTTP:
TOKEN='local-development-token-32-bytes'
curl -sS http://127.0.0.1:8443/v1/graphs/default/query \
-H "Authorization: Bearer $TOKEN" \
-H 'X-Graph-Namespace: default' \
-H 'Content-Type: application/json' \
--data '{"cell_id":"cell-0","query":"CREATE (a {id: 1})-[:FOLLOWS]->(b {id: 2})"}'
curl -sS http://127.0.0.1:8443/v1/graphs/default/query \
-H "Authorization: Bearer $TOKEN" \
-H 'X-Graph-Namespace: default' \
-H 'Content-Type: application/json' \
--data '{"cell_id":"cell-0","query":"MATCH (a {id: 1})-[:FOLLOWS]->(b) RETURN b.id AS id"}'
The second call returns one row containing
{"type":"vertex_id","value":2}. A listening port is not proof the node works;
a round-tripped write is.
| Symptom | Cause and fix |
|---|---|
No available formula with the name "libcypher-parser" |
Use the tap: brew install cleishm/neo4j/libcypher-parser |
command not found: rustup-init |
Homebrew's rustup is keg-only and no longer ships it; use the official installer above |
invalid environment variable CLOUD_PROVIDER value \null`` |
CLOUD_PROVIDER is unset — null means absent, not the string. local also needs LOCAL_PATH, pointing at a directory that already exists |
wrapper.h:4:10: fatal error: 'cypher-parser.h' file not found |
BINDGEN_EXTRA_CLANG_ARGS unset while invoking cargo directly on macOS. Prefer just, which exports it |
Node answers /readyz, then aborts with has overflowed its stack on the first query |
RUST_MIN_STACK unset; export 33554432 |
curl: (7) Failed to connect ... port 9090 |
The node is not running. graph-node holds the foreground, so start it in its own shell |
Agents working in this repository should read AGENTS.md, which carries the same sequence plus repository conventions and failure modes. Contributors building HydraDB should also read DEVELOPMENT.md for the full recipe, harness, and script surface.
HydraDB supports a practical OpenCypher subset for graph reads and mutations,
including typed relationships, bounded variable-length paths, property and
label predicates, ordering, pagination, aggregation, OPTIONAL MATCH, UNION,
and batched UNWIND writes.
Applications can connect with a Neo4j driver using a routed URI:
neo4j://127.0.0.1:7687
Use neo4j+s:// with a publicly trusted certificate or neo4j+ssc:// for a
self-signed development certificate. Direct bolt:// node addresses are for
diagnostics and targeted failure tests; write-capable clustered clients should
use routing.
HydraDB includes native snapshot-scoped path procedures:
algo.SPpaths finds bounded paths between one source and one target.algo.SSpaths finds bounded paths from one source.algo.MSpaths resolves many indexed source and target values and evaluates
them together, avoiding client-side query fan-out.CALL algo.MSpaths({
sourceLabel: 'Entity',
sourceProperty: 'name',
sourceValues: ['alpha', 'beta', 'gamma'],
targetValues: ['alpha', 'beta', 'gamma'],
pairwise: true,
relTypes: ['RELATES'],
relDirection: 'both',
m
No open issues yet, or sync has not completed.