Persistence & Durability
Persistence & Durability
JetGraph is an in-memory engine — all data lives in RAM for sub-millisecond access. Durability is provided by a layered persistence model that writes to disk asynchronously without pausing query processing.
How Data Survives Restarts
Full Snapshots
Complete graph serialized to disk nightly (default 02:30 AM) or on a configurable interval. Compressed with zstd.
Delta Files
Incremental changes written every 30–60 seconds. Applied on top of the base snapshot at startup for fast recovery.
Checkpoints
When the delta chain grows long, deltas are merged offline into a single checkpoint file — no engine pause required.
Shutdown Delta
On graceful stop, a final delta is written before the full snapshot — so recent changes are safe even if the snapshot is interrupted.
Recovery Sequence at Startup
-
1
Load latest full snapshot
The most recent
snapshot-*.binfile is deserialized into memory. The engine is marked ready immediately so queries can start while the Cuckoo filter rebuilds in the background. -
2
Apply checkpoint (if any)
If a checkpoint file exists for this snapshot base, it is applied first to skip replaying the earliest deltas.
-
3
Replay remaining delta files
Any delta files newer than the checkpoint are applied in order, bringing the graph fully up to date.
/health while loading a snapshot. Polls will succeed once loading completes — typically within seconds for small graphs, longer for very large ones. This is why start_period in the health check should be set generously.Emergency Snapshot (Memory Pressure)
When RSS memory exceeds memory_limit_bytes in config.toml, the engine automatically writes an emergency full snapshot, blocks new ingest, and logs an error. This protects data before the container OOM-killer fires. Ingest resumes once memory drops back below the limit.
Manual Snapshot
Trigger a snapshot on demand via Cypher without restarting:
Graceful Shutdown
Send SIGTERM (what docker compose stop / docker compose down sends) and the engine will write a final delta then a full snapshot before exiting. The stop_grace_period: 300s in the compose file gives it up to 5 minutes for very large graphs. Never use SIGKILL directly — it bypasses the shutdown snapshot.
Storage Modes — Full Memory vs Mixed
JetGraph ships one binary with two storage modes, selected by the STORAGE_MODE environment variable (or the mode field in config.toml). The default is full-memory mode: every row lives in RAM. Mixed mode keeps hot data in RAM and pages cold data to an NVMe page store, so a single node can serve graphs larger than RAM. Durability is identical in both modes — snapshots plus deltas, as described above. The page store is a cache, not the source of truth.
Full memory (default)
All forward rows, histograms, and properties stay resident. Lowest latency, simplest operations. Right for evaluation, small graphs, and any graph that fits comfortably in RAM.
Mixed (STORAGE_MODE=mixed)
Three page classes (forward rows, histograms, properties) share a RAM budget and evict cold entries to NVMe pages; reads fault them back on demand. Right for production graphs approaching or beyond RAM size.
Where Each Mode Keeps Data
| Component | Full memory | Mixed |
|---|---|---|
| Forward rows (edges) | RAM | Hot set in RAM, cold rows on NVMe pages (always pageable) |
| Inverse index (edge → sources) | RAM | RAM always — never paged, ~8 bytes per edge pair plus overhead. This is the capacity ceiling of one node. |
| Node records, external IDs, masks | RAM | RAM (unpaged) |
| Node histograms (velocity timelines) | RAM | RAM unless STORAGE_PAGE_HISTOGRAMS=true (default false) |
| Node properties | RAM | RAM unless STORAGE_PAGE_PROPERTIES=true (default false) |
| Durability | Snapshots + deltas | Snapshots + deltas (page store is a rebuildable cache) |
The Three Page Classes and the Budget Split
In mixed mode, STORAGE_MEMORY_BUDGET_BYTES (default: 25% of memory_limit_bytes) is the RAM budget for all materialized pages, split between three classes:
- forward — edge rows. Always paged; receives whatever the other two classes do not take.
- histogram — velocity timelines. Initial share
STORAGE_HISTOGRAM_BUDGET_PCT(default 20%). Only active whenSTORAGE_PAGE_HISTOGRAMS=true. - property — node property rows. Initial share
STORAGE_PROPERTY_BUDGET_PCT(default 40%). Only active whenSTORAGE_PAGE_PROPERTIES=true.
Shares drift slowly toward the class under the most refault pressure (bounded by a 10% floor and 90% ceiling), so the split self-tunes to the workload. A sweep runs every STORAGE_SWEEP_INTERVAL_SECS seconds (default 10): above the high watermark (default 90% of budget) it evicts cold entries to pages until usage drops below the low watermark (default 75%). Because eviction only triggers under budget pressure, a class that fits in its share stays fully resident — paged_entries 0 in the metrics then means "fits in RAM", not "paging is broken".
Storage Reference (Environment Variables)
An unset variable keeps the config.toml value; a malformed value fails startup instead of being silently ignored. Confirm the effective config in the startup log (storage env override … lines followed by Config loaded: storage.…).
| Variable | Default | What it does |
|---|---|---|
STORAGE_MODE |
full_memory |
full_memory or mixed. The only switch most deployments need. |
STORAGE_PAGE_STORE_DIR |
<snapshot_dir>/pages |
NVMe page store (segment files + manifest + indexes). Must be a persistent volume; wiping it forces a cold rebuild from snapshot + deltas. |
STORAGE_MEMORY_BUDGET_BYTES |
25% of memory_limit_bytes |
Total RAM budget for materialized pages across all three classes. |
STORAGE_HIGH_WATERMARK_PCT / STORAGE_LOW_WATERMARK_PCT |
90 / 75 |
Eviction starts above the high share of the budget and stops below the low share. |
STORAGE_PAGE_HISTOGRAMS / STORAGE_HISTOGRAM_BUDGET_PCT |
false / 20 |
Page velocity timelines; initial budget share in percent. Prefer src-side histograms on 1–2 edge types — dst bins saturate at 255. |
STORAGE_PAGE_PROPERTIES / STORAGE_PROPERTY_BUDGET_PCT |
false / 40 |
Page node property rows; initial budget share in percent. |
STORAGE_PAGE_SIZE_BYTES |
8192 |
Uncompressed page target. 8 KiB faults ~1.7× faster than 32 KiB for point lookups at ~40% more disk. |
STORAGE_SWEEP_INTERVAL_SECS |
10 |
Base eviction-sweep interval; shortens automatically under pressure. |
STORAGE_COLD_RESTORE |
true |
Restore snapshot rows straight into pages (bounded RAM) instead of RAM on cold start. |
STORAGE_WARM_RESTART |
true |
Reuse the page store across restarts when its checkpoint matches the snapshot (rows=0-style installs, no page reads). Off = wipe and rebuild every start. |
STORAGE_CHECKPOINT_INTERVAL_SECS |
3600 |
Hourly checkpoint flushes dirty rows, writes the row/histogram indexes, and runs segment GC. A checkpoint is also written at shutdown — restarts minutes after a bulk load lean on snapshot + delta replay. At 1B-edge scale raise to 7200–14400 and watch jetgraph_storage_last_checkpoint_phase_ms — a checkpoint passing half its interval logs a warning. |
STORAGE_GC_LIVE_PCT |
50 |
Sealed segments below this live-record share are relocated and dropped at checkpoint. |
Sizing Guidance
- Start full-memory. If the graph fits in RAM with headroom, mixed mode buys nothing.
- The inverse index sets the ceiling. It is never paged (~8 bytes per edge pair plus overhead), so size RAM for it first, then set the page budget from what remains. History that doesn't fit the index doesn't fit one node.
- NVMe is mandatory for mixed mode. Page faults are on the read path; spinning disks will show up directly as query p99.
- Raise
max_nodes/max_edges(defaults 80M / 100M) before bank-scale loads — they cap the ID space, not memory.
Restart Behavior in Mixed Mode
Warm restart (page store intact, checkpoint matches): forward rows and histograms reinstall as sentinels from the row/histogram indexes without reading any page — the log shows warm restart installed the row index (no page was read). Properties restore from the snapshot into RAM and the sweeps page cold rows back out over the following minutes. Cold restart (page store wiped or mismatch): everything rebuilds from snapshot + delta replay. If the engine reports Mixed · degraded on the dashboard, a page fault hit I/O or corruption — rows serve empty until restart; check jetgraph_storage_degraded and the segment GC metrics before restarting.
Verifying Paging Is Working
jetgraph_storage_class_budget_bytesvsjetgraph_storage_class_resident_bytesper class (forward,histogram,property) — the budget split and actual footprint.jetgraph_storage_prop_paged_entries/jetgraph_storage_hist_paged_entries{edge_type,side}— rows currently on disk. Zero with everything resident is healthy.…_evicted_totalrising while…_faults_totalstays flat — the sweep is draining cold data nobody reads. Both rising together — working set churn; watch p99.…_fault_failures_totaland…_writes_dropped_totalmust stay 0.
paged_entries > 0 proves bytes are on disk; evicted_total climbing proves the sweeps are writing them; faults_total climbing proves reads come back intact. See Metrics and Memory & Compression for the surrounding instrumentation.