Design your graph database

A graph database is a stack of trades you choose deliberately — linked-list vs CSR, native vs graph-on-KV, supernode mitigations, ACID-on-one-primary vs distributed predicate-sharding, and traversal-local vs whole-graph workloads — and the same primitives configure radically different systems depending on which one the workload demands.

Previously

You've felt every force: index-free adjacency makes local traversal fly, the supernode and the partition cut break it, ACID anchors correctness on one machine, and Pregel inverts the payoff for whole-graph work. The capstone is choosing deliberately for a real workload — every knob traceable to the scene that justified it, with the local-vs-global spine as your compass.

Scene 14

Design your graph database

  1. Watch
  2. Try it
  3. Predict
  4. Capture
ACTIVEFraud-ring traver…real-time multi-hop · constantl…fraud-traversalKnowledge graphbillions of typed edges · type-…knowledge-graphSocial feedcelebrity supernodes · follow-f…social-feedNightly PageRankwhole-graph scan · read-mostly …pagerank-batchDESIGN PALETTEStorage layoutdoubly-linked chains (…CSR array (compact / i…Index strategyB-tree anchor on the l…full-text (Lucene) anc…no anchor index (label…Supernode han…expand the low-degree …dense-node relationshi…relationship-chain loc…Consistencysingle-primary ACID (l…distributed / weaker i…Distributionsingle node (no partit…edge-cut partitionvertex-cut partitionpredicate sharding (ed…Workload modelon-demand local traver…whole-graph Pregel / B…VERIFIER✓FITSdoubly-linked chains (mutable / O…↳ Doubly-linked chains take cheap edge inserts on a … · scene 6.✓FITSB-tree anchor on the lookup prope…↳ You still SEEK the flagged account once via an index before… · scene 5.✓FITSsingle-primary ACID (logical log)↳ A mutating fraud graph on one box gets full ACID with no … · scene 10.✓FITSsingle node (no partition)↳ It fits one big box, so every hop stays a pointer dereference — no… · scenes 3, 11.✓FITSon-demand local traversal↳ k-hop 'who-transacted-with-whom' lights a few nodes — the local … · scene 3.Configuring: Fraud-ring traversal — real-time multi-hop · constantly mutating · fits one big box
four workloads, one toolkit
the spine is your compass: local lights a few · global lights all
every knob traces back to a scene →
What to watch for

Here is everything the arc built, in one place. Four workload cards are docked across the top — fraud-ring traversal, knowledge graph, social feed, nightly PageRank — and each card states its real constraints: does the graph mutate or sit read-mostly, does it fit one box or sprawl across machines, are there celebrity supernodes, and does a query touch a few nodes or every node. Down the side is the palette: storage layout (mutable doubly-linked chains vs compact CSR), the index anchor you SEEK before you EXPAND, supernode handling, consistency (single-primary ACID vs distributed), distribution (single node, edge-cut, vertex-cut, predicate sharding), and the workload model (on-demand local traversal vs whole-graph Pregel). Every one of those knobs was earned in an earlier scene. The job now is not to learn anything new — it's to read each workload off the local-vs-global spine and pick the honest configuration. The same toolkit; four different right answers.

Where this sits in Build a graph database (Neo4j / Dgraph-style)

Scene 14 of 16, in the The payoff act — No locality for whole-graph work — then design it.. Every graph-DB deployment is a deliberate set of choices — storage layout, index strategy, supernode handling, single-node ACID vs distributed, and traversal-heavy vs aggregate-heavy workload — and the right configuration for a fraud-ring traversal is wrong for a PageRank pipeline even though the primitives are identical.

All 16 scenes in Build a graph database (Neo4j / Dgraph-style) · Every curriculum

Built with Arqly
Every scene in Build a graph database (Neo4j / Dgraph-style) builds on the one before it.All 16 Build a graph database (Neo4j / Dgraph-style) scenes