Design your graph database
A graph database is a stack of trades you choose deliberately — linked-list vs CSR, native vs graph-on-KV, supernode mitigations, ACID-on-one-primary vs distributed predicate-sharding, and traversal-local vs whole-graph workloads — and the same primitives configure radically different systems depending on which one the workload demands.
You've felt every force: index-free adjacency makes local traversal fly, the supernode and the partition cut break it, ACID anchors correctness on one machine, and Pregel inverts the payoff for whole-graph work. The capstone is choosing deliberately for a real workload — every knob traceable to the scene that justified it, with the local-vs-global spine as your compass.
Scene 14
Design your graph database
- Watch
- Try it
- Predict
- Capture
Here is everything the arc built, in one place. Four workload cards are docked across the top — fraud-ring traversal, knowledge graph, social feed, nightly PageRank — and each card states its real constraints: does the graph mutate or sit read-mostly, does it fit one box or sprawl across machines, are there celebrity supernodes, and does a query touch a few nodes or every node. Down the side is the palette: storage layout (mutable doubly-linked chains vs compact CSR), the index anchor you SEEK before you EXPAND, supernode handling, consistency (single-primary ACID vs distributed), distribution (single node, edge-cut, vertex-cut, predicate sharding), and the workload model (on-demand local traversal vs whole-graph Pregel). Every one of those knobs was earned in an earlier scene. The job now is not to learn anything new — it's to read each workload off the local-vs-global spine and pick the honest configuration. The same toolkit; four different right answers.
- docRobinson, Webber & Eifrem — Graph Databases (O'Reilly, 2nd ed)
- blogRelationship Chain Locks: Don't Block the Rock (Neo4j)
- blogPowerGraph: distributed graph-parallel computation (the morning paper)
- docDgraph design concepts — minimizing network calls
- docPregel: A System for Large-Scale Graph Processing (Malewicz et al.)
Where this sits in Build a graph database (Neo4j / Dgraph-style)
Scene 14 of 16, in the The payoff act — No locality for whole-graph work — then design it.. Every graph-DB deployment is a deliberate set of choices — storage layout, index strategy, supernode handling, single-node ACID vs distributed, and traversal-heavy vs aggregate-heavy workload — and the right configuration for a fraud-ring traversal is wrong for a PageRank pipeline even though the primitives are identical.
All 16 scenes in Build a graph database (Neo4j / Dgraph-style) · Every curriculum