Build a coordination service (ZooKeeper / etcd style)

No scenes authored for this problem yet.

About Build a coordination service (ZooKeeper / etcd style)

Raft alone isn't enough. Build the service layer on top of Raft that the rest of your infrastructure depends on: a hierarchical namespace, watches, ephemeral nodes, leases, and the recipe library (locks, leader election, queues) that everyone reimplements badly. Internalize why Kubernetes, Kafka, and Consul all run on something shaped like this.

Difficulty
advanced
Time
about 90 minutes
Stages
9
Topic
Consensus, Coordination & Durable Execution

How this problem is worked

Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.

  1. 01Purpose & invariantsWhat is this for, and what must always be true of it?
  2. 02Workload characterizationWho writes, who reads, and in what shapes?
  3. 03Data model & on-disk formatWhat does the data look like at rest?
  4. 04Core algorithmsHow do the write path and the read path actually work?
  5. 05Distribution & replicationHow does this scale out and survive losing a machine?
  6. 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
  7. 07Failure modes & recoveryWhat actually happens when each part fails?
  8. 08Operational characteristicsCan a human run this at three in the morning?
  9. 09Trade-offs & comparisonWhere does this sit against the alternatives?

Primary sources for this problem

  • Burrows — The Chubby lock service for loosely-coupled distributed systems (OSDI 2006)
  • Hunt et al. — ZooKeeper: wait-free coordination for Internet-scale systems (USENIX ATC 2010)
  • Junqueira & Reed — ZAB: A simple totally ordered broadcast protocol (LADIS 2008)
  • etcd docs — Architecture, gRPC API, watch implementation
  • ZooKeeper recipes (zookeeper.apache.org/doc/current/recipes.html)
  • Reed, Junqueira — Apache ZooKeeper (book chapter)

More in Consensus, Coordination & Durable Execution

Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.

Browse the full problem catalog, or see what the simulator does and does not model.