23
Build Raft — consensus you can defend
Replicate a deterministic state machine across N servers with safety as a theorem and liveness under partial synchrony. Build the protocol from term to commit to safety proof to reads, and feel why etcd, Cockroach, and TiKV ship slightly different Rafts.SavedSaved on this device — Saved on this device
AI staff engineer
Enter to send · Shift+Enter for a new line
About Build Raft — consensus you can defend
Replicate a deterministic state machine across N servers with safety as a theorem and liveness under partial synchrony. Build the protocol from term to commit to safety proof to reads, and feel why etcd, Cockroach, and TiKV ship slightly different Rafts.
- Difficulty
- advanced
- Time
- about 100 minutes
- Stages
- 9
- Topic
- Consensus, Coordination & Durable Execution
How this problem is worked
Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01Purpose & invariantsWhat is this for, and what must always be true of it?
- 02Workload characterizationWho writes, who reads, and in what shapes?
- 03Data model & on-disk formatWhat does the data look like at rest?
- 04Core algorithmsHow do the write path and the read path actually work?
- 05Distribution & replicationHow does this scale out and survive losing a machine?
- 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
- 07Failure modes & recoveryWhat actually happens when each part fails?
- 08Operational characteristicsCan a human run this at three in the morning?
- 09Trade-offs & comparisonWhere does this sit against the alternatives?
Primary sources for this problem
- Ongaro & Ousterhout — In Search of an Understandable Consensus Algorithm (USENIX ATC 2014)
- Diego Ongaro — Consensus: Bridging Theory and Practice (Stanford PhD thesis, 2014)
- etcd-io/raft — design doc + raft.go state machine
- TiKV — Implementing Raft in Rust (PingCAP blog series)
- hashicorp/raft — README + protocol notes
- raft.github.io — canonical visualization
More in Consensus, Coordination & Durable Execution
Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.
- Build Build a workflow engine (Temporal / Airflow / Cadence style)A function that survives crashes, restarts, and re-deploys — and still finishes. Build a durable execution engine where workflow code is replayed deterministically from an event history, activities retry with exponential backoff, sagas compensate on failure, and the same workflow definition runs identically a year later. Internalize why 'just retry the cron job' breaks at the second step.
- Distributed LockRedlock controversy, fencing tokens, lease vs lock.
- Build Build a production AI agent (from one API call up)A model API is a function from text to text — it remembers nothing, does nothing and fails without saying so. Everything an agent actually does lives in the program around it: the loop, the tools, the window, the budgets, the gates, the sandbox, the run log. Build that program, one mechanism at a time, until it can resolve a support ticket and issue a refund exactly once.
- AI Agent PlatformLong-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.
Browse the full problem catalog, or see what the simulator does and does not model.