29
Distributed Job Scheduler
Cron at scale: leader, exactly-once, DAGs.SavedSaved on this device — Saved on this device
01Clarifications
What would you ask before drawing a single box?
Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.
AI staff engineer
Enter to send · Shift+Enter for a new line
About Distributed Job Scheduler
Cron at scale: leader, exactly-once, DAGs.
- Difficulty
- intermediate
- Time
- about 50 minutes
- Stages
- 10
- Topic
- Consensus, Coordination & Durable Execution
How this problem is worked
Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.
- 01ClarificationsWhat would you ask before drawing a single box?
- 02Functional reqsWhat must this system actually do?
- 03Non-functionalWhat must it promise about speed, uptime and correctness?
- 04Capacity estimationHow much load and data does this have to hold?
- 05API designWhat does the outside world call, and what comes back?
- 06Data modelWhat gets stored, and what is it looked up by?
- 07Use-case breakdownHow does each requirement actually get served?
- 08High-level designWhich components handle a request, and in what order?
- 09Deep divesWhich part breaks first, and what do you do about it?
- 10Trade-offsWhat did this design cost, and what breaks at 10×?
Primary sources for this problem
- Reliable Cron Across the Planet (Google SRE)
More in Consensus, Coordination & Durable Execution
Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.
- Build Build Raft — consensus you can defendReplicate a deterministic state machine across N servers with safety as a theorem and liveness under partial synchrony. Build the protocol from term to commit to safety proof to reads, and feel why etcd, Cockroach, and TiKV ship slightly different Rafts.
- Build Build a workflow engine (Temporal / Airflow / Cadence style)A function that survives crashes, restarts, and re-deploys — and still finishes. Build a durable execution engine where workflow code is replayed deterministically from an event history, activities retry with exponential backoff, sagas compensate on failure, and the same workflow definition runs identically a year later. Internalize why 'just retry the cron job' breaks at the second step.
- Distributed LockRedlock controversy, fencing tokens, lease vs lock.
- Build Build a production AI agent (from one API call up)A model API is a function from text to text — it remembers nothing, does nothing and fails without saying so. Everything an agent actually does lives in the program around it: the loop, the tools, the window, the budgets, the gates, the sandbox, the run log. Build that program, one mechanism at a time, until it can resolve a support ticket and issue a refund exactly once.
- AI Agent PlatformLong-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.
Browse the full problem catalog, or see what the simulator does and does not model.