Build your own workflow engine (Temporal / Airflow / Cadence style)
A function that survives crashes, restarts, and re-deploys — and still finishes. Build a durable execution engine where workflow code is replayed deterministically from an event history, activities retry with exponential backoff, sagas compensate on failure, and the same workflow definition runs identically a year later. Internalize why 'just retry the cron job' breaks at the second step.
- Scenes
- 13 interactive scenes
- Time
- about 91 minutes
- Topic
- Consensus, Coordination & Durable Execution
What you are building, and why
You have written a function that does a few things in order — charge a card, reserve some inventory, ship a package, send an email. You have written a cron job that retries something that failed. That is your starting point. By the end of this curriculum you will have built — conceptually, layer by layer — a Temporal/Cadence-style durable execution engine, and you will be able to reason about a real workflow: where its state lives, what happens to it when the process crashes mid-run, when a retry is safe, how it waits 30 days without holding a thread, and how to change its code a year after it started — and choose, for a given workload, whether a durable engine, a DAG scheduler, or a state machine is the right tool.
We start from one uncomfortable, concrete fact. Scene 1 runs ORDER #1001: ChargeCard($42) → ReserveInventory(1 Widget) → ShipPackage → SendConfirmationEmail. The process crashes right after the card is charged but before the inventory is reserved. A plain function kept its progress (charged = true) only in RAM, and the crash erased it — so a naive cron retry re-runs from the top and charges the card a second time: $84. That is what "just retry the cron job breaks at the second step" actually looks like. Every later mechanism in the stack exists to plug a worse-placed crash, and we re-ask the same question against each one: did we double-charge, and does ORDER #1001 still finish?
This is a curriculum about a strict density-of-vocabulary discipline: across 13 scenes, each scene introduces at most two new terms, and the diagram carries the load. One running example — ORDER #1001 — is threaded through every scene: you watch its progress get wiped from RAM, then written to a durable event history, then replayed to resume at step 2, then quarantined inside activities, then survive a redeploy, a retry storm, a crash inside the charge, a 30-day wait, a cancel signal, a mid-order failure, an endless billing loop, and finally a code change deployed while the order was still sleeping. The same blue-is-workflow / red-is-activity coloring and the same append-only history strip appear in every scene so the arc reads as one continuous story.
Resist the urge to "describe Temporal." Build each layer because the previous layer's crash forced it, feel the failure mode each one prevents, and let the design canvas at the end push back on whether durable execution was even the right tool.
What you will be able to explain afterwards
- durable execution: a function whose progress survives crashes, reboots, and redeploys — not just an in-memory call
- the event history is the source of truth: append every step's result to a durable log, never trust RAM
- event sourcing: rebuild current state by folding the log of events, instead of saving variables directly
- deterministic replay: re-run the workflow code against the recorded history to resume — so the code must be pure
- activities are the boundary to non-determinism: side effects (RPC, DB, time, randomness) run once and the RESULT is recorded
- task queues + workers that PULL: the engine is decoupled from the worker, so a redeploy is just 'no worker for a moment'
- retries with exponential backoff + per-activity policies (max attempts, backoff coefficient, jitter, non-retryable errors)
- idempotency keys: the real defense against the double-charge under at-least-once delivery
- durable timers as first-class events: 'sleep 30 days' survives ten restarts because the timer lives in history, not a thread
- signals & queries: a running workflow is an addressable actor you can message and read
- the saga pattern: compensating actions undo completed steps in reverse on partial failure (refund, not rollback)
- child workflows + ContinueAsNew: compose sub-units and keep an endless workflow's history bounded
- versioning workflows safely: a year-old in-flight execution must still replay identically against today's code
Durability
Why RAM double-charges; the event history.
- 01The crash that charges you twice — durable execution and at-least-once deliveryA plain function keeps its progress in RAM, so a crash after step one erases it — and a naive cron retry re-runs from the top and charges the card a second time.~7 min
- 02The event history is the source of truthStop trusting RAM: append every step's result to a durable, append-only event history, so a crash that wipes memory leaves the record of what already happened intact.~7 min
Replay
Rebuild state by replay; quarantine in activities.
- 03Replay: re-run the code against the history — determinism and the non-determinism errorTo resume, the engine re-runs your code from the top and hands back the recorded results instead of redoing them — which only works if the code is deterministic.~7 min
- 04Activities: quarantine for side effects — commands, events, and the activity boundaryReplay re-runs workflow code, so every side effect must move into an activity whose result is recorded — replay hands the result back instead of charging again.~7 min
Runtime
Workers pull; retries with backoff; idempotency.
- 05Task queues: workers pull, so redeploys are safeThe engine never pushes work; stateless workers pull tasks from a queue, so a redeploy is just 'no worker for a moment' and the task simply waits to be picked up.~7 min
- 06Retries and exponential backoffThe engine retries a failed activity on its own, widening the gap between attempts so a sick downstream can recover instead of being pinned down by a retry storm.~7 min
- 06aIdempotency keys: the last hole in the double-chargeAn activity can run twice if it succeeds then crashes before recording — a stable idempotency key lets the downstream recognize the repeat and refuse the second charge.~7 min
Time & actors
Durable timers and the workflow as an actor.
- 07Durable timers: sleep 30 days on zero computeA thread that sleeps for a month dies on the first crash; a durable timer records the wait as an event, so the workflow goes dormant until the engine fires the wake-up.~7 min
- 08Signals and queries: the workflow as an actorA running workflow is an addressable actor: a signal delivers external input durably and can change its path; a query reads its state without mutating it.~7 min
Scale & undo
Sagas compensate; children + ContinueAsNew.
- 09The saga: compensate partial failure in reverseYou can't wrap steps across services in one transaction; a saga gives each step an undo and runs the compensators for completed steps in reverse — a refund, not a rollback.~7 min
- 10Child workflows and ContinueAsNewChild workflows isolate sub-units with their own histories, and ContinueAsNew restarts an endless workflow with a fresh history but the same ID before it hits the limit.~7 min
Evolve & ship
Version safely, then choose the right engine.
- 11Versioning: the year-later replay — getVersion gates for in-flight executionsA year-old in-flight execution can't be hot-fixed; a version gate routes old executions down the old path and new ones down the new path, so both replay deterministically.~7 min
- 12Design canvas: choose the right engineMatch each workload to its model — code-as-workflow replay, a DAG scheduler, or a state machine — and defend every choice with the scene that taught the requirement.~7 min
More in Consensus, Coordination & Durable Execution
Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.
- Build Build Raft — consensus you can defendReplicate a deterministic state machine across N servers with safety as a theorem and liveness under partial synchrony. Build the protocol from term to commit to safety proof to reads, and feel why etcd, Cockroach, and TiKV ship slightly different Rafts.
- Distributed LockRedlock controversy, fencing tokens, lease vs lock.
- Build Build a production AI agent (from one API call up)A model API is a function from text to text — it remembers nothing, does nothing and fails without saying so. Everything an agent actually does lives in the program around it: the loop, the tools, the window, the budgets, the gates, the sandbox, the run log. Build that program, one mechanism at a time, until it can resolve a support ticket and issue a refund exactly once.
- AI Agent PlatformLong-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build a workflow engine (Temporal / Airflow / Cadence style) workspace