Replay: re-run the code against the history — determinism and the non-determinism error
On restart the engine re-runs your workflow from the top and hands back each recorded result from history instead of re-executing the step, so it fast-forwards to exactly where the crash hit — but only if the code is deterministic and issues the same steps every time.
We have a durable history; replay re-runs the code against it, handing back recorded results so ORDER #1001 fast-forwards to step 2 without re-charging. But that only works if the code is deterministic — and charging a card, calling a payment API, sending an email are all non-deterministic side effects. Where do those live so replay doesn't repeat them?
Scene 03
Replay: re-run the code against the history
- Watch
- Try it
- Predict
- Capture
You restart after the crash with the durable history from last scene, but a record isn't a running program. So the engine does something clever: it re-runs your workflow function from the very top — and for every step that already has a recorded result, it hands your code that recorded value INSTEAD of doing the step again. So ChargeCard returns its recorded "ok" without touching the card a second time. The head fast-forwards step by step until it reaches the first step with no recorded result yet — ReserveInventory — which now runs for real. This re-run-against-history move is called replay: re-executing the code from the start, substituting each step's recorded result, to fast-forward back to exactly where the crash hit. Drag the scrub head across the strip and watch the gray 'replayed' steps fast-forward to the live one.
Highlighted lines are the ones running in the diagram right now.
def replay(workflow_fn, history):cursor = 0 # position in recorded historyrun(workflow_fn):on command issued by code:cmd = next command from codeapply(cmd, history, cursor)cursor += 1# history exhausted -> caught up; go liveresume_live_execution()
def apply(cmd, history, cursor):if cursor < len(history):event = history[cursor]assert_matches(cmd, event) # determinism gate# NOT re-executed: ChargeCard is not re-chargedreturn event.recorded_resultelse:# first un-recorded step -> execute for realresult = execute_activity(cmd)history.append(record(cmd, result))return result
def assert_matches(cmd, event):# the re-run must issue the SAME command it did beforeif cmd == event.command:return # reconciled, continue replay# Math.random()/clock returned a different value,# so the code branched off the recorded pathraise NonDeterminismError(expected = event.command,got = cmd,) # workflow freezes -- no safe merge
Where this sits in Build a workflow engine (Temporal / Airflow / Cadence style)
Scene 03 of 13, in the Replay act — Rebuild state by replay; quarantine in activities.. To resume, the engine re-runs your code from the top and hands back the recorded results instead of redoing them — which only works if the code is deterministic.
Up next. Replay must never re-charge the card. The only way to do that is to move every side effect out of the replayed workflow code and into a special boundary whose result gets recorded — so on replay the engine hands back the recorded result instead of doing the effect again. That boundary is the activity.
All 13 scenes in Build a workflow engine (Temporal / Airflow / Cadence style) · Every curriculum