Durable timers: sleep 30 days on zero compute
A workflow sleep records a TimerStarted event and the workflow goes fully dormant on zero compute until the engine fires TimerFired at the deadline — so a long wait survives any number of crashes and redeploys, because the timer is an event in history, not a sleeping thread.
We've now survived crashes at every step of a FAST order. But real orders aren't fast — ORDER #1001 must wait 7 days for the 'rate this product' email. A thread that sleeps for a week dies on the first crash in that week. So the engine records the wait as a durable event in the history and the workflow goes completely dormant, using zero compute, until the engine fires the wake-up event. Time becomes a first-class durable object: the timer.
Scene 07
Durable timers: sleep 30 days on zero compute
- Watch
- Try it
- Predict
- Capture
ORDER #1001 has shipped. Now it must wait 7 days before sending the "rate this product" email. The obvious way — call sleep(7 days) — holds a real worker process for the whole week, which is what the TOP model shows. But the engine has a better way. When your workflow code asks to wait, the engine doesn't block a thread; it appends a TimerStarted event to ORDER #1001's history and the workflow box goes completely dark: dormant — holding no worker and burning essentially zero compute, because the only thing tracking the deadline now is the engine itself. When the countdown hits zero, the engine appends a TimerFired event and wakes the workflow via the task queue. That whole mechanism — a wait recorded as a TimerStarted/TimerFired event pair that the engine owns — is a durable timer: time is stored as an event in history, not held in a sleeping thread. Watch the durable model start the wait and go dormant, then fire on schedule.
Highlighted lines are the ones running in the diagram right now.
def waitForRating(order):ship(order)# BROKEN: held thread, countdown lives in RAMsleep(days=7) # dies on first crash# DURABLE: yields, recording a TimerStarted eventawait workflow.sleep(days=7) # returns control; goes dormantsend_rating_email(order)
def onTimerStarted(wf_id, duration):fire_at = now() + durationhistory.append(wf_id, TimerStarted(fire_at))schedule.add(wf_id, fire_at) # engine owns the countdownunload(wf_id) # dormant: no worker held
def fireTimers(): # also runs on recovery after a crashfor wf_id, fire_at in schedule.due(now()):history.append(wf_id, TimerFired())task_queue.put(wf_id) # wake itreplay(wf_id) # resumes right after the sleep
Where this sits in Build a workflow engine (Temporal / Airflow / Cadence style)
Scene 07 of 13, in the Time & actors act — Durable timers and the workflow as an actor.. A thread that sleeps for a month dies on the first crash; a durable timer records the wait as an event, so the workflow goes dormant until the engine fires the wake-up.
Up next. Waiting on a clock is one thing; reacting to the outside world is another. A running workflow is a long-lived addressable thing you can send messages to — a fire-and-forget signal that changes its path (cancel the order), or a read-only query that peeks at its state. The workflow becomes an actor.
All 13 scenes in Build a workflow engine (Temporal / Airflow / Cadence style) · Every curriculum