Retries and exponential backoff
The engine automatically retries a failed activity on a schedule that grows exponentially — 1s, 2s, 4s, 8s — so a transient blip self-heals without hammering a sick downstream, while a permanent error can be flagged non-retryable to fail fast instead.
Retries let a transient 503 self-heal — but retries also mean an activity can RUN more than once. Picture the worst crash yet: the worker charges the card, then dies before recording the result. The engine, seeing no result, retries — and charges again. History and replay can't help here, because the effect already happened outside the recorded boundary. What stops THAT double-charge?
Scene 06
Retries and exponential backoff
- Watch
- Try it
- Predict
- Capture
ORDER #1001 reaches step 1, ChargeCard $42 — and the Payment API returns 503: briefly overloaded, not refusing the card. A naive system would fail the whole order. Instead the engine re-attempts just the ACTIVITY, on a growing schedule: attempt 2 after 1s, attempt 3 after 2s, attempt 4 after 4s. That rule — initial gap, how fast the gap grows, the cap, how many attempts — is the activity's retry policy (read it off the chip top-left). And those gaps don't stay flat; each is bigger than the last, so the re-tries stop pounding a service that's already struggling. Growing the wait between attempts like 1s, 2s, 4s, 8s is exponential backoff (a small random offset, called jitter, is added so thousands of workflows don't all re-fire on the same tick). Watch the blip self-heal on a later attempt — and notice each retry is scheduled as a durable timer+task, so even a crash mid-wait can't lose it.
Highlighted lines are the ones running in the diagram right now.
def executeActivity(activity, policy):for attempt in 1 .. policy.maximumAttempts:result = try_run(activity)if result.ok:return result # blip self-healedif result.errorType in policy.nonRetryableErrorTypes:raise result.error # fail fast, no scheduledelay = policy.nextDelay(attempt)scheduleRetry(activity, after = delay)
def nextDelay(attempt):raw = initialInterval * backoffCoefficient ** (attempt - 1)gap = min(raw, maximumInterval) # cap a single waitreturn gap + jitter() # desync the herd
def scheduleRetry(activity, after):fireAt = now() + afterhistory.append(TimerStarted(fireAt)) # durable# ...engine may crash here; on recovery the timer# is replayed from history and still fires...on fireAt:enqueue(activity, taskQueue) # worker re-runs it
Where this sits in Build a workflow engine (Temporal / Airflow / Cadence style)
Scene 06 of 13, in the Runtime act — Workers pull; retries with backoff; idempotency.. The engine retries a failed activity on its own, widening the gap between attempts so a sick downstream can recover instead of being pinned down by a retry storm.
Up next. Retrying an activity is only safe if running it twice does no extra harm. For a charge, that means the downstream must recognize a repeat and refuse to charge again — by attaching a stable label to the request so the second one is deduplicated. That label is an idempotency key.
All 13 scenes in Build a workflow engine (Temporal / Airflow / Cadence style) · Every curriculum