Idempotency keys: the last hole in the double-charge
History and replay stop the workflow from re-issuing the charge, but a retried activity can still charge the card a second time when the crash lands between 'money moved' and 'result recorded' — so the caller attaches a stable label the downstream deduplicates on.
Retries let a transient 503 self-heal — but a retry also means an activity can RUN more than once. Picture the worst crash yet: the worker charges the card, then dies before recording the result. The engine, seeing no result, retries — and charges again. History and replay can't help, because the effect already happened outside the recorded boundary. What stops THAT double-charge?
Scene 06a
Idempotency keys: the last hole in the double-charge
- Watch
- Try it
- Predict
- Capture
We already moved every side effect into an activity, so replay never re-charges the card. But zoom all the way into the ChargeCard activity itself. It has three moments: the request is sent, the money actually moves at the card network, and finally the result is recorded back in history. The danger lives in the gap between the last two. If the worker crashes AFTER the money moved but BEFORE the result was recorded — the CRASH WINDOW, shaded red — the engine sees no recorded result and does exactly what we taught it to: it retries the activity. The money already moved once; the retry moves it again. This is the one hole history and replay cannot reach, because the effect happened outside the recorded boundary. The banner above the strip names the deal you actually get: the engine gives you effectively-once workflow LOGIC — your code's command is never re-issued — but activity EFFECTS are at-least-once, meaning the real-world charge can apply more than once. The only lever that closes this is the chip on the request: an idempotency key — a stable label the caller attaches (order-1001-charge) so the receiver can recognize a repeat and refuse to charge twice. Drop the skull into the red window and watch the statement.
Highlighted lines are the ones running in the diagram right now.
def runActivity(activity, input):while True:try:result = activity(input) # money moves herehistory.record(result) # result recorded herereturn resultexcept NoResultReported:# crash landed after the call, before recordsleep(backoff()) # then retry the call
def charge_card(order):# key is derived from data that survives a crash,# so the retry carries the SAME label as the first trykey = idempotency_key(order) # 'order-1001-charge'return payment_api.charge(amount = order.total,idempotency_key = key, # null when off)
def charge(amount, idempotency_key):if idempotency_key in dedup_store:# repeat recognized: return the first outcomereturn dedup_store[idempotency_key]result = card_network.move(amount) # applies effectif idempotency_key is not None:dedup_store[idempotency_key] = resultreturn result
Where this sits in Build a workflow engine (Temporal / Airflow / Cadence style)
Scene 06a of 13, in the Runtime act — Workers pull; retries with backoff; idempotency.. An activity can run twice if it succeeds then crashes before recording — a stable idempotency key lets the downstream recognize the repeat and refuse the second charge.
Up next. The crash window between 'money moved' and 'result recorded' was the last double-charge hole, and an idempotency key plugged it. We've now survived crashes at every step of a FAST order. But real orders aren't fast — ORDER #1001 might wait 7 days for a 'rate this product' email. A thread that sleeps for a week dies on the first crash in that week. How does a workflow wait days and survive every restart in between?
All 13 scenes in Build a workflow engine (Temporal / Airflow / Cadence style) · Every curriculum