The saga: compensate partial failure in reverse
Because you can't wrap steps across separate services in one transaction, a saga gives each step a compensating action that semantically undoes it, and on a mid-order failure it runs the compensations for completed steps in reverse order — a refund, not a rollback to a prior byte-state.
The saga safely backs ORDER #1001 out of partial work, refunding and releasing in reverse. But we've kept ORDER #1001 as ONE giant workflow with one growing history — and real orders split across warehouses, run for months, and bill forever. A single history can't grow without bound, and one monolithic workflow is hard to reason about. How do we compose and how do we run forever?
Scene 09
The saga: compensate partial failure in reverse
- Watch
- Try it
- Predict
- Capture
ORDER #1001 has four steps that live on four different services: ChargeCard ($42), ReserveInventory (one Widget), ShipPackage, then SendConfirmationEmail. Watch them stack into a tower as each one completes — first the charge clears, then the Widget is reserved. Then ShipPackage can't complete: the warehouse is out of stock, permanently. Now you're stuck in the worst place: you've already taken $42 and reserved a Widget, but you can't ship — and there is no single database transaction wrapping these four separate services that you could simply roll back. Notice the ghost block sitting next to each completed step — RefundCard beside ChargeCard, ReleaseInventory beside ReserveInventory. That's a sequence of steps where every step carries its own undo. The systems-design name for modeling a long operation this way — local steps, each with a compensating action — is a saga (Garcia-Molina & Salem, 1987): the practical substitute for a distributed transaction when you can't get one.
Highlighted lines are the ones running in the diagram right now.
def run(order):done = [] # completed steps, in orderfor step in [Charge, Reserve, Ship, Email]:try:step.do(order) # an activity on its servicedone.append(step)except StepFailed:compensate(done) # reverse-undo what completedraise
def compensate(done):for step in reversed(done): # backward recoveryif step.compensator is None:continue # no clean undo (Ship/Email)compensateStep(step.compensator)
def compensateStep(comp):delay = 1.0 # initialIntervalwhile True:try:comp.do(idempotency_key) # semantic undo, dedup-safereturn # append Compensated(comp)except ActivityFailed:sleep(delay)delay *= 2.0 # backoffCoefficient
Where this sits in Build a workflow engine (Temporal / Airflow / Cadence style)
Scene 09 of 13, in the Scale & undo act — Sagas compensate; children + ContinueAsNew.. You can't wrap steps across services in one transaction; a saga gives each step an undo and runs the compensators for completed steps in reverse — a refund, not a rollback.
Up next. Sagas handle one order failing, but a real order can split across warehouses and a subscription can bill forever. The answers are composition — a parent workflow starting child workflows it awaits — and a way to keep an endless workflow's history from growing without bound by restarting it fresh while keeping its ID. That restart is ContinueAsNew.
All 13 scenes in Build a workflow engine (Temporal / Airflow / Cadence style) · Every curriculum