Retries: idempotency and a token budget
You can only safely retry a method whose re-execution is harmless (idempotent), and even then you must cap retries with a token-bucket budget — or blind retries amplify load on a struggling service into a self-sustaining storm.
We learned to abort doomed work with deadlines and cancellation; now the mirror problem — work that FAILED and might be worth re-sending. But scene 1 warned we often can't tell whether the first attempt already ran, so a retry is a loaded gun.
Scene 07
Retries: idempotency and a token budget
- Watch
- Try it
- Predict
- Capture
The backend is browning out — slow and dropping calls. With no budget and 3 retries per failed call, every client re-sends at once. Watch the OFFERED-LOAD meter on the backend climb. Partway through, the original slowness clears (the trigger turns off) — but the load DOESN'T drop, because the retries are now feeding themselves. That self-sustaining overload, where the service stays down after its own cause is gone, is a retry storm: blind retries pile on exactly when a service can least afford it. This scene is about the two things that make retrying safe.
Highlighted lines are the ones running in the diagram right now.
def callWithRetry(method, req):attempt = 0while attempt < maxAttempts: # slider: retries + 1status = send(method, req)if status == OK:budget.onSuccess() # refill tokenRatioreturnif status not in retryableStatusCodes:raise # e.g. not UNAVAILABLEbudget.onFailure() # drain one tokenif not budget.allow(): # bucket below halfraisesleep(backoffWithJitter(attempt))attempt += 1
tokens = maxTokens # full bucketdef onFailure():tokens = max(0, tokens - 1)def onSuccess():tokens = min(maxTokens, tokens + tokenRatio)def allow():if not enabled:return True # no budget: never pausereturn tokens >= maxTokens / 2
# greet is idempotent: re-running returns the same valuedef greet(name):return 'hello ' + name# charge is NOT: each call moves moneydef charge(name, amount):account[name].balance -= amount # replay double-billsreturn receipt()# the leak: a retry can't tell if the first attempt ran
Where this sits in Build a gRPC-style RPC framework
Scene 07 of 14, in the Reliability act — Deadlines, cancellation, retries, and the interceptor onion.. Only an idempotent method is safe to auto-retry, and even then a token-bucket budget must cap retries — or a brownout turns into a self-sustaining retry storm.
Up next. Retries, deadlines, auth, metrics — these all have to wrap EVERY call, and copy-pasting them into every method is madness; we need one place where such cross-cutting concerns compose as layers.
All 14 scenes in Build a gRPC-style RPC framework · Every curriculum