Timeout and retry budget — bounded patience

Retrying without a budget turns one slow backend into a 243× retry storm; a retry budget caps total retry concurrency to a small fraction of normal traffic so retries cannot become the outage.

Previously

Once the proxy picked a replica and the replica is slow, the proxy has to decide how long to wait and whether to try again — and that decision, scaled across the fleet, is where outages are born.

Scene 06

Timeout and retry budget — bounded patience

  1. Watch
  2. Try it
  3. Predict
  4. Capture
EFFECTIVE LOAD ON S5243xretries/hop =3 · budget = offClienthealthyS1healthyS2healthyS3healthyS4healthyS5slow / failing32 retries96 retries2718 retries8154 retries243162 retriesRETRY BUDGET · offno cap — retries can dominate the clusternormal trafficretries67% of cluster traffic is retriesNo budget: each hop retries independently, multiplying load by 3^5 = 243× on S5.
clock = timeout (the bound that fires the retry)
each hop retries independently → 3^5 = 243
retry budget — cap retries as % of normal traffic
What to watch for

S5 is failing. Each upstream hop retries 3 times per attempt. Watch the per-hop counter climb back up the chain — the badge above the diagram lands on 243× (3^5), the load S5 actually sees per single client request.

Implementation

Highlighted lines are the ones running in the diagram right now.

Proxy.sendWithRetry
the naive per-hop retry every proxy runs by default
def send_with_retry(req):
for attempt in range(retries + 1):
resp = try_send(
req, deadline = per_try_timeout,
)
if resp.ok:
return resp
sleep(backoff_with_jitter(attempt))
return ERROR
// Why naive retry storms
the compounding that makes 5 hops × 3 retries = 243x
# Hop 1 retries 3x on failure.
# Hop 2 inherits all of hop 1's load,
# and retries 3x of each failure too.
# Every hop multiplies, not adds.
#
# load_on_tail = retries ^ hops
# = 3 ^ 5
# = 243x per single client request
Cluster.admitRetry
Envoy's token-bucket admission: retry only if below cap
def admit_retry(cluster):
budget = budget_percent / 100 # e.g. 0.20
active = cluster.active_request_count
retrying = cluster.active_retry_count
cap = max(
active * budget,
min_retry_concurrency,
)
if retrying >= cap:
return DROP_RETRY # honor the cap
return ALLOW_RETRY
route.yaml
the config that wires timeout, per-try, and budget together
route:
timeout: 30s # whole call, incl. retries
retry_policy:
retries: 3
per_try_timeout: 5s # bound on one attempt
retry_budget:
budget_percent: 20.0 # cap retries at 20% of traffic
min_retry_concurrency: 3

Where this sits in Build a Service Mesh (Envoy / Istio style)

Scene 06 of 13. Naive multi-hop retries amplify load 243x on a failing backend. A retry budget caps total retries as a fraction of normal traffic so retries can't become the outage.

Up next. A budget caps retries fleet-wide, but it doesn't tell the proxy when to stop talking to one specific replica that's broken.

All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum

Built with Arqly
Every scene in Build a Service Mesh (Envoy / Istio style) builds on the one before it.All 13 Build a Service Mesh (Envoy / Istio style) scenes