Fifty services, fifty broken retry policies

When every service re-implements retry, timeout, breaker, mTLS, and tracing in its own language, the network becomes non-uniform and a single wobbly dependency cascades into a fleet-wide outage.

Scene 01

Fifty services, fifty broken retry policies

  1. Watch
  2. Try it
  3. Predict
  4. Capture
FLEET CALL GRAPH12 services · 12 different retry/timeout libraries · S7 just got slowDISTINCT RETRY POLICIES12hover boxesS1checkoutokhttp retry=3healthyS2cartaxios retry=5 exphealthyS3inventorygrpc-go defaulthealthyS4pricingfetch no retryhealthyS5searchrequests retry=∞healthyS6rankingnet/http no retryhealthyS7profilerest-template defaultSLOWdownS8sessionfeign retry=5healthyS9recshttp.client retry=3healthyS10notifyktor retry=2healthyS11billingnode-fetch retry=4healthyS12ledgeruplink no retryhealthystage 0 · S7 is slow; everyone else looks fine
What to watch for

Look at the badges. Twelve services, twelve different retry libraries — okhttp, axios, requests with unbounded retries, fetch with none. S7 in the middle just got slow (its box is red, tagged SLOW). Its three callers — S2, S5, S9 — feed straight into it. Nothing has cascaded yet.

Implementation

Highlighted lines are the ones running in the diagram right now.

Team checkout (Java) — retry in the app
OkHttp, exponential backoff, three tries — somebody's flavor
Response call_profile(Request req):
for attempt in 1..3:
try:
return okhttp.newCall(req).execute()
except IOException:
sleep((2 ** attempt) * 100ms)
raise UpstreamFailed()
# no shared budget · no breaker · no deadline
Team search (Python) — retry in the app
requests, bare except, infinite retries — different shop
def call_profile(req):
while True:
try:
return requests.get(req.url, timeout=None)
except Exception:
continue # try again forever
# no backoff · no cap · no jitter · no timeout
Team recs (Node) — retry in the app
axios, fixed retry=5, no timeout — common footgun
async function callProfile(req) {
return axios.request({
url: req.url,
// timeout: undefined // forgot to set one
'axios-retry': { retries: 5 },
})
}
// fires 5 extra calls at a peer that's already slow
After the mesh — what the app keeps
the app becomes trivial · policy lives somewhere else
async function callProfile(req) {
// localhost · the same in every language
return fetch('http://localhost/profile' + req.path)
}
# retry / timeout / breaker / mtls / tracing —
# owned by the thing sitting next to the app

Where this sits in Build a Service Mesh (Envoy / Istio style)

Scene 01 of 13. Every team picks its own retry, timeout, breaker, and mTLS library. One slow dependency turns into a fleet-wide outage.

Up next. If the problem is that fifty app teams cannot agree on how the network should behave, the fix is to take the network out of the app's hands and put it in something all fifty share.

All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum

Built with Arqly
Every scene in Build a Service Mesh (Envoy / Istio style) builds on the one before it.All 13 Build a Service Mesh (Envoy / Istio style) scenes