Fifty services, fifty broken retry policies
When every service re-implements retry, timeout, breaker, mTLS, and tracing in its own language, the network becomes non-uniform and a single wobbly dependency cascades into a fleet-wide outage.
Scene 01
Fifty services, fifty broken retry policies
- Watch
- Try it
- Predict
- Capture
Look at the badges. Twelve services, twelve different retry libraries — okhttp, axios, requests with unbounded retries, fetch with none. S7 in the middle just got slow (its box is red, tagged SLOW). Its three callers — S2, S5, S9 — feed straight into it. Nothing has cascaded yet.
Highlighted lines are the ones running in the diagram right now.
Response call_profile(Request req):for attempt in 1..3:try:return okhttp.newCall(req).execute()except IOException:sleep((2 ** attempt) * 100ms)raise UpstreamFailed()# no shared budget · no breaker · no deadline
def call_profile(req):while True:try:return requests.get(req.url, timeout=None)except Exception:continue # try again forever# no backoff · no cap · no jitter · no timeout
async function callProfile(req) {return axios.request({url: req.url,// timeout: undefined // forgot to set one'axios-retry': { retries: 5 },})}// fires 5 extra calls at a peer that's already slow
async function callProfile(req) {// localhost · the same in every languagereturn fetch('http://localhost/profile' + req.path)}# retry / timeout / breaker / mtls / tracing —# owned by the thing sitting next to the app
Where this sits in Build a Service Mesh (Envoy / Istio style)
Scene 01 of 13. Every team picks its own retry, timeout, breaker, and mTLS library. One slow dependency turns into a fleet-wide outage.
Up next. If the problem is that fifty app teams cannot agree on how the network should behave, the fix is to take the network out of the app's hands and put it in something all fifty share.
All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum