Circuit breaker — the state machine
A circuit breaker is a three-state machine — closed, open, half-open — that fast-fails requests to a broken dependency and periodically probes it for recovery, so the caller stops waiting on timeouts it can already predict will fail.
A retry budget caps how much extra load the fleet generates, but each individual caller is still spending its own time and connections waiting for a backend it could already tell is broken. What's missing is a switch that says 'stop calling this dependency for a while' — and that switch is the circuit breaker.
Scene 07
Circuit breaker — the state machine
- Watch
- Try it
- Predict
- Capture
Watch the breaker trip. Errors climb past the threshold; CLOSED hands off to OPEN; new requests fast-fail in milliseconds. After a cooldown, HALF-OPEN admits exactly one probe — and the result decides whether the breaker closes or stays open.
Highlighted lines are the ones running in the diagram right now.
state = CLOSEDerror_rate_window = sliding(60s)open_started_at = 0def on_request(req):if state == CLOSED:outcome = forward(req, timeout=T)error_rate_window.record(outcome)if error_rate_window.error_rate() > THRESHOLD:state = OPENopen_started_at = now()return outcomeif state == OPEN:if now() - open_started_at > COOLDOWN:state = HALF_OPENelse:return fast_fail_503() # ~1ms, no backend callif state == HALF_OPEN:outcome = forward(req, timeout=T) # single PROBEstate = CLOSED if outcome.ok else OPENopen_started_at = now()return outcome
# Without breaker: every caller waits the full timeout.# N callers * T seconds = N*T thread-seconds blocked.# Caller's thread pool fills with stuck requests.# With breaker OPEN: every caller returns in ~1ms.# Caller frees the resource and degrades gracefully.# Traffic to OTHER (healthy) dependencies keeps flowing.
# Envoy's `circuit_breakers` config is connection-pool# ceilings per cluster, not closed/open/half-open.# Per-replica state-machine semantics live in# outlier_detection (next scene).cluster:circuit_breakers:max_connections: 1024max_pending_requests: 1024max_requests: 1024max_retries: 3
Where this sits in Build a Service Mesh (Envoy / Istio style)
Scene 07 of 13. Closed → open → half-open. Fast-fail to a known-broken dependency and periodically probe for recovery, so callers stop wasting resources on guaranteed failures.
Up next. A breaker per cluster decides 'is the dependency broken' as a whole — but a cluster usually has many replicas, and the bad apple is just one of them. That finer-grained decision is the next scene.
All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum