Outlier detection — eject one bad replica

Outlier detection ejects a single misbehaving replica from the load-balancing pool based on the real traffic it serves, while an active health check independently probes a known endpoint on a schedule — you need both, because each catches a failure the other can't see.

Previously

A circuit breaker tripping the whole cluster is too coarse when only replica R3 is broken. The proxy needs a finer move: pull just R3 out of the pool.

Scene 08

Outlier detection — eject one bad replica

  1. Watch
  2. Try it
  3. Predict
  4. Capture
SERVICE CLUSTER · 4 REPLICASLoad Balancerround-robin · in-pool onlyPool Manageroutlier detectionpassivereal-traffic 5xxactive/healthz probesmode: both0/5 5xxR1traffic: ok/healthz: pass/healthz0/5 5xxR2traffic: ok/healthz: pass/healthz5/5 5xxR3traffic: 5xx/healthz: passEJECTED/healthzeject 7.0s0/5 5xxR4traffic: ok/healthz: pass/healthzboth: passive catches R3's real-traffic 5xx; active would catch a replica with no traffic.
passive outlier detection — real traffic 5xx accumulation
active health check — scheduled /healthz probe
ejected for 10s, then re-admitted
What to watch for

A cluster of 4 replicas. R3 is the bad apple: it returns 200 on /healthz but 500s on real /checkout traffic. Watch the red dots above R3 climb to 5 — that's the consecutive-5xx threshold — then R3 grays out and the LB stops routing to it. The active probe arrow keeps pinging /healthz on every replica; R3's probe never went red.

Implementation

Highlighted lines are the ones running in the diagram right now.

PoolManager.on_real_response
passive outlier detection — trips on consecutive 5xx of real traffic
def on_real_response(replica, response):
if response.status >= 500:
replica.consecutive_5xx += 1
else:
replica.consecutive_5xx = 0
if replica.consecutive_5xx >= consecutive_5xx_threshold:
eject(replica,
for=base_ejection_time * replica.ejection_count)
replica.ejection_count += 1
# max_ejection_percent guards against ejecting everyone.
PoolManager.probe_loop
active health check — periodic /healthz probe on every replica
every interval seconds:
for replica in cluster.replicas:
resp = http.get(
replica.address + '/healthz',
timeout=2s,
)
if resp.status != 200:
eject(replica)
else:
readmit_if_ejected(replica)
Why both signals coexist
the failures each one misses on its own
# Active alone misses:
# /healthz returns 200, but /checkout returns 500.
# (probe path is fine; business path is broken.)
# Passive alone misses:
# a replica with no real traffic yet has zero
# 5xx samples — but its probe will show it failing.
# Use both. They catch disjoint failures and share
# one eject() path into the load-balancing pool.

Where this sits in Build a Service Mesh (Envoy / Istio style)

Scene 07a of 13. Don't trip the whole cluster — pull just the misbehaving replica from the pool. Passive (real 5xx) catches what active /healthz probes miss.

Up next. Ejecting a misbehaving replica protects the cluster from itself, but not from a flood of perfectly well-formed requests arriving faster than the cluster can serve them — next, rate limiting.

All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum

Built with Arqly
Every scene in Build a Service Mesh (Envoy / Istio style) builds on the one before it.All 13 Build a Service Mesh (Envoy / Istio style) scenes