Cluster and load balancing — pick one of many
A cluster is the group of replicas behind a destination name, and the cluster's load-balancing policy decides which replica each request goes to — round-robin is 'fair' by request count, not by work, which is why it's the wrong default under heterogeneous latency.
Picking 'one of many replicas' is its own design decision: the route's destination name resolves to a cluster, and inside that cluster the proxy applies a load-balancing policy.
Scene 05
Cluster and load balancing — pick one of many
- Watch
- Try it
- Predict
- Capture
Watch the proxy fan requests across the four replicas behind the checkout-svc cluster. The policy is round-robin — every replica gets the same share. R3 is slow (800 ms vs 50 ms). Watch R3's in-flight tower climb while the others stay flat, and the tail-latency gauge climb with it.
Highlighted lines are the ones running in the diagram right now.
def round_robin(cluster):i = (cluster.cursor + 1) % len(cluster.replicas)cluster.cursor = ireturn cluster.replicas[i]# fair by request COUNT — never reads in_flight,# so a slow replica drains slower than it fills# and its queue grows without bound.
Where this sits in Build a Service Mesh (Envoy / Istio style)
Scene 05 of 13. Behind one destination name is a cluster of replicas; round-robin, least-request, or ring-hash decides who serves each request. Round-robin is the wrong default under heterogeneous latency.
Up next. Picking a replica is only half the question — the other half is what happens when the replica you picked is slow or doesn't answer.
All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum