The L4 pinning trap: balance requests, not connections

Because HTTP/2 keeps one long-lived connection, an L4 load balancer balances CONNECTIONS not requests — so every RPC pins to a single backend; the fix is client-side load balancing that discovers all backends and picks one per call.

Previously

The interceptor chain handed our call to the transport — but the very property that made HTTP/2 great, ONE long-lived connection, means a connection-level load balancer pins all our calls to one unlucky backend.

Scene 09

The L4 pinning trap: balance requests, not connections

  1. Watch
  2. Try it
  3. Predict
  4. Capture
gRPC client1 HTTP/2 conn0 RPCsone connectionL4 LB (ELB/nginx)balances CONNECTIONSpick_first→ pinned to 1 podBACKEND PODS · 4pod #10% load0 RPCs0pod #20% load0 RPCs0pod #3100% load0 RPCsPINNEDpod #40% load0 RPCs0One HTTP/2 connection pinned to pod #3 — every greet("Ada") lands there.
What to watch for

A quick refresher: an RPC (a function call dressed up to run on another machine) rides exactly one HTTP/2 stream, and many streams share ONE long-lived TCP connection — that single-connection trick is what made HTTP/2 fast. Now put a load balancer in front. The kind here is an L4 balancer (it works at the TCP layer — it sees connections, not the individual calls inside them), like a typical cloud ELB or nginx in TCP mode. Watch: the client opens ONE connection, the L4 LB sends that connection to pod #3, and then every greet("Ada") the client fires rides that same pinned connection straight to pod #3. The other three pods stay dark. When one backend gets stuck with all the work because the connection never moves, that's connection pinning — defined below.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Channel.dial
what the channel resolves and connects to at startup
def dial(target, policy):
if policy == "pick_first":
# one VIP address -> the L4 LB picks the backend
addrs = [resolve_one(target)]
else: # round_robin
# headless service / xDS returns every backend IP
addrs = resolve_all(target)
# one subchannel == one long-lived HTTP/2 connection
self.subchannels = [Subchannel(a) for a in addrs]
self.next = 0
Channel.pickSubchannel
which connection this one call rides
def pickSubchannel():
if policy == "pick_first":
# only one subchannel exists -> always the same
return self.subchannels[0]
# round_robin: rotate per call across all subchannels
sc = self.subchannels[self.next % len(self.subchannels)]
self.next += 1
return sc
Channel.send
the call opens one HTTP/2 stream on the picked connection
def send(method, request):
sc = pickSubchannel()
# one RPC == one HTTP/2 stream on sc's connection
stream = sc.conn.new_stream()
stream.send_headers(method)
stream.send_data(encode(request))
return await stream.recv_response()

Where this sits in Build a gRPC-style RPC framework

Scene 09 of 14, in the Resilience act — Client-side balancing, backpressure, head-of-line blocking.. HTTP/2's one long-lived connection means an L4 load balancer pins every RPC to one backend. Client-side balancing discovers all backends and picks one per request.

Up next. Spreading calls across backends keeps any one pod from drowning in REQUESTS — but a single streaming RPC can still drown one pod in DATA if a fast sender outruns a slow reader.

All 14 scenes in Build a gRPC-style RPC framework · Every curriculum

Built with Arqly
Every scene in Build a gRPC-style RPC framework builds on the one before it.All 14 Build a gRPC-style RPC framework scenes