The L4 pinning trap: balance requests, not connections
Because HTTP/2 keeps one long-lived connection, an L4 load balancer balances CONNECTIONS not requests — so every RPC pins to a single backend; the fix is client-side load balancing that discovers all backends and picks one per call.
The interceptor chain handed our call to the transport — but the very property that made HTTP/2 great, ONE long-lived connection, means a connection-level load balancer pins all our calls to one unlucky backend.
Scene 09
The L4 pinning trap: balance requests, not connections
- Watch
- Try it
- Predict
- Capture
A quick refresher: an RPC (a function call dressed up to run on another machine) rides exactly one HTTP/2 stream, and many streams share ONE long-lived TCP connection — that single-connection trick is what made HTTP/2 fast. Now put a load balancer in front. The kind here is an L4 balancer (it works at the TCP layer — it sees connections, not the individual calls inside them), like a typical cloud ELB or nginx in TCP mode. Watch: the client opens ONE connection, the L4 LB sends that connection to pod #3, and then every greet("Ada") the client fires rides that same pinned connection straight to pod #3. The other three pods stay dark. When one backend gets stuck with all the work because the connection never moves, that's connection pinning — defined below.
Highlighted lines are the ones running in the diagram right now.
def dial(target, policy):if policy == "pick_first":# one VIP address -> the L4 LB picks the backendaddrs = [resolve_one(target)]else: # round_robin# headless service / xDS returns every backend IPaddrs = resolve_all(target)# one subchannel == one long-lived HTTP/2 connectionself.subchannels = [Subchannel(a) for a in addrs]self.next = 0
def pickSubchannel():if policy == "pick_first":# only one subchannel exists -> always the samereturn self.subchannels[0]# round_robin: rotate per call across all subchannelssc = self.subchannels[self.next % len(self.subchannels)]self.next += 1return sc
def send(method, request):sc = pickSubchannel()# one RPC == one HTTP/2 stream on sc's connectionstream = sc.conn.new_stream()stream.send_headers(method)stream.send_data(encode(request))return await stream.recv_response()
Where this sits in Build a gRPC-style RPC framework
Scene 09 of 14, in the Resilience act — Client-side balancing, backpressure, head-of-line blocking.. HTTP/2's one long-lived connection means an L4 load balancer pins every RPC to one backend. Client-side balancing discovers all backends and picks one per request.
Up next. Spreading calls across backends keeps any one pod from drowning in REQUESTS — but a single streaming RPC can still drown one pod in DATA if a fast sender outruns a slow reader.
All 14 scenes in Build a gRPC-style RPC framework · Every curriculum