Rate limiting — the token bucket

A token bucket holds N tokens refilling at R/s; each request consumes one and an empty bucket returns 429. Local enforcement is cheap but drifts under autoscaling (cap × sidecars); global enforcement holds the fleet cap at the configured number but adds an RPC hop.

Previously

Ejecting a bad replica handles broken backends; the next failure shape is a HEALTHY backend being asked to do more than it can. The proxy needs a way to say 'no, slow down' — that policy is rate limiting.

Scene 09

Rate limiting — the token bucket

  1. Watch
  2. Try it
  3. Predict
  4. Capture
SCOPE · LOCAL (per-sidecar bucket)sidecarcheckout-sc-1req+1 / 0.50sR = 2 rpsbucket · checkout-sc…3 / 5 tokensCONFIGURED PER-BUCKET1,000 rps× 1 sidecars (independent)fleet cap driftsEFFECTIVE FLEET RPS1,000 rpsOne proxy, one bucket — configured cap and fleet cap are the same number.
↓ token — one allowed request
← faucet drips at the refill rate R
What to watch for

One sidecar, one bucket — 3 of 5 tokens right now. The faucet drips +1 token every 0.5s (R = 2 rps). The green arrow is a request that grabbed a token and passed, and each pass costs the bucket one token. Drain a bucket to empty and the next request bounces back a 429 instead — you'll watch that happen on the scope slider. That policy — capping the accept rate — is rate limiting; the mechanism on screen — N tokens + refill — is the token bucket.

Implementation

Highlighted lines are the ones running in the diagram right now.

TokenBucket.on_request
the per-bucket filter: refill, then take one or 429
bucket = { tokens: CAPACITY, last_refill: now() }
def on_request():
elapsed = now() - bucket.last_refill
bucket.tokens = min(
CAPACITY,
bucket.tokens + elapsed * RATE,
)
bucket.last_refill = now()
if bucket.tokens >= 1:
bucket.tokens -= 1
return ALLOW
return DENY_429 # bucket empty
Local scope — each proxy keeps its own bucket
fleet cap = configured cap × sidecar count (drift)
# configured: 1000 rps per proxy
# 1 sidecar -> effective fleet = 1000 rps
# 3 sidecars -> effective fleet = 3000 rps
# 15 sidecars -> effective fleet = 15000 rps
def on_request(): # runs in every sidecar
return bucket.on_request() # no RPC hop
Global scope — one shared bucket on a side service
every request makes one gRPC hop to the RL service
def on_request(): # runs in every sidecar
resp = grpc.call(
RATE_LIMIT_SERVICE, # e.g. Lyft ratelimit / Redis
descriptors = [
('client_id', req.client_id),
],
)
if resp.code == OVER_LIMIT:
return DENY_429
return ALLOW
# fleet cap == configured cap, regardless of sidecar count
Choosing scope
the one-line rule of thumb
# Local: cheap, no extra hop;
# drifts under autoscaling (cap * sidecars).
# Good for coarse 'be polite' spike absorbers.
# Global: exact aggregate; adds 1 RPC per request and a
# dependency whose outage matters at every hop.
# Required for per-tenant SLA / quota enforcement.

Where this sits in Build a Service Mesh (Envoy / Istio style)

Scene 08 of 13. Per-client token bucket: each request takes a token, an empty bucket returns 429. Local is cheap and drifts; global stays exact via a coordinator.

Up next. Saying no to extra requests defends capacity. But before the proxy can rate-limit or authorize, it needs to know WHO each request is from — workload identity, not just an IP.

All 13 scenes in Build a Service Mesh (Envoy / Istio style) · Every curriculum

Built with Arqly
Every scene in Build a Service Mesh (Envoy / Istio style) builds on the one before it.All 13 Build a Service Mesh (Envoy / Istio style) scenes