You can't average a p99

Per-pod percentiles can't be combined, but histogram buckets are plain counters that add up across pods, so the percentile is estimated after summing, with bucket boundaries setting accuracy and each bucket costing a series.

Previously

Counts combine by rate then sum; latency is the harder case, because a percentile isn't a count.

Scene 06

You can't average a p99

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Latency per pod, bucketedpod-ap99 ≈ 100 ms498 requests (49.8%)390.0545720.250.751.52.5+Infpod-bp99 ≈ 100 ms497 requests (49.7%)390.0545620.250.751.52.5+Infpod-cp99 ≈ 2 s5 requests (0.5%)0.050.250.7511.542.5+Infaverage the three p99 badgesFleet: bucket counts added across podsbar height ∝ √countfast podsslow pod780.059130.140.250.50.75111.5422.55+Inflele = bucket upper bound in seconds; +Inf catches everything slowerp99 estimatenot used: only the p99 badges are averagedFleet p99, three waysaverage of pods' p99sshown as answer≈ 733 msestimated from summed buckets≈ 100 mstrue fleet p99, all 1,000 reques…ground truth≈ 100 msSERIES COSTexample default layout13 series per pod (11 buckets + _sum + _count)= 39 series across 3 podsAveraging the pods' p99s ignores how many requests each pod actually served.
What to watch for

How slow is checkout for the unluckiest users, across many pods? An average latency won't say — the fast requests drag it to the middle. Watch three checkout pods appear: pod-a and pod-b answer in about 100 ms, pod-c in about 2 s. Each pod reports the latency that 99 of every 100 of its own requests beat — that number is its p99, a percentile — and beside the three badges sits the fleet's own p99, computed over all 1,000 requests. Read those two numbers against each other before you continue.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

App.observe
every request bumps each bucket whose bound it beats
http_duration = Histogram(
"http_request_duration_seconds",
buckets=LE_BOUNDS,
)
def observe(seconds):
for le in LE_BOUNDS:
if seconds <= le:
_bucket[le] += 1 # le="0.5"
_bucket["+Inf"] += 1
_sum += seconds
_count += 1
Query.fleet_quantile
histogram_quantile(0.99, sum by (le) (rate(_bucket[5m])))
def fleet_quantile(q):
total = {}
for le in LE_BOUNDS + ["+Inf"]:
total[le] = sum(rate(p._bucket[le]) for p in pods)
rank = q * total["+Inf"]
lo, seen = 0, 0
for le, count in sorted(total.items()):
if count >= rank:
frac = (rank - seen) / (count - seen)
return lo + (le - lo) * frac
lo, seen = le, count
Query.average_of_p99s
each pod's own p99, divided by how many pods there are
def pod_p99(p):
xs = sorted(p.recent_latencies)
return xs[int(0.99 * len(xs))]
def average_of_p99s():
badges = [pod_p99(p) for p in pods]
return sum(badges) / len(badges) # len(badges) = 3

Where this sits in Metrics / Monitoring System

Scene 06 of 18, in the Ask act — rate() before sum(); add buckets, never average p99s.. Per-pod percentiles can't be combined, but histogram buckets are plain counters that add up across pods, so the fleet percentile is estimated after summing — and bucket boundaries set its accuracy.

Up next. Now that we can compute honest error rates and latency across the fleet, the next question is how one of those numbers becomes a page without paging on every blip.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes