You can't average a p99
Per-pod percentiles can't be combined, but histogram buckets are plain counters that add up across pods, so the percentile is estimated after summing, with bucket boundaries setting accuracy and each bucket costing a series.
Counts combine by rate then sum; latency is the harder case, because a percentile isn't a count.
Scene 06
You can't average a p99
- Watch
- Try it
- Predict
- Capture
How slow is checkout for the unluckiest users, across many pods? An average latency won't say — the fast requests drag it to the middle. Watch three checkout pods appear: pod-a and pod-b answer in about 100 ms, pod-c in about 2 s. Each pod reports the latency that 99 of every 100 of its own requests beat — that number is its p99, a percentile — and beside the three badges sits the fleet's own p99, computed over all 1,000 requests. Read those two numbers against each other before you continue.
Highlighted lines are the ones running in the diagram right now.
http_duration = Histogram("http_request_duration_seconds",buckets=LE_BOUNDS,)def observe(seconds):for le in LE_BOUNDS:if seconds <= le:_bucket[le] += 1 # le="0.5"_bucket["+Inf"] += 1_sum += seconds_count += 1
def fleet_quantile(q):total = {}for le in LE_BOUNDS + ["+Inf"]:total[le] = sum(rate(p._bucket[le]) for p in pods)rank = q * total["+Inf"]lo, seen = 0, 0for le, count in sorted(total.items()):if count >= rank:frac = (rank - seen) / (count - seen)return lo + (le - lo) * fraclo, seen = le, count
def pod_p99(p):xs = sorted(p.recent_latencies)return xs[int(0.99 * len(xs))]def average_of_p99s():badges = [pod_p99(p) for p in pods]return sum(badges) / len(badges) # len(badges) = 3
Where this sits in Metrics / Monitoring System
Scene 06 of 18, in the Ask act — rate() before sum(); add buckets, never average p99s.. Per-pod percentiles can't be combined, but histogram buckets are plain counters that add up across pods, so the fleet percentile is estimated after summing — and bucket boundaries set its accuracy.
Up next. Now that we can compute honest error rates and latency across the fleet, the next question is how one of those numbers becomes a page without paging on every blip.
All 18 scenes in Metrics / Monitoring System · Every curriculum