rate() first, then sum() — counter resets and aggregation across pods
Because a pod's counter drops to zero when it restarts, you compute each series' per-second rate first, which detects and repairs the drop, and only then sum across pods, since summing first hides the reset and draws a false spike.
The monitor now collects every pod's running totals; turning them into one fleet line needs the right order of operations.
Scene 05
rate() first, then sum()
- Watch
- Try it
- Predict
- Capture
How do you turn ever-growing counters into 'requests per second across the fleet' without the graph lying? Three checkout pods keep running totals of requests served (example values around 1,000, 1,200 and 800). A total that only climbs doesn't say how busy a pod is right now, so watch each pod's climb get turned into requests per second over a short window, then watch the three pods' numbers get added into one line for the whole checkout service.
Highlighted lines are the ones running in the diagram right now.
http_requests_total = 0def on_request():http_requests_total += 1def on_scrape():return http_requests_totaldef restart():http_requests_total = 0
def rate(series, window):samples = series.in_window(window) # >= 4x scrapeif len(samples) < 2:return Noneincrease = 0for prev, cur in pairs(samples):if cur < prev:increase += curelse:increase += cur - prevreturn increase / span_seconds(samples)
def rate_then_sum(pods, window):# sum(rate(http_requests_total[1m]))per_pod = []for p in pods:r = rate(p.series, window)if r is None:return Noneper_pod.append(r)return sum(per_pod)
def sum_then_rate(pods, window):# rate(sum(http_requests_total)[1m:15s])combined = []for t in scrape_times:total = sum(p.series.at(t) for p in pods)combined.append(total)return rate(combined, window)
Where this sits in Metrics / Monitoring System
Scene 05 of 18, in the Ask act — rate() before sum(); add buckets, never average p99s.. A pod's counter drops to zero when it restarts, so compute each series' per-second rate first and sum after: summing the raw counters first hides the reset and draws a false spike.
Up next. Now that request counts combine honestly with rate-then-sum, the next question is whether latency percentiles combine the same way across pods.
All 18 scenes in Metrics / Monitoring System · Every curriculum