rate() first, then sum() — counter resets and aggregation across pods

Because a pod's counter drops to zero when it restarts, you compute each series' per-second rate first, which detects and repairs the drop, and only then sum across pods, since summing first hides the reset and draws a false spike.

Previously

The monitor now collects every pod's running totals; turning them into one fleet line needs the right order of operations.

Scene 05

rate() first, then sum()

  1. Watch
  2. Try it
  3. Predict
  4. Capture
TIME TO DETECT≈ 1 minexampleORDERrate then sum: rate each pod, then add the rateswindow 1m = 4× the 15 s scrape intervalwindow 1m = 4× the 15 s scrape intervalpod-1 raw counterpod-1 raw counterclimbs; drops on restartpod-1 rate /spod-1 rate /sraw –rate no datapod-2 raw counterpod-2 raw counterclimbs; drops on restartpod-2 rate /spod-2 rate /sraw –rate no datapod-3 raw counterpod-3 raw counterclimbs; drops on restartpod-3 rate /spod-3 rate /sraw –rate no datacheckout fleetcheckout fleet requests/srequests/scheckout fleet requests/ssum of the per-pod ratessum(rate(http_requests_total[sum(rate(http_requests_total[1m]))1m]))sum(rate(http_requests_total[1m]))01:58:0001:58:3001:59:0001:59:3002:00:0002:00:3002:01:00No restart: every counter only climbs, so both orders draw the same example 2.4 requests/s.
What to watch for

How do you turn ever-growing counters into 'requests per second across the fleet' without the graph lying? Three checkout pods keep running totals of requests served (example values around 1,000, 1,200 and 800). A total that only climbs doesn't say how busy a pod is right now, so watch each pod's climb get turned into requests per second over a short window, then watch the three pods' numbers get added into one line for the whole checkout service.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Pod.restart
the counter climbs in memory and starts again from 0
http_requests_total = 0
def on_request():
http_requests_total += 1
def on_scrape():
return http_requests_total
def restart():
http_requests_total = 0
Query.rate
per-second increase over a window, repairing drops
def rate(series, window):
samples = series.in_window(window) # >= 4x scrape
if len(samples) < 2:
return None
increase = 0
for prev, cur in pairs(samples):
if cur < prev:
increase += cur
else:
increase += cur - prev
return increase / span_seconds(samples)
Query.rate_then_sum
one rate per pod series, then a single addition
def rate_then_sum(pods, window):
# sum(rate(http_requests_total[1m]))
per_pod = []
for p in pods:
r = rate(p.series, window)
if r is None:
return None
per_pod.append(r)
return sum(per_pod)
Query.sum_then_rate
pods added per scrape first, then one rate of that series
def sum_then_rate(pods, window):
# rate(sum(http_requests_total)[1m:15s])
combined = []
for t in scrape_times:
total = sum(p.series.at(t) for p in pods)
combined.append(total)
return rate(combined, window)

Where this sits in Metrics / Monitoring System

Scene 05 of 18, in the Ask act — rate() before sum(); add buckets, never average p99s.. A pod's counter drops to zero when it restarts, so compute each series' per-second rate first and sum after: summing the raw counters first hides the reset and draws a false spike.

Up next. Now that request counts combine honestly with rate-then-sum, the next question is whether latency percentiles combine the same way across pods.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes