Running totals survive missed reads — counters, gauges and reset-on-read

A counter reports a running total that only goes up and a gauge reports a current level, and because totals are reported instead of counts since the last read, a missed or doubled read loses at most detail.

Previously

Each series is one number reported every interval; now we decide what that number means so lost or repeated reads can't corrupt it.

Scene 03

Running totals survive missed reads

  1. Watch
  2. Try it
  3. Predict
  4. Capture
TIME TO DETECT≈ 30 sexampleREPORT STYLErunning total: pod's number only goes upcheckout pod: requestscheckout pod: requests numbernumbercheckout pod: requests numberclimbs, never resetscollector A readscollector A readsone read per stepCOLLECTOR A RECORDED153 of 153 · all counted153 of 153 · all countedcollector B readscollector B readsoffcollector B is not readingCOLLECTOR B RECORDEDnot readingnot readingTRUE REQUESTS AFTER FIRST READ153 (example)153 (example)compare each collector below0s15s30s45s60s75s90s105s120s135s150s165sin-flight requestsin-flight requestsside panel: a levela level: rises and fallsa level: rises and falls↕ goes down too; summing it means nothingEvery read returns checkout's number so far: 1,193, 1,207, 1,224… (example). It only goes up, and returns to 0 only on a restart.
What to watch for

What shapes can a metric take, and why report 'total so far' instead of 'since last time'? Checkout holds a number of requests served, and a collector reads it every 15 seconds (example). Watch the number climb from read to read, then look at the panel below: in-flight requests, the requests being handled right now, rise and fall with load.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

App.handle_request
one number that only climbs, one that rises and falls
def handle_request(req):
in_flight += 1
requests_served += 1
since_last_read += 1
resp = checkout(req)
in_flight -= 1
return resp
App.on_read
what the pod hands back, in each report style
def on_read():
if report_running_total:
return requests_served
count = since_last_read
since_last_read = 0
return count
Collector.read
each collector asks the pod once per interval
def read(collector): # every 15 s (example)
reply = pod.on_read()
if reply is None:
return
collector.recorded.append(
Sample(now(), reply),
)
Collector.counted
turning stored reads into a request count
def counted(collector):
reads = collector.recorded
if report_running_total:
return reads[-1].value - reads[0].value
return sum(r.value for r in reads)

Where this sits in Metrics / Monitoring System

Scene 03 of 18, in the Collect act — Running totals, and who starts the conversation.. A counter reports a running total that only goes up and a gauge a current level; because totals are reported, a missed or doubled read loses at most detail, never the count itself.

Up next. Now that running totals survive missed and repeated reads, the next question is who starts each read: does the monitor fetch the numbers, or does the app send them?

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes