Running totals survive missed reads — counters, gauges and reset-on-read
A counter reports a running total that only goes up and a gauge reports a current level, and because totals are reported instead of counts since the last read, a missed or doubled read loses at most detail.
Each series is one number reported every interval; now we decide what that number means so lost or repeated reads can't corrupt it.
Scene 03
Running totals survive missed reads
- Watch
- Try it
- Predict
- Capture
What shapes can a metric take, and why report 'total so far' instead of 'since last time'? Checkout holds a number of requests served, and a collector reads it every 15 seconds (example). Watch the number climb from read to read, then look at the panel below: in-flight requests, the requests being handled right now, rise and fall with load.
Highlighted lines are the ones running in the diagram right now.
def handle_request(req):in_flight += 1requests_served += 1since_last_read += 1resp = checkout(req)in_flight -= 1return resp
def on_read():if report_running_total:return requests_servedcount = since_last_readsince_last_read = 0return count
def read(collector): # every 15 s (example)reply = pod.on_read()if reply is None:returncollector.recorded.append(Sample(now(), reply),)
def counted(collector):reads = collector.recordedif report_running_total:return reads[-1].value - reads[0].valuereturn sum(r.value for r in reads)
Where this sits in Metrics / Monitoring System
Scene 03 of 18, in the Collect act — Running totals, and who starts the conversation.. A counter reports a running total that only goes up and a gauge a current level; because totals are reported, a missed or doubled read loses at most detail, never the count itself.
Up next. Now that running totals survive missed and repeated reads, the next question is who starts each read: does the monitor fetch the numbers, or does the app send them?
All 18 scenes in Metrics / Monitoring System · Every curriculum