Find out before your customers do — the monitoring pipeline and time-to-detect

A monitoring system turns 'what is our software doing right now?' into numbers measured on a fixed schedule, kept in one place and checked by rules on a fixed schedule, so a machine rather than a customer is first to notice checkout is broken.

Scene 01

Find out before your customers do

  1. Watch
  2. Try it
  3. Predict
  4. Capture
measurenothing countedcollectcoming laterkeepquerycoming laterchecknobody checksnotifycoming later02:14 — Shopfront checkout starts failingCustomers complainselectedSomeone glances at a dashboardA rule checks every 15 s051015202530354045505560minutes after 02:14 (example)First run: every failed request is written to a log, but nothing is counting or checking anything.
What to watch for

What does it take to find out your site is broken before your customers tell you? At 02:14 Shopfront's checkout servers start answering requests with errors. Every failed request is written to a log, and nobody is counting anything. Watch how long it takes before anyone notices. Then the same night replays, this time with the servers counting their errors every 15 s and a rule checking that count every 15 s (an example setup). Watch which lane lights first, and the time-to-detect readout that appears top-right.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

App.handle_checkout
each checkout request; failures logged and counted
def handle_checkout(request):
result = checkout(request)
if result.failed:
log.write(request, result.error)
checkout_errors += 1
return result
App.measure
turns the error count into one number per interval
every 15 s: # measure interval (example)
value = checkout_errors
checkout_errors = 0
store.append("checkout_errors", now(), value)
Dashboard.watch
a person looking at the numbers on a screen
def watch_dashboard():
while person.at_screen():
screen.draw(store.series("checkout_errors"))
if person.glances() and screen.shows_spike():
person.page(on_call)
Rule.check
a machine compares the latest number to a limit
every 15 s: # check interval (example)
latest = store.last("checkout_errors")
if latest > limit:
pager.page(on_call, "checkout is failing")

Where this sits in Metrics / Monitoring System

Scene 01 of 18, in the Why measure act — Numbers beat complaints; labels, not traffic, set the bill.. Monitoring means numbers measured on a schedule and checked by rules on a schedule, so a machine — not a customer — is first to notice checkout is broken.

Up next. Now that a machine checks numbers on a schedule, the next question is what keeping those numbers costs, and what makes them expensive.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes