An alert is a query with patience — the for duration, pending and firing states

An alerting rule is a query re-run at a fixed interval whose for duration holds the alert pending until the condition has stayed true continuously that long, trading a known extra delay for far fewer false pages.

Previously

We can compute honest numbers; now a machine must decide when one of them is bad enough, and long enough, to wake someone.

Scene 07

An alert is a query with patience

  1. Watch
  2. Try it
  3. Predict
  4. Capture
ALERTING RULEalerting rule — re-run "5xx ratio > 5%" every 15 s, page only if it stays true 2 minTIME TO DETECT≈ 2.5 minexample≤ 2 minpages5xx ratiocheckout really is failing — 6 minutes5xx ratio > 5%eval ticksone evaluation every 15 s (example) · 48 ticks = 12 minutesfor: 2minactivependingfiringpending = true, not long enough yet · firing = pages1 page/week (example)for: 2m — spikes under 2 minutes are dropped; detect ≈ 2.5 min (example): 30 s plus the wait.
What to watch for

How does a number become a page without paging on every blip? A machine has to watch checkout's 5xx ratio for us. Watch the ticks march along the bottom: at every one of them the same query runs again and the answer is compared to the threshold. The strip underneath records what the alert decided after each evaluation. Two brief spikes and one real outage go past — see which of them actually rings a bell.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Rule.evaluate
one query, one comparison, re-run on a timer
def evaluate(rule): # every evaluation_interval
value = query(rule.expr)
if value > rule.threshold:
update_state(rule, now())
else:
clear(rule)
def run_rules(rules):
while True:
for rule in rules:
evaluate(rule)
sleep(evaluation_interval) # 1m default
Rule.update_state
pending while the run is short, firing once it is long enough
def update_state(rule, now):
if rule.active is None:
rule.active = Alert(started_at=now)
held = now - rule.active.started_at
if held >= rule.for_duration: # `for`
rule.active.state = FIRING
page(rule.active)
else:
rule.active.state = PENDING
def clear(rule):
del rule.active
Pager.page
only a firing alert sends, and only once per run
def page(alert):
if alert.state != FIRING:
return
if alert.already_paged:
return
alert.already_paged = True
send_page(alert)
def worst_case_ttd(rule):
return (scrape_interval
+ evaluation_interval
+ rule.for_duration)

Where this sits in Metrics / Monitoring System

Scene 07 of 18, in the Wake a human act — Patient rules, one page, and a heartbeat for the monitor.. An alerting rule is a query re-run every interval; its waiting period holds the alert pending until the condition has stayed true that long, trading a known delay for far fewer false pages.

Up next. Now that rules decide something is wrong, the next question is who gets told, and how 500 firing alerts avoid becoming 500 pages.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes