500 alerts, one page — alert grouping, inhibition and the shared notification log

Rules only decide that something is wrong; a separate alert router bundles alerts sharing labels into one notification, drops identical copies from redundant monitors, and holds back symptoms of a known cause, so a zone outage becomes one page.

Previously

Rules decide something is wrong; a separate system must decide who hears about it and how many times.

Scene 08

500 alerts, one page

  1. Watch
  2. Try it
  3. Predict
  4. Capture
instrumentcollectstorequeryalertnotifywho gets toldTIME TO DETECT≈ 2.5 minexample500 pods down in zone-b · ZoneDown firingmonitor 1sends 0 alertsmonitor 2sends 0 alertsshared log of what was sentalert router 1received alertsalert router 2received alertsLBno load balancer: every monitor sends to every routerno grouping: one notification per alertPodDown · monitor 1×500name tag attached→ notifyPodDown · monitor 2×500name tag attached→ notifyZoneDown · both copies×2name tag attached→ notifyinhibition (mute rule): offZoneDown mutes PodDown, same zonePAGES SENT0nothing sent yet11,000to the on-call phoneNothing is configured yet: every alert becomes its own notification, and each monitor's copy carries its own name tag, so copies never match.
What to watch for

What decides who gets told, how often, and what gets held back? Zone-b has just failed. Two monitors watch the same pods, so each of them fires a PodDown alert for all 500 — plus one ZoneDown alert naming the cause — and each sends every alert to both alert routers. Nothing is configured on the routers yet. Watch the pages-sent counter and let it finish.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Monitor.send_alerts
every firing alert goes to every router, name tag stripped
def send_alerts(firing):
for alert in firing:
# alert_relabel_configs drops the tag
# naming which monitor sent it
if drop_monitor_tag:
del alert.labels["monitor"]
for router in all_routers:
# never behind a load balancer
router.post(alert)
Router.on_alert
each arriving alert lands in a bucket keyed by its labels
def on_alert(alert):
if group_by: # alertname, zone
key = tuple(alert.labels[l] for l in group_by)
else:
key = alert.fingerprint # ['...'] = no bundle
group = groups.setdefault(key, [])
group.append(alert)
if group.first_flush:
# window for a cause alert to arrive
schedule(flush_group, key, after=group_wait)
else:
schedule(flush_group, key, after=group_interval)
Router.flush_group
mute check first, then the shared log, then notify
def flush_group(key):
alerts = groups.pop(key)
for rule in inhibit_rules: # source -> target
if rule.source_is_firing and \
rule.equal_labels_match(alerts):
return # held back, no notify
wait(self.position * 15) # cluster.peer-timeout
if notification_log.has(key):
return # a peer already sent it
notify(receiver, alerts)
notification_log.add(key)

Where this sits in Metrics / Monitoring System

Scene 08 of 18, in the Wake a human act — Patient rules, one page, and a heartbeat for the monitor.. Rules only decide something is wrong: a separate router bundles alerts that share labels, drops identical copies from redundant monitors and holds back symptoms of a known cause, so a zone outage is one page.

Up next. Now that alerts reliably become one page, the next question is what happens when the pipeline delivering them is itself broken.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes