500 alerts, one page — alert grouping, inhibition and the shared notification log
Rules only decide that something is wrong; a separate alert router bundles alerts sharing labels into one notification, drops identical copies from redundant monitors, and holds back symptoms of a known cause, so a zone outage becomes one page.
Rules decide something is wrong; a separate system must decide who hears about it and how many times.
Scene 08
500 alerts, one page
- Watch
- Try it
- Predict
- Capture
What decides who gets told, how often, and what gets held back? Zone-b has just failed. Two monitors watch the same pods, so each of them fires a PodDown alert for all 500 — plus one ZoneDown alert naming the cause — and each sends every alert to both alert routers. Nothing is configured on the routers yet. Watch the pages-sent counter and let it finish.
Highlighted lines are the ones running in the diagram right now.
def send_alerts(firing):for alert in firing:# alert_relabel_configs drops the tag# naming which monitor sent itif drop_monitor_tag:del alert.labels["monitor"]for router in all_routers:# never behind a load balancerrouter.post(alert)
def on_alert(alert):if group_by: # alertname, zonekey = tuple(alert.labels[l] for l in group_by)else:key = alert.fingerprint # ['...'] = no bundlegroup = groups.setdefault(key, [])group.append(alert)if group.first_flush:# window for a cause alert to arriveschedule(flush_group, key, after=group_wait)else:schedule(flush_group, key, after=group_interval)
def flush_group(key):alerts = groups.pop(key)for rule in inhibit_rules: # source -> targetif rule.source_is_firing and \rule.equal_labels_match(alerts):return # held back, no notifywait(self.position * 15) # cluster.peer-timeoutif notification_log.has(key):return # a peer already sent itnotify(receiver, alerts)notification_log.add(key)
Where this sits in Metrics / Monitoring System
Scene 08 of 18, in the Wake a human act — Patient rules, one page, and a heartbeat for the monitor.. Rules only decide something is wrong: a separate router bundles alerts that share labels, drops identical copies from redundant monitors and holds back symptoms of a known cause, so a zone outage is one page.
Up next. Now that alerts reliably become one page, the next question is what happens when the pipeline delivering them is itself broken.
All 18 scenes in Metrics / Monitoring System · Every curriculum