Silence looks exactly like health — the dead man's switch heartbeat alert
When the monitoring pipeline dies no alerts fire, which looks exactly like health, so you keep an always-firing heartbeat alert that an outside service expects to hear and pages when it stops.
Alerts reliably become one page, but only while every stage of the pipeline is alive.
Scene 08a
Silence looks exactly like health
- Watch
- Try it
- Predict
- Capture
How do you find out that the monitoring itself is broken? Everything you have learned so far assumes the pipeline is running. Watch one Friday night: the monitor's memory bar climbs, it dies at 02:00, and five minutes later checkout starts failing. Keep your eyes on the two panels at the bottom — the alerts dashboard and the on-call phone — and ask yourself how this night differs from a night where nothing happened at all.
Highlighted lines are the ones running in the diagram right now.
watchdog = Rule(expr = "vector(1)", # always trueroute_to = "outside-check",)def evaluate(rules): # every evaluation_intervalfor rule in rules:if rule.expr_is_true():a = rule.fire()a.valid_until = now() + 4 * max(evaluation_interval, resend_delay)router.post(a)# all of this runs inside the monitor process
def on_alert(alert): # posted by the monitorgroups[alert.key].append(alert)schedule(flush, alert.key, after=group_wait)def flush(key):for alert in groups.pop(key):if now() > alert.valid_until:resolve(alert) # no renewal arrivedcontinuereceiver_of(alert).send(alert)# the watchdog's receiver is the outside check# both halves run inside your cluster
last_seen = None # state at a different providerdef on_pulse(alert):if alert.is_firing:last_seen = now()def on_tick():if not configured:returnif now() - last_seen <= timeout: # 5 min (example)returnpage(on_call, path=outside_the_cluster)
Where this sits in Metrics / Monitoring System
Scene 08a of 18, in the Wake a human act — Patient rules, one page, and a heartbeat for the monitor.. When the monitoring pipeline dies no alerts fire, which looks exactly like health — so an always-firing heartbeat leaves the pipeline and an outside service pages when it stops arriving.
Up next. Now that we can tell when the monitor dies, the next question is the most common reason it dies at scale: one server can't hold every series.
All 18 scenes in Metrics / Monitoring System · Every curriculum