Silence looks exactly like health — the dead man's switch heartbeat alert

When the monitoring pipeline dies no alerts fire, which looks exactly like health, so you keep an always-firing heartbeat alert that an outside service expects to hear and pages when it stops.

Previously

Alerts reliably become one page, but only while every stage of the pipeline is alive.

Scene 08a

Silence looks exactly like health

  1. Watch
  2. Try it
  3. Predict
  4. Capture
instrumentcollectstorequeryalertnotifyTIME TO DETECT—nothing t…YOUR CLUSTERmonitorrunningalert routerrunningpager integrationrunningmonitor memorycustomer_idFriday deployTIMELINE01:58all six stages runningoutside check (not set up)offexpects a pulse every 5 min (example)WHAT THE ON-CALL ENGINEER SEESalerts dashboardalerts firing: 0on-call phonequiet — no pages01:58 — every stage is running. Checkout is healthy, the alerts dashboard is empty, and an empty dashboard is exactly what a good night looks like.
What to watch for

How do you find out that the monitoring itself is broken? Everything you have learned so far assumes the pipeline is running. Watch one Friday night: the monitor's memory bar climbs, it dies at 02:00, and five minutes later checkout starts failing. Keep your eyes on the two panels at the bottom — the alerts dashboard and the on-call phone — and ask yourself how this night differs from a night where nothing happened at all.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Rule.watchdog
an alerting rule whose condition is always true
watchdog = Rule(
expr = "vector(1)", # always true
route_to = "outside-check",
)
def evaluate(rules): # every evaluation_interval
for rule in rules:
if rule.expr_is_true():
a = rule.fire()
a.valid_until = now() + 4 * max(
evaluation_interval, resend_delay)
router.post(a)
# all of this runs inside the monitor process
Router.notify
delivers firing alerts, resolves the ones that stop renewing
def on_alert(alert): # posted by the monitor
groups[alert.key].append(alert)
schedule(flush, alert.key, after=group_wait)
def flush(key):
for alert in groups.pop(key):
if now() > alert.valid_until:
resolve(alert) # no renewal arrived
continue
receiver_of(alert).send(alert)
# the watchdog's receiver is the outside check
# both halves run inside your cluster
OutsideCheck.on_tick
a timer at another provider, watching for the pulse
last_seen = None # state at a different provider
def on_pulse(alert):
if alert.is_firing:
last_seen = now()
def on_tick():
if not configured:
return
if now() - last_seen <= timeout: # 5 min (example)
return
page(on_call, path=outside_the_cluster)

Where this sits in Metrics / Monitoring System

Scene 08a of 18, in the Wake a human act — Patient rules, one page, and a heartbeat for the monitor.. When the monitoring pipeline dies no alerts fire, which looks exactly like health — so an always-firing heartbeat leaves the pipeline and an outside service pages when it stops arriving.

Up next. Now that we can tell when the monitor dies, the next question is the most common reason it dies at scale: one server can't hold every series.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes