Metrics / Monitoring System
18 scenes · ~126 min · build the primitive

Build your own Metrics / Monitoring System

Time-series at scale. Cardinality is the enemy.

Scenes
18 interactive scenes
Time
about 126 minutes
Topic
Observability: Metrics, Logs & Traces

What you are building, and why

You are building the monitoring platform for Shopfront, an online store on Kubernetes whose checkout service must never fail silently. At 02:14 on a Tuesday checkout starts returning errors. Either a machine notices within a couple of minutes and pages the on-call engineer, or customers notice first and complain for forty minutes. Everything in this problem exists to make the first outcome true — for one team with two hundred pods, and then for thirty teams with fifty thousand.

build-tsdb built the storage engine that keeps one server's samples. This problem is the product and platform on top of it: how numbers are instrumented and collected, how they are queried without lying, how a query becomes exactly one page, how the whole thing scales past one server's memory, and how it survives the Friday deploy that adds a customer_id label.

The thesis runs through every stage: cardinality is the enemy. A metric's cost ignores traffic and multiplies with every label. Every design decision below is either a way to get more signal per series or a door that stops one team's label explosion from silencing everyone else's alerts. Make each decision yourself, defend it against the workload, and keep one number honest the whole way: time to detect.

What you will be able to explain afterwards

  • metrics vs logs: cost follows label combinations, not traffic
  • counters, gauges, histograms and cumulative reporting
  • pull scraping, the up metric, and when push fits
  • rate() per series before sum(); percentiles from summed buckets
  • alerting rules with for, routing with grouping and inhibition
  • meta-monitoring: the always-firing heartbeat and dead man's switch
  • federation vs remote-write out of one server's memory wall
  • distributor and ingesters: full-label-set hashing, 3 copies, 2 confirmations
  • HA pair deduplication: write-time election vs query-time merge
  • immutable 2 h blocks on object storage and the compactor
  • query frontend: split by day, shard, cache finished days
  • churn, relabeling, sample limits, tenant limits, pre-aggregation and exemplars
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 08a
  10. 09
  11. 10
  12. 10a
  13. 11
  14. 12
  15. 13
  16. 14
  17. 15
  18. 16

Why measure

Numbers beat complaints; labels, not traffic, set the bill.

  1. 01
    Find out before your customers do — the monitoring pipeline and time-to-detect
    Monitoring means numbers measured on a schedule and checked by rules on a schedule, so a machine — not a customer — is first to notice checkout is broken.
    ~7 min
  2. 02
    Cost grows with labels, not traffic — label cardinality and the series count
    The app reports one number per label combination per interval, so a metric's cost ignores traffic and multiplies with every label: one unbounded label like customer_id dwarfs a 100× traffic spike.
    ~7 min

Collect

Running totals, and who starts the conversation.

  1. 03
    Running totals survive missed reads — counters, gauges and reset-on-read
    A counter reports a running total that only goes up and a gauge a current level; because totals are reported, a missed or doubled read loses at most detail, never the count itself.
    ~7 min
  2. 04
    Pull or push: who starts the conversation — scrape targets, the up metric and service discovery
    When the monitor pulls each target on its own schedule, a failed fetch itself says the target is down; when apps push, silence becomes ambiguous — dead, or just quiet?
    ~7 min

Ask

rate() before sum(); add buckets, never average p99s.

  1. 05
    rate() first, then sum() — counter resets and aggregation across pods
    A pod's counter drops to zero when it restarts, so compute each series' per-second rate first and sum after: summing the raw counters first hides the reset and draws a false spike.
    ~7 min
  2. 06
    You can't average a p99
    Per-pod percentiles can't be combined, but histogram buckets are plain counters that add up across pods, so the fleet percentile is estimated after summing — and bucket boundaries set its accuracy.
    ~7 min

Wake a human

Patient rules, one page, and a heartbeat for the monitor.

  1. 07
    An alert is a query with patience — the for duration, pending and firing states
    An alerting rule is a query re-run every interval; its waiting period holds the alert pending until the condition has stayed true that long, trading a known delay for far fewer false pages.
    ~7 min
  2. 08
    500 alerts, one page — alert grouping, inhibition and the shared notification log
    Rules only decide something is wrong: a separate router bundles alerts that share labels, drops identical copies from redundant monitors and holds back symptoms of a known cause, so a zone outage is one page.
    ~7 min
  3. 08a
    Silence looks exactly like health — the dead man's switch heartbeat alert
    When the monitoring pipeline dies no alerts fire, which looks exactly like health — so an always-firing heartbeat leaves the pipeline and an outside service pages when it stops arriving.
    ~7 min

Scale out

Remote-write, ingesters, dedup, blocks, split queries.

  1. 09
    One monitor hits a wall — splitting targets, federation and remote-write
    One server's memory holds every active series, so a growing fleet has three exits: split targets and query them all, federate sums and lose per-pod detail, or remote-write to a cluster you must now run.
    ~7 min
  2. 10
    Hash every series to three ingesters
    The cluster's front door hashes each series' full label set onto a ring and writes to three ingesters, succeeding once two confirm — no popular metric lands on one machine, but every query asks them all.
    ~7 min
  3. 10a
    Two scrapers, one copy of the truth — replica labels and deduplication at write or query
    A redundant scraper pair sends every sample twice, so the cluster keeps one stream: elect one replica at write time and accept a short failover gap, or store both and merge at query time for twice the storage.
    ~7 min
  4. 11
    Keep a year cheaply: blocks on object storage
    Ingesters keep only recent hours in memory and cut immutable two-hour blocks to object storage, where a compactor merges the copies replication created — retention becomes a storage and query-speed problem.
    ~7 min
  5. 12
    Split, shard, and cache the past
    A query frontend splits a long range into day-sized pieces, spreads them across workers and caches finished days, which almost never change, so a refresh recomputes only the newest slice.
    ~7 min

Cardinality

Churn, defenses, aggregation, then design it.

  1. 13
    Churn: yesterday's pods still cost you — active series versus series held in memory
    Every rolling deploy gives pods new names and brand-new series, while the old ones go quiet but linger in memory and in every block index — cost follows series seen recently, not the active count.
    ~7 min
  2. 14
    Stop the bomb at every door — relabeling, per-scrape sample and per-tenant series limits
    Each door stops a cardinality explosion by failing something different: relabeling drops the metric, a per-scrape limit fails the whole scrape and silences that target, a tenant limit refuses only new series.
    ~7 min
  3. 15
    Keep the answer, drop the series — pre-aggregation on arrival and bucket exemplars
    Summing away labels nobody queries as samples arrive cuts series by orders of magnitude, while deleting every distinguishing label makes counters collide — and an exemplar keeps one example customer.
    ~7 min
  4. 16
    Design your monitoring platform
    A monitoring platform is a chain of decisions — collection, alerting, scaling out, retention and label budget — each right only for a workload whose series count, teams, retention and detection target justify it.
    ~7 min

Prefer to design it yourself?

The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.

Open the Metrics / Monitoring System workspace