Build your own Metrics / Monitoring System
Time-series at scale. Cardinality is the enemy.
- Scenes
- 18 interactive scenes
- Time
- about 126 minutes
- Topic
- Observability: Metrics, Logs & Traces
What you are building, and why
You are building the monitoring platform for Shopfront, an online store on Kubernetes whose checkout service must never fail silently. At 02:14 on a Tuesday checkout starts returning errors. Either a machine notices within a couple of minutes and pages the on-call engineer, or customers notice first and complain for forty minutes. Everything in this problem exists to make the first outcome true — for one team with two hundred pods, and then for thirty teams with fifty thousand.
build-tsdb built the storage engine that keeps one server's samples. This problem is the product and platform on top of it: how numbers are instrumented and collected, how they are queried without lying, how a query becomes exactly one page, how the whole thing scales past one server's memory, and how it survives the Friday deploy that adds a customer_id label.
The thesis runs through every stage: cardinality is the enemy. A metric's cost ignores traffic and multiplies with every label. Every design decision below is either a way to get more signal per series or a door that stops one team's label explosion from silencing everyone else's alerts. Make each decision yourself, defend it against the workload, and keep one number honest the whole way: time to detect.
What you will be able to explain afterwards
- metrics vs logs: cost follows label combinations, not traffic
- counters, gauges, histograms and cumulative reporting
- pull scraping, the up metric, and when push fits
- rate() per series before sum(); percentiles from summed buckets
- alerting rules with for, routing with grouping and inhibition
- meta-monitoring: the always-firing heartbeat and dead man's switch
- federation vs remote-write out of one server's memory wall
- distributor and ingesters: full-label-set hashing, 3 copies, 2 confirmations
- HA pair deduplication: write-time election vs query-time merge
- immutable 2 h blocks on object storage and the compactor
- query frontend: split by day, shard, cache finished days
- churn, relabeling, sample limits, tenant limits, pre-aggregation and exemplars
Why measure
Numbers beat complaints; labels, not traffic, set the bill.
- 01Find out before your customers do — the monitoring pipeline and time-to-detectMonitoring means numbers measured on a schedule and checked by rules on a schedule, so a machine — not a customer — is first to notice checkout is broken.~7 min
- 02Cost grows with labels, not traffic — label cardinality and the series countThe app reports one number per label combination per interval, so a metric's cost ignores traffic and multiplies with every label: one unbounded label like customer_id dwarfs a 100× traffic spike.~7 min
Collect
Running totals, and who starts the conversation.
- 03Running totals survive missed reads — counters, gauges and reset-on-readA counter reports a running total that only goes up and a gauge a current level; because totals are reported, a missed or doubled read loses at most detail, never the count itself.~7 min
- 04Pull or push: who starts the conversation — scrape targets, the up metric and service discoveryWhen the monitor pulls each target on its own schedule, a failed fetch itself says the target is down; when apps push, silence becomes ambiguous — dead, or just quiet?~7 min
Ask
rate() before sum(); add buckets, never average p99s.
- 05rate() first, then sum() — counter resets and aggregation across podsA pod's counter drops to zero when it restarts, so compute each series' per-second rate first and sum after: summing the raw counters first hides the reset and draws a false spike.~7 min
- 06You can't average a p99Per-pod percentiles can't be combined, but histogram buckets are plain counters that add up across pods, so the fleet percentile is estimated after summing — and bucket boundaries set its accuracy.~7 min
Wake a human
Patient rules, one page, and a heartbeat for the monitor.
- 07An alert is a query with patience — the for duration, pending and firing statesAn alerting rule is a query re-run every interval; its waiting period holds the alert pending until the condition has stayed true that long, trading a known delay for far fewer false pages.~7 min
- 08500 alerts, one page — alert grouping, inhibition and the shared notification logRules only decide something is wrong: a separate router bundles alerts that share labels, drops identical copies from redundant monitors and holds back symptoms of a known cause, so a zone outage is one page.~7 min
- 08aSilence looks exactly like health — the dead man's switch heartbeat alertWhen the monitoring pipeline dies no alerts fire, which looks exactly like health — so an always-firing heartbeat leaves the pipeline and an outside service pages when it stops arriving.~7 min
Scale out
Remote-write, ingesters, dedup, blocks, split queries.
- 09One monitor hits a wall — splitting targets, federation and remote-writeOne server's memory holds every active series, so a growing fleet has three exits: split targets and query them all, federate sums and lose per-pod detail, or remote-write to a cluster you must now run.~7 min
- 10Hash every series to three ingestersThe cluster's front door hashes each series' full label set onto a ring and writes to three ingesters, succeeding once two confirm — no popular metric lands on one machine, but every query asks them all.~7 min
- 10aTwo scrapers, one copy of the truth — replica labels and deduplication at write or queryA redundant scraper pair sends every sample twice, so the cluster keeps one stream: elect one replica at write time and accept a short failover gap, or store both and merge at query time for twice the storage.~7 min
- 11Keep a year cheaply: blocks on object storageIngesters keep only recent hours in memory and cut immutable two-hour blocks to object storage, where a compactor merges the copies replication created — retention becomes a storage and query-speed problem.~7 min
- 12Split, shard, and cache the pastA query frontend splits a long range into day-sized pieces, spreads them across workers and caches finished days, which almost never change, so a refresh recomputes only the newest slice.~7 min
Cardinality
Churn, defenses, aggregation, then design it.
- 13Churn: yesterday's pods still cost you — active series versus series held in memoryEvery rolling deploy gives pods new names and brand-new series, while the old ones go quiet but linger in memory and in every block index — cost follows series seen recently, not the active count.~7 min
- 14Stop the bomb at every door — relabeling, per-scrape sample and per-tenant series limitsEach door stops a cardinality explosion by failing something different: relabeling drops the metric, a per-scrape limit fails the whole scrape and silences that target, a tenant limit refuses only new series.~7 min
- 15Keep the answer, drop the series — pre-aggregation on arrival and bucket exemplarsSumming away labels nobody queries as samples arrive cuts series by orders of magnitude, while deleting every distinguishing label makes counters collide — and an exemplar keeps one example customer.~7 min
- 16Design your monitoring platformA monitoring platform is a chain of decisions — collection, alerting, scaling out, retention and label budget — each right only for a workload whose series count, teams, retention and detection target justify it.~7 min
More in Observability: Metrics, Logs & Traces
The three pillars, built rather than bought: a time-series database, a log pipeline, a tracing system, and the metrics platform on top.
- Build Build a distributed tracing system (Jaeger / Zipkin style)Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'
- Build Build a distributed logging stack (ELK / Loki)Ship lines off N hosts, choose what to index, age data through tiers, retain or delete on schedule, and survive a chatty service — built one decision at a time.
- Build Build a Prometheus-style time-series databaseThe simplest database that can absorb a 1M-points-per-second firehose and still answer `sum(rate(http_requests_total{status="500"}[5m]))` in milliseconds — built bit by bit, literally.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Metrics / Monitoring System workspace