Build your own Prometheus-style time-series database
The simplest database that can absorb a 1M-points-per-second firehose and still answer `sum(rate(http_requests_total{status="500"}[5m]))` in milliseconds — built bit by bit, literally.
- Scenes
- 12 interactive scenes
- Time
- about 84 minutes
- Topic
- Observability: Metrics, Logs & Traces
What you are building, and why
You are designing the simplest database that can stomach a metrics firehose: 10,000 hosts each emitting 50 numbers every 15 seconds, forever, without losing precision and without needing a rack of disks. Clients push points in (scrape), or run PromQL-shaped queries pulling them out (sum(rate(metric{labels}[range]))); the engine packs every sample into a few bits on disk and answers a 5-minute alert in single-digit milliseconds.
Prometheus and Gorilla are the canonical worked examples. Once you build the core, the design space of every metrics store collapses to a handful of named trades — the bits-per-point you accept, the cardinality budget you defend, the depth-of-retention you afford, and where you put the line between "single-node" and "somebody else's problem."
Resist the urge to "describe Prometheus." Make decisions yourself, defend them, and let the workload push back. The point is to feel why each price is paid and which workload pays it.
What you will be able to explain afterwards
- the time-series point shape: (metric, label_set, timestamp, value)
- delta-of-delta encoding for fixed-cadence timestamps
- XOR encoding for adjacent IEEE-754 floats
- chunks: 120 samples or 2 hours of compressed points
- head chunk in RAM + write-ahead log + sealed mmapped chunks
- inverted index: postings lists per (label, value), intersection at query time
- cardinality as the master operational variable
- downsampling and the retention pyramid
- single-node TSDB + HA via parallel scrapers + remote-write
- 01What a metric point actually isA point is (metric, label_set, ts, value). The first two parts are the series identity — change one label and you've named a different stream.~7 min
- 02Why one row per point is wrong — per-row label repetition and 60-byte overheadStoring each point as a SQL row spends 60+ bytes of metadata to carry an 8-byte float — the labels JSON repeats on every row of the firehose.~7 min
- 03Delta-of-delta crushes timestampsScrapers run on a fixed cadence, so the second derivative of timestamps is almost always zero — encoding ~96% of timestamps in a single bit.~7 min
- 04XOR crushes adjacent floatsAdjacent float64s share most of their IEEE-754 bits. XOR them and the leftover bits are tiny — an unchanged value costs 1 bit.~7 min
- 05A chunk: 120 points, packedBundle ~120 consecutive points into a single bit-packed blob. Gorilla: 16 B/point → 1.37 B/point — about 12× compression.~7 min
- 06Head chunks, WAL, and flushingActive chunk lives in RAM (the head); a write-ahead log on disk catches every sample so a crash mid-chunk loses nothing.~7 min
- 07How a query becomes points — the four-stage read path, parse to aggregateA read is four stages — parse, resolve label-selectors to series IDs, decompress the matching chunks, then aggregate. Stage 3 dominates.~7 min
- 08Inverted index — labels to seriesFor each (label, value) pair the database stores a sorted list of series IDs (a postings list). A multi-label query is the intersection.~7 min
- 09Cardinality is the killerEach unique label-set is one series with its own head chunk in RAM. Add an unbounded label like user_id and you OOM in minutes.~7 min
- 10Downsampling — a retention pyramidAggregate old chunks into 5-minute, then 1-hour buckets, dropping originals as you go. Recording rules materialize — they cost storage.~7 min
- 10aSingle-node by design; HA is somebody else's problemThe TSDB itself isn't replicated. HA = two parallel scrapers; durability = ship every sample to a remote-write receiver that dedups.~7 min
- 11Design canvas: pick a workload, ship a configCapstone: alerting, tracing, or business KPIs — the verifier turns scrape interval, label set, retention, and rules into projected RAM, disk, and a fits/refuses verdict.~7 min
More in Observability: Metrics, Logs & Traces
The three pillars, built rather than bought: a time-series database, a log pipeline, a tracing system, and the metrics platform on top.
- Build Metrics / Monitoring SystemTime-series at scale. Cardinality is the enemy.
- Build Build a distributed tracing system (Jaeger / Zipkin style)Metrics tell you that checkout is slow and logs tell you what one service printed; neither can say which hop of one request consumed the time. Build the third pillar: one record per hop tied together by one id — span propagation, sampling that doesn't lie, an export path that never blocks a customer, storage keyed by trace id, and the waterfall that finally answers 'where did the latency go?'
- Build Build a distributed logging stack (ELK / Loki)Ship lines off N hosts, choose what to index, age data through tiers, retain or delete on schedule, and survive a chatty service — built one decision at a time.
Prefer to design it yourself?
The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.
Open the Build a Prometheus-style time-series database workspace