Build a Prometheus-style time-series database
12 scenes · ~84 min · build the primitive

Build your own Prometheus-style time-series database

The simplest database that can absorb a 1M-points-per-second firehose and still answer `sum(rate(http_requests_total{status="500"}[5m]))` in milliseconds — built bit by bit, literally.

Scenes
12 interactive scenes
Time
about 84 minutes
Topic
Observability: Metrics, Logs & Traces

What you are building, and why

You are designing the simplest database that can stomach a metrics firehose: 10,000 hosts each emitting 50 numbers every 15 seconds, forever, without losing precision and without needing a rack of disks. Clients push points in (scrape), or run PromQL-shaped queries pulling them out (sum(rate(metric{labels}[range]))); the engine packs every sample into a few bits on disk and answers a 5-minute alert in single-digit milliseconds.

Prometheus and Gorilla are the canonical worked examples. Once you build the core, the design space of every metrics store collapses to a handful of named trades — the bits-per-point you accept, the cardinality budget you defend, the depth-of-retention you afford, and where you put the line between "single-node" and "somebody else's problem."

Resist the urge to "describe Prometheus." Make decisions yourself, defend them, and let the workload push back. The point is to feel why each price is paid and which workload pays it.

What you will be able to explain afterwards

  • the time-series point shape: (metric, label_set, timestamp, value)
  • delta-of-delta encoding for fixed-cadence timestamps
  • XOR encoding for adjacent IEEE-754 floats
  • chunks: 120 samples or 2 hours of compressed points
  • head chunk in RAM + write-ahead log + sealed mmapped chunks
  • inverted index: postings lists per (label, value), intersection at query time
  • cardinality as the master operational variable
  • downsampling and the retention pyramid
  • single-node TSDB + HA via parallel scrapers + remote-write
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 10a
  12. 11
  1. 01
    What a metric point actually is
    A point is (metric, label_set, ts, value). The first two parts are the series identity — change one label and you've named a different stream.
    ~7 min
  2. 02
    Why one row per point is wrong — per-row label repetition and 60-byte overhead
    Storing each point as a SQL row spends 60+ bytes of metadata to carry an 8-byte float — the labels JSON repeats on every row of the firehose.
    ~7 min
  3. 03
    Delta-of-delta crushes timestamps
    Scrapers run on a fixed cadence, so the second derivative of timestamps is almost always zero — encoding ~96% of timestamps in a single bit.
    ~7 min
  4. 04
    XOR crushes adjacent floats
    Adjacent float64s share most of their IEEE-754 bits. XOR them and the leftover bits are tiny — an unchanged value costs 1 bit.
    ~7 min
  5. 05
    A chunk: 120 points, packed
    Bundle ~120 consecutive points into a single bit-packed blob. Gorilla: 16 B/point → 1.37 B/point — about 12× compression.
    ~7 min
  6. 06
    Head chunks, WAL, and flushing
    Active chunk lives in RAM (the head); a write-ahead log on disk catches every sample so a crash mid-chunk loses nothing.
    ~7 min
  7. 07
    How a query becomes points — the four-stage read path, parse to aggregate
    A read is four stages — parse, resolve label-selectors to series IDs, decompress the matching chunks, then aggregate. Stage 3 dominates.
    ~7 min
  8. 08
    Inverted index — labels to series
    For each (label, value) pair the database stores a sorted list of series IDs (a postings list). A multi-label query is the intersection.
    ~7 min
  9. 09
    Cardinality is the killer
    Each unique label-set is one series with its own head chunk in RAM. Add an unbounded label like user_id and you OOM in minutes.
    ~7 min
  10. 10
    Downsampling — a retention pyramid
    Aggregate old chunks into 5-minute, then 1-hour buckets, dropping originals as you go. Recording rules materialize — they cost storage.
    ~7 min
  11. 10a
    Single-node by design; HA is somebody else's problem
    The TSDB itself isn't replicated. HA = two parallel scrapers; durability = ship every sample to a remote-write receiver that dedups.
    ~7 min
  12. 11
    Design canvas: pick a workload, ship a config
    Capstone: alerting, tracing, or business KPIs — the verifier turns scrape interval, label set, retention, and rules into projected RAM, disk, and a fits/refuses verdict.
    ~7 min

Prefer to design it yourself?

The same subject as a staged workspace: draw the architecture, and a simulator traces requests through the boxes you drew.

Open the Build a Prometheus-style time-series database workspace