Cardinality is the killer

Each unique label-set is a brand-new time series with its own head chunk, index entry, and postings inserts — so adding a label whose values are unbounded multiplies series count and pierces the head's RAM budget within minutes.

Previously

The index is fast — until you make it index a million tiny postings lists. Cardinality is where the architecture meets reality.

Scene 09

Cardinality is the killer

  1. Watch
  2. Try it
  3. Predict
  4. Capture
LABEL-SET EXPLOSION — every unique combo is a seriesmethod:2status:2methodstatus200500GETPOST#1#2#3#4series:4HEAD RAMBUDGET (7.81 GB)Series4RAM0 MB / 7.81 GBBaseline: 2 methods × 2 statuses = 4 series. The head chunk is tiny. RAM bar is calm.
What to watch for

Start with two safe labels — method and status. Four series, head chunk barely visible on the RAM bar. Then add path with 7 bounded values: 28 series, still calm. Both labels are enumerable, and that is what keeps the curve flat.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Series.id
the full label-set IS the series identity
def series_id(metric_name, labels):
# canonical: name + sorted (k, v) pairs
parts = [metric_name]
for k in sorted(labels.keys()):
parts.append(k + '=' + labels[k])
# one distinct value of ANY label
# mints a brand-new series id
return hash('|'.join(parts))
Head.onSample
lookup-or-create per scraped sample — create is the OOM path
def on_sample(metric_name, labels, ts, value):
sid = series_id(metric_name, labels)
series = head.index.get(sid)
if series is None:
# NEW series: head chunk + index entry +
# one postings insert per label.
series = HeadSeries(labels)
head.index[sid] = series
for k, v in labels.items():
postings[(k, v)].append(sid)
series.append(ts, value)
Head.estimateRamMb
cardinality is the PRODUCT of label cardinalities
HEAD_CHUNK_BYTES = 12_288 # ~12 KB live chunk
INDEX_ENTRY_BYTES = 4_096 # postings + label strings
BUDGET_MB = 8_000 # head RAM budget
def estimate_ram_mb(dimensions, churn):
series = 1
for d in dimensions:
series *= len(d.values) # MULTIPLY, not add
bytes_per = HEAD_CHUNK_BYTES + INDEX_ENTRY_BYTES
stale = 1.25 if churn else 1.0
mb = series * bytes_per * stale / (1024 * 1024)
assert mb < BUDGET_MB, 'OOM-killed'

Where this sits in Build a Prometheus-style time-series database

Scene 09 of 12. Each unique label-set is one series with its own head chunk in RAM. Add an unbounded label like user_id and you OOM in minutes.

Up next. Cardinality bounds the width of the database. Now look at the depth — every series accumulates points forever, and a 30-day dashboard can't decode billions of them on the fly.

All 12 scenes in Build a Prometheus-style time-series database · Every curriculum

Built with Arqly
Every scene in Build a Prometheus-style time-series database builds on the one before it.All 12 Build a Prometheus-style time-series database scenes