Cardinality is the killer
Each unique label-set is a brand-new time series with its own head chunk, index entry, and postings inserts — so adding a label whose values are unbounded multiplies series count and pierces the head's RAM budget within minutes.
The index is fast — until you make it index a million tiny postings lists. Cardinality is where the architecture meets reality.
Scene 09
Cardinality is the killer
- Watch
- Try it
- Predict
- Capture
Start with two safe labels — method and status. Four series, head chunk barely visible on the RAM bar. Then add path with 7 bounded values: 28 series, still calm. Both labels are enumerable, and that is what keeps the curve flat.
Highlighted lines are the ones running in the diagram right now.
def series_id(metric_name, labels):# canonical: name + sorted (k, v) pairsparts = [metric_name]for k in sorted(labels.keys()):parts.append(k + '=' + labels[k])# one distinct value of ANY label# mints a brand-new series idreturn hash('|'.join(parts))
def on_sample(metric_name, labels, ts, value):sid = series_id(metric_name, labels)series = head.index.get(sid)if series is None:# NEW series: head chunk + index entry +# one postings insert per label.series = HeadSeries(labels)head.index[sid] = seriesfor k, v in labels.items():postings[(k, v)].append(sid)series.append(ts, value)
HEAD_CHUNK_BYTES = 12_288 # ~12 KB live chunkINDEX_ENTRY_BYTES = 4_096 # postings + label stringsBUDGET_MB = 8_000 # head RAM budgetdef estimate_ram_mb(dimensions, churn):series = 1for d in dimensions:series *= len(d.values) # MULTIPLY, not addbytes_per = HEAD_CHUNK_BYTES + INDEX_ENTRY_BYTESstale = 1.25 if churn else 1.0mb = series * bytes_per * stale / (1024 * 1024)assert mb < BUDGET_MB, 'OOM-killed'
Where this sits in Build a Prometheus-style time-series database
Scene 09 of 12. Each unique label-set is one series with its own head chunk in RAM. Add an unbounded label like user_id and you OOM in minutes.
Up next. Cardinality bounds the width of the database. Now look at the depth — every series accumulates points forever, and a 30-day dashboard can't decode billions of them on the fly.
All 12 scenes in Build a Prometheus-style time-series database · Every curriculum