Downsampling — a retention pyramid
Long-range queries stay tractable by pre-aggregating old chunks into coarser tiers (5 m, 1 h) and dropping the originals — but each tier is materialized as a real series with its own head chunk, postings, and disk chunks, not a query-time hint.
Cardinality bounds width. Retention bounds depth — and downsampling is how depth stays affordable.
Scene 10
Downsampling: a retention pyramid
- Watch
- Try it
- Predict
- Capture
Watch a query range slide across three tiers. A 1-day range is served by the raw tier; a 60-day range falls through to the 5 m rollup; a 1-year range drops to the 1 h rollup. The engine always picks the coarsest tier that fully covers the range.
Highlighted lines are the ones running in the diagram right now.
def evaluateRecordingRule(rule, window):# rule.expr e.g. rate(http_requests_total[5m])samples = promql.eval(rule.expr, window)for labels, value in samples:series = tsdb.getOrCreateSeries(name = rule.name, # job:http_requests:rate5mlabels = labels,)series.headChunk.append(now(), value)postings.index(series.id, labels)# real bytes, real cardinality, real cost
def pickTier(query):# tiers ordered finest -> coarsesttiers = [raw, rate5m, rate1h]for tier in reversed(tiers): # try coarsest firstif tier.retentionDays >= query.range.days:covers = tier# coarsest tier that still reaches back far enoughreturn covers.lookupSeries(query.matchers)
def dropExpiredChunks(tier):cutoff = now() - tier.retentionDays * dayfor series in tier.allSeries():for chunk in series.chunks:if chunk.maxTime < cutoff:series.chunks.remove(chunk)disk.unlink(chunk.file)postings.maybeGC(series.id)# raw: 7d | rate5m: 90d | rate1h: 730d
Where this sits in Build a Prometheus-style time-series database
Scene 10 of 12. Aggregate old chunks into 5-minute, then 1-hour buckets, dropping originals as you go. Recording rules materialize — they cost storage.
Up next. We have a complete single-node TSDB. But what about replication, HA, and 100-node clusters? That story is shorter than you'd think.
All 12 scenes in Build a Prometheus-style time-series database · Every curriculum