Downsampling — a retention pyramid

Long-range queries stay tractable by pre-aggregating old chunks into coarser tiers (5 m, 1 h) and dropping the originals — but each tier is materialized as a real series with its own head chunk, postings, and disk chunks, not a query-time hint.

Previously

Cardinality bounds width. Retention bounds depth — and downsampling is how depth stays affordable.

Scene 10

Downsampling: a retention pyramid

  1. Watch
  2. Try it
  3. Predict
  4. Capture
RETENTION PYRAMID · raw → downsampled tiersRaw / 15 s scrape15 s · 7d · 40.3 k pts/seriesdense5-minute rollup5 m · 90d · 25.9 k pts/series+ job:http_requests:rate5mmedium1-hour rollup1 h · 2y · 17.5 k pts/series+ job:http_requests:rate1hsparsequery: 1d ago → nownow (0d)← 2y ago90d1.0yQUERY READOUTpoints decoded5.8 kactive tierRaw / 15 s scrapeSTORAGE PER TIER15 s55.0 MB5 m35.0 MB1 h24.0 MBtotal114.0 MBQuery: last 1 day → raw tier serves (15 s resolution, every scrape preserved).
What to watch for

Watch a query range slide across three tiers. A 1-day range is served by the raw tier; a 60-day range falls through to the 5 m rollup; a 1-year range drops to the 1 h rollup. The engine always picks the coarsest tier that fully covers the range.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

RuleEngine.evaluateRecordingRule
every rule.interval, write the result back as a NEW series
def evaluateRecordingRule(rule, window):
# rule.expr e.g. rate(http_requests_total[5m])
samples = promql.eval(rule.expr, window)
for labels, value in samples:
series = tsdb.getOrCreateSeries(
name = rule.name, # job:http_requests:rate5m
labels = labels,
)
series.headChunk.append(now(), value)
postings.index(series.id, labels)
# real bytes, real cardinality, real cost
QueryEngine.pickTier
find the coarsest tier that fully covers query.range
def pickTier(query):
# tiers ordered finest -> coarsest
tiers = [raw, rate5m, rate1h]
for tier in reversed(tiers): # try coarsest first
if tier.retentionDays >= query.range.days:
covers = tier
# coarsest tier that still reaches back far enough
return covers.lookupSeries(query.matchers)
Retention.dropExpiredChunks
per-tier retention — old chunks get unlinked
def dropExpiredChunks(tier):
cutoff = now() - tier.retentionDays * day
for series in tier.allSeries():
for chunk in series.chunks:
if chunk.maxTime < cutoff:
series.chunks.remove(chunk)
disk.unlink(chunk.file)
postings.maybeGC(series.id)
# raw: 7d | rate5m: 90d | rate1h: 730d

Where this sits in Build a Prometheus-style time-series database

Scene 10 of 12. Aggregate old chunks into 5-minute, then 1-hour buckets, dropping originals as you go. Recording rules materialize — they cost storage.

Up next. We have a complete single-node TSDB. But what about replication, HA, and 100-node clusters? That story is shorter than you'd think.

All 12 scenes in Build a Prometheus-style time-series database · Every curriculum

Built with Arqly
Every scene in Build a Prometheus-style time-series database builds on the one before it.All 12 Build a Prometheus-style time-series database scenes