Design canvas: pick a workload, ship a config

Every TSDB choice — scrape interval, label set, chunk size, retention, recording rules, downsampling — is a knob, and the right settings depend entirely on the workload's cardinality, write rate, and query range; some workloads have no honest TSDB configuration at all.

Previously

We have all the pieces. Now you build.

Scene 12

Design canvas: pick a workload, ship a config

  1. Watch
  2. Try it
  3. Predict
  4. Capture
WORKLOADCONFIGVERIFIERACTIVE1000-host fleet alerting1000 targets × 50 metrics × 15s scrape — bou…1000 targets · 50/tgt @ 15sRAM≈2.9 GB · disk≈6.0 GB/dCARDINALITY LOWPer-request distributed tracing50 services × 10 metrics × 1s scrape — but t…50 targets · 10/tgt @ 1sRAM≈78.1 GB · disk≈195.3 GB/dCARDINALITY HIGHBusiness KPIs · 1-week retenti…12 services × 200 metrics × 60s scrape — slo…12 targets · 200/tgt @ 60sRAM≈500 MB · disk≈200 MB/dCARDINALITY MEDscrape interval15schunk samples120head retention3hlabel setjobinstancemethodstatusrecording rules• job:http_request_duration_seconds:…downsampling tiers• raw 15s → 15d• 5m rollup → 90d✓FITS BUDGETprojection within budgetRAM2.9 GB / 8.0 GBDISK6.0 GB/dayQUERY50 ms typicalWARNINGSCardinality ≈ 1000 hosts × 50 metrics × ~4…Chunks: 120 samples / 15s = 30 min per chu…Honest fit for fleet alerting — every knob traces to a scene.
What to watch for

Workload A is on the canvas: 1000-host fleet alerting, 50 metrics per target, 15s scrape. Default config loads — chunk samples 120 (scene 4), head retention 3h (scene 5), one recording rule (scene 9), two downsampling tiers (scene 9). The verifier turns it green: ~3 GB RAM, ~6 GB/day disk, ~50 ms typical query.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Canvas.verify
the top-level fits-or-doesnt verdict the canvas renders
def verify(workload, config):
bomb = detectCardinalityBomb(workload, config)
if bomb:
return refuse(bomb.reason) # wrong tool
ram = projectRam(workload, config)
disk = projectDisk(workload, config)
latency = projectQueryLatency(config, queryRange)
fits = ram.mb <= ram.budget and disk.ok
warnings = ram.warnings + disk.warnings
return Verdict(fits, ram, disk, latency, warnings)
verifier.projectRam
head-block RAM is dominated by series count, not sample rate
def projectRam(workload, config):
series_count = product(
cardinality(label) for label in config.labelSet
) * workload.metricsPerTarget
head_bytes = series_count * BYTES_PER_HEAD_CHUNK
index_bytes = series_count * BYTES_PER_POSTINGS_ENTRY
ram_mb = (head_bytes + index_bytes) / MB
if ram_mb > RAM_BUDGET_MB:
return overshoot(ram_mb, RAM_BUDGET_MB)
return ok(ram_mb)
verifier.detectCardinalityBomb
scan the proposed label set for unbounded identifiers
UNBOUNDED = {
'user_id', 'request_id', 'trace_id',
'session_id', 'email', 'ip',
}
def detectCardinalityBomb(workload, config):
for label in config.labelSet:
if label in UNBOUNDED:
return Bomb(
reason=f'{label} is unbounded — wrong tool',
)
return None
verifier.projectQueryLatency
downsampling tiers turn long-range queries from O(M) into O(k)
def projectQueryLatency(config, queryRange):
tier = pickTier(config.downsamplingTiers, queryRange)
points = queryRange.seconds / tier.resolutionSeconds
return points * DECODE_COST_PER_POINT_MS
def pickTier(tiers, queryRange):
# coarsest tier whose retention covers the range
for t in sorted(tiers, by=resolution, desc=True):
if t.retentionDays * DAY >= queryRange.seconds:
return t
return tiers[0] # fall back to raw

Where this sits in Build a Prometheus-style time-series database

Scene 11 of 12. Capstone: alerting, tracing, or business KPIs — the verifier turns scrape interval, label set, retention, and rules into projected RAM, disk, and a fits/refuses verdict.

All 12 scenes in Build a Prometheus-style time-series database · Every curriculum

Built with Arqly
Every scene in Build a Prometheus-style time-series database builds on the one before it.All 12 Build a Prometheus-style time-series database scenes