Why one row per point is wrong — per-row label repetition and 60-byte overhead
Storing each point as a SQL row spends 60+ bytes of metadata to carry an 8-byte float, because the label set is repeated on every row of an append-only firehose.
We have a clean four-part tuple — now we try to store it the obvious way and watch the bytes explode.
Scene 02
Why one row per point is wrong
- Watch
- Try it
- Predict
- Capture
Six scrapes of the same series stream in one at a time. Watch the labels JSON column (red) take more bytes than the row header, metric, ts, and value combined — every single row.
Highlighted lines are the ones running in the diagram right now.
def naiveInsert(point):# rowHeader 8 + metric 22 + labels 60 + ts 8 + value 8# = ~106 B per row to carry an 8-byte floatdb.execute('INSERT INTO points',' (metric, labels, ts, value)',' VALUES (?, ?, ?, ?)',point.metric, # 'http_requests_total'json(point.labels), # full label set, every rowpoint.ts,point.value,)
labelDict = {} # labelSet -> seriesId (16-bit)def dedupeInsert(point):key = canonical(point.labels)if key not in labelDict:labelDict[key] = nextSeriesId()sid = labelDict[key] # 2-byte refdb.execute('INSERT INTO points (sid, ts, value)',' VALUES (?, ?, ?)',sid, point.ts, point.value,)# row shrinks to ~28 B: header + 2B ref + ts + value
def projectStorage(seriesCount, labelBytes, scrapesPerDay,):rowHeader, metric, ts, value = 8, 22, 8, 8perRow = rowHeader + metric + labelBytes + ts + valueperSeriesPerDay = perRow * scrapesPerDayfleet = perSeriesPerDay * seriesCountreturn perSeriesPerDay, fleet# perRow is paid on every scrape — it doesn't amortize.# fleet scales the per-series number by seriesCount.
Where this sits in Build a Prometheus-style time-series database
Scene 02 of 12. Storing each point as a SQL row spends 60+ bytes of metadata to carry an 8-byte float — the labels JSON repeats on every row of the firehose.
Up next. The labels are the easy waste to fix — store them once. But timestamps are 8 bytes each forever, and there are billions of them. Can we shrink those too?
All 12 scenes in Build a Prometheus-style time-series database · Every curriculum