A chunk: 120 points, packed

A CHUNK (the new term) is a fixed-size bundle of consecutive points from one series — bit-packed into one blob with one header — and Prometheus closes it at 120 samples or 2 hours, whichever comes first, because that is where Gorilla's 1.37 B/point compression averages out (vs 16 B/point naive: a ~12× win).

Previously

We have two encodings, each almost magical alone. A chunk is where they live together.

Scene 05

A chunk: 120 points, packed

  1. Watch
  2. Try it
  3. Predict
  4. Capture
CHUNKcpu_usage_seconds_totalseries#S#73a1 · {job=api,instance=10.0.0.7,mode=user}0/120 samples · cap 120mTIMELINE · 0 samples (cap 120 or 120m)0306090120cap = 120PACKED BITSTRINGHEADERsid S#73a1ts0 20:06:40v0 0.42full encoding · once per chunkfirst 1 samples · 144 bitstotal 144 bits packed · 18.0 BCOST · 0 samplesNaive (uncompressed)16 B/point × 01 BCompressed (ΔΔ + XOR)~0.00 B/point × 00 BRATIO0.00×smallerOne series, one chunk-to-be. The timeline is empty; the bitstring shows only the 16-byte header.
What to watch for

New term this scene: CHUNK — a fixed-size bundle of compressed points from one series. Watch 120 samples accumulate, the bitstring pack in lockstep, and the chunk close. Then the bars draw and the 12× ratio appears.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Chunk
one series, one header, one bit-packed body
class Chunk:
# --- header (16 B, written once) ---
series_id: uint64 # which series
first_ts: int64 # anchor for delta-of-delta
first_value: float64 # anchor for XOR
# --- body (bit-packed, grows append-only) ---
bits: BitBuffer # (ts_bits, val_bits) pairs
samples: int = 0 # cap at SAMPLES_CAP (120)
span_min: int = 0 # cap at TIME_CAP_MIN (120)
closed: bool = False # immutable once true
Chunk.append
encode dod + xor; close on whichever cap fires first
def append(chunk, ts, value):
dod = (ts - chunk.last_ts) - chunk.last_delta
chunk.bits.write(encode_dod(dod)) # ~1 bit on cadence
xor = float_bits(value) ^ chunk.last_value_bits
chunk.bits.write(encode_xor(xor)) # 1 bit if unchanged
chunk.samples += 1
chunk.span_min = (ts - chunk.first_ts) // 60
if chunk.samples >= SAMPLES_CAP \
or chunk.span_min >= TIME_CAP_MIN:
close_chunk(chunk)
close_chunk
finalize, mark immutable, hand to mmap
def close_chunk(chunk):
chunk.bits.flush_byte_aligned() # body ends on byte boundary
chunk.closed = True # no more appends, ever
# one chunk = the unit of compression / mmap / I/O
file = chunks_dir / f"{chunk.series_id}-{chunk.first_ts}"
file.write(chunk.header || chunk.bits)
mmap_readonly(file) # query path reads via mmap
# head opens a fresh chunk for this series and keeps scraping
return Chunk(series_id=chunk.series_id)

Where this sits in Build a Prometheus-style time-series database

Scene 05 of 12. Bundle ~120 consecutive points into a single bit-packed blob. Gorilla: 16 B/point → 1.37 B/point — about 12× compression.

Up next. A chunk is great when it's full. But points arrive one at a time — where does a not-yet-full chunk live, and what happens if the process crashes mid-chunk?

All 12 scenes in Build a Prometheus-style time-series database · Every curriculum

Built with Arqly
Every scene in Build a Prometheus-style time-series database builds on the one before it.All 12 Build a Prometheus-style time-series database scenes