Head chunks, WAL, and flushing

Every sample lands in two places before the ack — an append-only WAL on disk (so a crash mid-chunk loses nothing) and the matching HEAD CHUNK in RAM (so writes are sub-millisecond and queryable); when a head fills to 120 samples it seals into a read-only mmap'd file on disk and a fresh head opens.

Previously

A chunk only compresses well after it is full — but writes don't wait. We need a place for the half-built chunk that's fast to append to AND survives a crash.

Scene 06

Head chunks, WAL, and the durable ack

  1. Watch
  2. Try it
  3. Predict
  4. Capture
RAM — head chunksS1http_requests_totalmethod=GET,status=20030/120S2http_requests_totalmethod=GET,status=50060/120S3http_requests_totalmethod=POST,status=200119/120S4node_cpu_seconds_totalcpu=0,mode=user90/120WAL — append-only log on diskS1S4S2S1S2S4S3S1S2S3S4S2S1S3Sealed chunks — mmapped read-onlyseal-S1-aS1120 samples165 B · mmapseal-S2-aS2120 samples158 B · mmapseal-S4-aS4120 samples172 B · mmapRAM (head chunks, one per active series) · disk WAL (every sample, append-only) · disk sealed chunks (mmap, read-only).
What to watch for

A new sample arrives for series S3. Watch it fork: one copy appends to the WAL on disk, one fills S3's head chunk in RAM. When S3 hits 120/120 it seals — the chunk slides down to the mmap'd disk row and a fresh empty head takes its place.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

TSDB.appendSample
fork: WAL on disk first, then head chunk in RAM, then ack
def appendSample(seriesId, ts, value):
entry = (seriesId, ts, value)
if wal_enabled:
walAppend(entry) # disk first
headAppend(seriesId, ts, value) # RAM second
return ok # ack: both have it
WAL.walAppend
sequential disk write; fsync per durability policy
def walAppend(entry):
record = encode(entry)
wal_file.write(record) # append-only, sequential
if fsync_policy == 'always':
wal_file.fsync() # durable before ack
# else: OS flushes on its own schedule
Head.headAppend
into the in-RAM chunk; if cap hit, seal and open fresh
def headAppend(seriesId, ts, value):
head = head_chunks[seriesId] # one per active series
head.samples.append((ts, value)) # sub-ms in-memory
if head.samples_in == head.capacity:
sealHeadChunk(seriesId) # write file + mmap RO
def sealHeadChunk(seriesId):
head = head_chunks[seriesId]
path = write_compressed(head) # gorilla-encoded bytes
sealed_chunks.append(mmap_readonly(path))
head_chunks[seriesId] = new_head() # fresh, empty
TSDB.recoverFromCrash
replay WAL into fresh heads; sealed chunks reopened mmap
def recoverFromCrash():
head_chunks = {} # RAM was wiped
for entry in wal.scan_sequentially():
headAppend(*entry) # rebuild fill levels
for path in chunks_dir.list():
sealed_chunks.append(mmap_readonly(path))
# without WAL: partial heads are gone forever

Where this sits in Build a Prometheus-style time-series database

Scene 06 of 12. Active chunk lives in RAM (the head); a write-ahead log on disk catches every sample so a crash mid-chunk loses nothing.

Up next. Writes are settled. Now flip the system around: a query arrives asking for {method=POST, status=500} over the last hour — how does the database even find the right series?

All 12 scenes in Build a Prometheus-style time-series database · Every curriculum

Built with Arqly
Every scene in Build a Prometheus-style time-series database builds on the one before it.All 12 Build a Prometheus-style time-series database scenes