Head chunks, WAL, and flushing
Every sample lands in two places before the ack — an append-only WAL on disk (so a crash mid-chunk loses nothing) and the matching HEAD CHUNK in RAM (so writes are sub-millisecond and queryable); when a head fills to 120 samples it seals into a read-only mmap'd file on disk and a fresh head opens.
A chunk only compresses well after it is full — but writes don't wait. We need a place for the half-built chunk that's fast to append to AND survives a crash.
Scene 06
Head chunks, WAL, and the durable ack
- Watch
- Try it
- Predict
- Capture
A new sample arrives for series S3. Watch it fork: one copy appends to the WAL on disk, one fills S3's head chunk in RAM. When S3 hits 120/120 it seals — the chunk slides down to the mmap'd disk row and a fresh empty head takes its place.
Highlighted lines are the ones running in the diagram right now.
def appendSample(seriesId, ts, value):entry = (seriesId, ts, value)if wal_enabled:walAppend(entry) # disk firstheadAppend(seriesId, ts, value) # RAM secondreturn ok # ack: both have it
def walAppend(entry):record = encode(entry)wal_file.write(record) # append-only, sequentialif fsync_policy == 'always':wal_file.fsync() # durable before ack# else: OS flushes on its own schedule
def headAppend(seriesId, ts, value):head = head_chunks[seriesId] # one per active serieshead.samples.append((ts, value)) # sub-ms in-memoryif head.samples_in == head.capacity:sealHeadChunk(seriesId) # write file + mmap ROdef sealHeadChunk(seriesId):head = head_chunks[seriesId]path = write_compressed(head) # gorilla-encoded bytessealed_chunks.append(mmap_readonly(path))head_chunks[seriesId] = new_head() # fresh, empty
def recoverFromCrash():head_chunks = {} # RAM was wipedfor entry in wal.scan_sequentially():headAppend(*entry) # rebuild fill levelsfor path in chunks_dir.list():sealed_chunks.append(mmap_readonly(path))# without WAL: partial heads are gone forever
Where this sits in Build a Prometheus-style time-series database
Scene 06 of 12. Active chunk lives in RAM (the head); a write-ahead log on disk catches every sample so a crash mid-chunk loses nothing.
Up next. Writes are settled. Now flip the system around: a query arrives asking for {method=POST, status=500} over the last hour — how does the database even find the right series?
All 12 scenes in Build a Prometheus-style time-series database · Every curriculum