Append, fsync, update, ack

Every put is four ordered steps — append at end-of-file, fsync per policy, update keydir in place, return ack — and the single open writer (enforced at open(), not at write time) makes that ordering the throughput ceiling of one Bitcask instance.

Previously

The split is the shape; the order is the contract. Append, then fsync, then keydir update, then ack — and one open writer at a time. That ordering is what makes the disk and the RAM agree.

Scene 02

Append, fsync, update, ack

  1. Watch
  2. Try it
  3. Predict
  4. Capture
DATA DIRECTORY001.dataclosed · immutablecrc:7a|t01|k…k0 = hellocrc:c2|t02|k…k1 = first002.dataclosed · immutablecrc:f3|t05|k…k2 = alpha003.dataactive · acceptingKEYDIR (in RAM)keyfile_idvalue_posvalue_szk000108Bk2002016BOPERATIONidle — no operation in flightTHROUGHPUT0 ops/s
What to watch for

A single producer issues PUT k1=v1. Watch the four steps light up in order — append at the tail, fsync, keydir row updated, ack to the client — and the throughput meter climb as more puts stream in.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Bitcask.put_record
the four ordered steps every put goes through
def put_record(key, value):
rec = encode(crc, tstamp, ksz, vsz, key, value)
# 1) append at end-of-file on the active file
pos = active_file.size
active_file.write(rec) # sequential append
# 2) fsync per sync_strategy (none | o_sync | interval)
if sync_strategy == 'o_sync':
active_file.fsync()
# 3) update keydir entry in place
keydir[key] = (active_file.id, vsz, pos, tstamp)
# 4) ack to the client
return ok
Bitcask.open_for_write
the writer lock is acquired once, at open() — not per write
def open(dir, mode):
if mode == read_write:
# one writer per directory, enforced here
lock = try_flock(dir + '/bitcask.write.lock')
if lock is None:
return error('already open for writing')
active_file = open_or_create_active(dir)
keydir = scan_or_load_hint_files(dir)
return Handle(dir, active_file, keydir, lock)
# read-only handles share freely
return ReadOnlyHandle(dir)
Bitcask.maybe_rotate
active file crosses max_file_size: close it, open the next one
def maybe_rotate(active_file):
if active_file.size < max_file_size: # default 2 GB
return active_file
# 1) close the current active file — now immutable
active_file.close()
immutable_files.append(active_file)
# 2) open a fresh active file for subsequent appends
next_id = active_file.id + 1
return create_active(dir, next_id)
# rotation is not merge: closed bytes are unchanged here

Where this sits in Build a Bitcask-style KV store

Scene 02 of 9, in the Anatomy act — Log on disk, hash in RAM — and the four-step write.. Every put is four ordered steps on one open writer. Concurrent writers are rejected at open(), not at write time.

Up next. Writes funnel through one append + one fsync. Reads don't share that bottleneck — the keydir is random-access, so a get is bounded by the laziest part of the OS, not by the writer.

All 9 scenes in Build a Bitcask-style KV store · Every curriculum

Built with Arqly
Every scene in Build a Bitcask-style KV store builds on the one before it.All 9 Build a Bitcask-style KV store scenes