fsync — pick two of safe, fast, simple

sync_strategy is a three-way trade: o_sync, interval=N, none. CRCs make the surviving log valid in every case — fsync controls recency, not corruption.

Previously

CRCs answered VALIDITY — the log is parseable no matter what. fsync answers RECENCY — how much of the most-recent tail you keep. Two separate guarantees, and Bitcask makes you pick the recency budget.

Scene 06a

fsync — pick two of safe, fast, simple

  1. Watch
  2. Try it
  3. Predict
  4. Capture
DATA DIRECTORY001.dataclosed · immutablecrc:7a|t01|k…k0 = hellocrc:c2|t02|k…k1 = first002.dataclosed · immutable003.dataactive · acceptingfsync: intervalKEYDIR (in RAM)keyfile_idvalue_posvalue_szk000108Bk10016012BOPERATION100,000 writes/sec → 003.datapolicy: interval=1scurrent policy →o_syncevery write fsyncedinterval=1snoneOS-controlledTHROUGHPUT100.0k ops/s100k writes/sec stream into the active file. Yellow glow = bytes the OS hasn't fsynced yet — the loss budget if the process dies.
What to watch for

A 100k writes/sec stream lands in the active file. Watch the durable-up-to marker walk behind the write head; the yellow glow between them is the vulnerable region — bytes the OS hasn't flushed to disk yet. The interval=1s policy is showing here; you'll switch policies in a moment.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Writer.append_with_policy
the put hot path — fsync is RECENCY, branched on sync_strategy
def append_with_policy(record, policy):
active_file.write(encode(record)) # append to tail
if policy == 'o_sync':
os.fsync(active_file.fd) # block until on platter
durable_up_to = active_file.tell()
elif policy == 'interval':
pass # interval_fsync_thread handles it
else: # 'none' — Riak default
pass # OS writeback decides; loss window 5-30s
return ack
IntervalFsync.run
background flusher — the recency window is now - last_fsync
def interval_fsync_thread(active_file, interval_s):
while running:
sleep(interval_s)
os.fsync(active_file.fd)
last_fsync_time = now()
durable_up_to = active_file.tell()
# vulnerable bytes = bytes appended since last_fsync_time;
# bounded by interval_s seconds at line rate
Recovery.validate_log
startup scan — CRC is VALIDITY, runs regardless of policy
def validate_log(file):
valid_end = 0
for record, offset in scan(file):
expected = crc32(record.body)
if record.crc != expected:
break # torn tail — drop everything from here
valid_end = offset + record.size
file.truncate(valid_end) # log is now parseable
return valid_end

Where this sits in Build a Bitcask-style KV store

Scene 06a of 9, in the Crash & sync act — Recovery, hint files, fsync — pick two.. sync_strategy is a three-way trade (o_sync / interval / none). fsync controls RECENCY; CRC controls VALIDITY — two separate guarantees.

Up next. fsync decides what survives the crash. The next question: what survives merge — and what's the bug when merge and the crash recovery scanner disagree about a deleted key?

All 9 scenes in Build a Bitcask-style KV store · Every curriculum

Built with Arqly
Every scene in Build a Bitcask-style KV store builds on the one before it.All 9 Build a Bitcask-style KV store scenes