RAM is paid per key, not per byte

Every live key consumes a fixed-size keydir entry (~44.5 B + key length) regardless of value size, so RAM scales with key count alone — not with how much data you store.

Previously

Reads were cheap because the keydir was in RAM. The bill for that line is paid in keys, not bytes — and most workloads don't realise which side of that asymmetry they're on until they OOM.

Scene 04

RAM is paid per key, not per byte

  1. Watch
  2. Try it
  3. Predict
  4. Capture
WORKLOAD SIZING — RAM vs DISK1.0 KB1.0 MB1.00 GB1.00 TBRAM (keydir): 0 B RAM = N × (44.5 + keylen)RAM (keydir)RAM = N × (44.5 + keylen)0 B0 BDISK (live records): 0 BDISK (live records)0 B0 Bbudget 256.0 MBFITSRAM 0 B vs budget 256.0 MBWORKLOAD KNOBSkey count01k1Bavg key length32 B8 B256 Bavg value size4.0 KB16 B1 MBkeydir overhead — 44.5 B/key (Riak capacity calculator)WORKLOAD PINS(driven by sliders — not clickable)session cachesession cache1M × 4 KB — Bitcask's sweet spot1M × 4 KB — Bitcask's sweet spotkeys: 1.00Mkeylen: 32 Bvalue: 4.0 KBACTIVEfat blobsfat blobs1M × 1 MB — RAM tiny, disk huge1M × 1 MB — RAM tiny, disk hugekeys: 1.00Mkeylen: 32 Bvalue: 1.0 MBACTIVEtiny tagstiny tags1B × 100 B — RAM blows the budget1B × 100 B — RAM blows the budgetkeys: 1.00Bkeylen: 16 Bvalue: 100 BACTIVEStreaming writes — RAM grows with keys, disk grows with key×value. Both fit.
What to watch for

Default workload: 1M keys, 32 B keys, 4 KB values. Watch the RAM bar settle around 76 MB while the disk bar climbs to ~4 GB — same key count, two very different bars. Both fit under the 256 MB RAM budget line.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

KeydirEntry (per live key)
what's actually stored in RAM for each key
struct KeydirEntry {
file_id: u32 # which data file holds the value
value_sz: u32 # bytes to pread
value_pos: u64 # offset within file_id
tstamp: u32 # for tombstone / merge ordering
# 4 + 4 + 8 + 4 = 20 B payload
# + ~24.5 B Erlang term + HashMap bookkeeping
# → 44.5 B static + len(key) per entry
}
estimateKeydirRam
the one-line sizing rule for a Bitcask deployment
def estimateKeydirRam(keyCount, avgKeyLen):
static_overhead = 44.5 # Riak capacity calculator figure
bytes_per_entry = static_overhead + avgKeyLen
# value_size DOES NOT appear — the keydir only stores
# a pointer to the value, never the value itself.
return keyCount * bytes_per_entry
verdict
the budget check operators run before deploying
def verdict(keyCount, avgKeyLen, ramBudget):
needed = estimateKeydirRam(keyCount, avgKeyLen)
if needed > ramBudget:
return OOM # use an LSM instead
if needed > 0.75 * ramBudget:
return TIGHT # one traffic spike from OOM
return FITS

Where this sits in Build a Bitcask-style KV store

Scene 04 of 9, in the Limits act — RAM is paid per key; merge gives back what's dead.. Every live key consumes a fixed-size keydir entry (~44.5 B + key length). RAM scales with key count alone — fat values are free, tiny tags OOM.

Up next. RAM tracks live keys. Disk tracks every write you've ever made — including the overwritten ones. Without something to reclaim, disk grows forever.

All 9 scenes in Build a Bitcask-style KV store · Every curriculum

Built with Arqly
Every scene in Build a Bitcask-style KV store builds on the one before it.All 9 Build a Bitcask-style KV store scenes