Design your LSM — and feel its trades

Every LSM knob — memtable size, level multiplier, leveled vs tiered, bloom bits-per-key, block size, compression — is a deliberate move on the same write-amp / read-amp / space-amp triangle, and the right configuration is dictated by the workload, not by a default.

Previously

All ten knobs are on the table. The capstone asks: given a workload, which ones do you turn — and which scene's insight justifies each turn?

Scene 11

Design your LSM — and feel its trades

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Workload — OLTP mixed (10M keys × 1 KB, p99 read <5ms, disk budget 2× live)Mixed read/write OLTP. Disk budget 2× live data — space-amp matters. p99 read <5ms — read-amp matters. Write rate ~10 MB/s — write-amp tolerable.CLUSTERRF=1 · partitions=11memtable (64 MB)2PWALmemtableL0 (just-flushed)4PL0/AL0/BL0/CL0/DL1+5PL1/0L1/1L1/2L1/3L1/…COMPACTION POLICY: LEVELED · 3 consumersreader (bloom 10 bits/key, blo…owns memtable, L0/A, L1/0writer (memtable + WAL)owns WAL, memtablecompactor (leveled, lz4)owns L0/A, L0/B, L0/C +4EOS LAYEROFFPIDproducer idEpochtxn epochGroup-Genrebalance genTRADE-OFFSwrite-amp4/5read-amp2/5space-amp2/5WARNINGSleveled matches the workload (scene lsm-08). leveled + 10 b/kbloom + 4 KB blocks + lz4 — the OLTP default.
What to watch for

OLTP mixed: 10M keys × 1KB values, p99 read <5ms, disk budget 2× live data. Default knobs are pre-loaded; the verifier traces each to its scene.

Implementation

Highlighted lines are the ones running in the diagram right now.

LSM design checklist
knob → workload axis → earlier scene
memtable size → recovery time vs flush frequency (scenes 2, 3)
compaction policy → write-amp vs read/space-amp (scene 8)
level multiplier → #levels vs per-compaction work (scene 7)
bloom bits/key → miss-cost vs RAM (scene 5)
block size → point-read I/O vs compression (scene 10)
compression → CPU vs disk bytes (scene 10)
block_cache_size → hot-working-set fits in RAM? (scene 10)
tombstone retention → safe deletion vs space (scene 9)
When LSM is the wrong answer
scene 1 is still in scope
# fat values, low key cardinality, all point reads
# → Bitcask wins (scene 1 contrast)
# transactional multi-key OLTP
# → B-tree DB (Postgres, MySQL) — LSM compaction is a tax here
# columnar OLAP
# → Parquet + columnar engine — LSM cannot beat sort-merge-with-pushdown
# tiny dataset that fits in RAM
# → just use a hash table

Where this sits in Build an LSM-tree storage engine (LevelDB / RocksDB style)

Scene 11 of 11, in the Sharp edges act — Tombstones, blocks, and the canvas.. Capstone: pick a workload, set knobs, watch the verifier trace each choice back to the scene that taught it. The amp triangle is the load-bearing trade.

All 11 scenes in Build an LSM-tree storage engine (LevelDB / RocksDB style) · Every curriculum

Built with Arqly
Every scene in Build an LSM-tree storage engine (LevelDB / RocksDB style) builds on the one before it.All 11 Build an LSM-tree storage engine (LevelDB / RocksDB style) scenes