Design your LSM — and feel its trades
Every LSM knob — memtable size, level multiplier, leveled vs tiered, bloom bits-per-key, block size, compression — is a deliberate move on the same write-amp / read-amp / space-amp triangle, and the right configuration is dictated by the workload, not by a default.
All ten knobs are on the table. The capstone asks: given a workload, which ones do you turn — and which scene's insight justifies each turn?
Scene 11
Design your LSM — and feel its trades
- Watch
- Try it
- Predict
- Capture
OLTP mixed: 10M keys × 1KB values, p99 read <5ms, disk budget 2× live data. Default knobs are pre-loaded; the verifier traces each to its scene.
Highlighted lines are the ones running in the diagram right now.
memtable size → recovery time vs flush frequency (scenes 2, 3)compaction policy → write-amp vs read/space-amp (scene 8)level multiplier → #levels vs per-compaction work (scene 7)bloom bits/key → miss-cost vs RAM (scene 5)block size → point-read I/O vs compression (scene 10)compression → CPU vs disk bytes (scene 10)block_cache_size → hot-working-set fits in RAM? (scene 10)tombstone retention → safe deletion vs space (scene 9)
# fat values, low key cardinality, all point reads# → Bitcask wins (scene 1 contrast)# transactional multi-key OLTP# → B-tree DB (Postgres, MySQL) — LSM compaction is a tax here# columnar OLAP# → Parquet + columnar engine — LSM cannot beat sort-merge-with-pushdown# tiny dataset that fits in RAM# → just use a hash table
Where this sits in Build an LSM-tree storage engine (LevelDB / RocksDB style)
Scene 11 of 11, in the Sharp edges act — Tombstones, blocks, and the canvas.. Capstone: pick a workload, set knobs, watch the verifier trace each choice back to the scene that taught it. The amp triangle is the load-bearing trade.
All 11 scenes in Build an LSM-tree storage engine (LevelDB / RocksDB style) · Every curriculum