Vectorized execution — process batches, not tuples
Pulling one row at a time through a chain of virtual function calls (Volcano) is interpreter overhead; processing 8192 column values per call lets the compiler emit SIMD inner loops and amortizes dispatch cost across the batch.
The bytes on disk are now tiny. But pulling tiny bytes through the CPU one row at a time still leaves a 50x performance win on the table — it depends on how the executor walks those bytes.
Scene 04
Vectorized execution: process batches, not tuples
- Watch
- Try it
- Predict
- Capture
Same query on both panels: scan a column, filter, aggregate over 1M rows. The top panel walks one tuple at a time; the bottom panel walks 8192-row batches. Watch the race bar — and the virtual-call counter on each side.
Highlighted lines are the ones running in the diagram right now.
def execute(plan):op = plan.root # Aggregate -> Filter -> Scanwhile True:row = op.next() # virtual dispatch, every rowif row is EOF:break# 3M virtual calls for 1M rows (3 operators deep)# SIMD lanes idle — no contiguous array to vectorizeaccumulate(row)
def execute(plan):op = plan.root # operators consume + produce Blockswhile True:block = op.nextBlock() # 8192 column values per callif block is EOF:break# one virtual call per BATCH, not per row# block.col is a contiguous array -> SIMD kernelaccumulateBlock(block)
def apply(block):if predicate.isNativeKernel():# tight loop -> AVX2 (8 int32) / AVX-512 (16)for i in range(block.n):out[i] = block.col[i] > thresholdreturn block.select(out)# opaque UDF: planner can't prove what it doesfor i in range(block.n):out[i] = python_udf(block.col[i]) # FFI per rowreturn block.select(out)
Where this sits in Build a columnar OLAP store (ClickHouse / Druid style)
Scene 04 of 13, in the Speedups act — Compression and vectorized execution — where the orders of magnitude live.. Tuple-at-a-time Volcano is interpreter overhead; processing 1024–8192 column values per call lets the CPU emit SIMD inner loops.
Up next. Vectorized execution wants thousands of values per call. That sets a hard rule for writes too — they have to arrive in batches, not one row at a time, or the per-column overhead crushes the engine.
All 13 scenes in Build a columnar OLAP store (ClickHouse / Druid style) · Every curriculum