A part — one batch, frozen on disk — sparse primary index inside an immutable part

Each batched write lands as a self-contained immutable directory of column files plus an index — a part — and a table is simply a collection of parts that a query has to ask one by one.

Previously

If every batch becomes its own self-contained file on disk, we need a name and a shape for that unit — because everything operational from here on out is about how many of these units a query has to touch.

Scene 06

A part: one batch, frozen on disk

  1. Watch
  2. Try it
  3. Predict
  4. Capture
TABLEevents0 partsPART DIRECTORYno part selectedtable is empty — about to write the first batch
What to watch for

Write one batch. It lands as a single self-contained part directory — one .bin per column, a primary.idx, and metadata. Write a second batch and a second part appears. The table is now exactly two parts; neither will ever be mutated in place.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

MergeTree.write_batch_to_part
one batch in → one self-contained part directory out
def write_batch_to_part(rows, columns, sort_key, part_id):
rows = sort(rows, by=sort_key)
part_dir = mkdir(f'{table_path}/{part_id}')
for col in columns: # one .bin per column
values = project(rows, col)
write_compressed(f'{part_dir}/{col}.bin', values)
write_marks(f'{part_dir}/{col}.mrk2', values)
write_primary_idx(f'{part_dir}/primary.idx',
first_of_each_granule(rows, sort_key))
write_text(f'{part_dir}/columns.txt', columns)
write_text(f'{part_dir}/checksums.txt', checksums(part_dir))
fsync_dir(part_dir) # part is now immutable
MergeTree.query_table
a table is a bag of parts — fan out, then merge results
def query_table(table, predicate, projection):
parts = list_parts(table) # every part on disk
streams = []
for part in parts: # ask every part
if part.minmax_excludes(predicate):
continue # pruned by metadata
granules = scan_primary_idx(part, predicate)
for col in projection:
streams.append(read_bin(part, col, granules))
return merge_streams(streams)
MergeTree.update_row
UPDATE is a part rewrite, not a row edit
def update_row(table, predicate, assignments):
# parts are immutable — we never reopen an existing .bin
matched = []
for part in list_parts(table):
rows = scan(part, predicate)
for r in rows:
matched.append(apply(r, assignments))
# write a NEW part containing the rewritten rows
new_id = next_part_id(table) # e.g. all_11_11_0
write_batch_to_part(matched, columns, sort_key, new_id)
# original part stays on disk, untouched, until merge

Where this sits in Build a columnar OLAP store (ClickHouse / Druid style)

Scene 06 of 13, in the Write side act — Bulk inserts → immutable parts → background merge → too-many-parts cliff.. Each batched write lands as an immutable directory of column files plus an index — a part. A table is a stack of parts.

Up next. Many small parts pile up fast. If the count grows unbounded, every query has to interrogate every part — there must be a background process that fuses them, or the engine collapses.

All 13 scenes in Build a columnar OLAP store (ClickHouse / Druid style) · Every curriculum

Built with Arqly
Every scene in Build a columnar OLAP store (ClickHouse / Druid style) builds on the one before it.All 13 Build a columnar OLAP store (ClickHouse / Druid style) scenes