A part — one batch, frozen on disk — sparse primary index inside an immutable part
Each batched write lands as a self-contained immutable directory of column files plus an index — a part — and a table is simply a collection of parts that a query has to ask one by one.
If every batch becomes its own self-contained file on disk, we need a name and a shape for that unit — because everything operational from here on out is about how many of these units a query has to touch.
Scene 06
A part: one batch, frozen on disk
- Watch
- Try it
- Predict
- Capture
Write one batch. It lands as a single self-contained part directory — one .bin per column, a primary.idx, and metadata. Write a second batch and a second part appears. The table is now exactly two parts; neither will ever be mutated in place.
Highlighted lines are the ones running in the diagram right now.
def write_batch_to_part(rows, columns, sort_key, part_id):rows = sort(rows, by=sort_key)part_dir = mkdir(f'{table_path}/{part_id}')for col in columns: # one .bin per columnvalues = project(rows, col)write_compressed(f'{part_dir}/{col}.bin', values)write_marks(f'{part_dir}/{col}.mrk2', values)write_primary_idx(f'{part_dir}/primary.idx',first_of_each_granule(rows, sort_key))write_text(f'{part_dir}/columns.txt', columns)write_text(f'{part_dir}/checksums.txt', checksums(part_dir))fsync_dir(part_dir) # part is now immutable
def query_table(table, predicate, projection):parts = list_parts(table) # every part on diskstreams = []for part in parts: # ask every partif part.minmax_excludes(predicate):continue # pruned by metadatagranules = scan_primary_idx(part, predicate)for col in projection:streams.append(read_bin(part, col, granules))return merge_streams(streams)
def update_row(table, predicate, assignments):# parts are immutable — we never reopen an existing .binmatched = []for part in list_parts(table):rows = scan(part, predicate)for r in rows:matched.append(apply(r, assignments))# write a NEW part containing the rewritten rowsnew_id = next_part_id(table) # e.g. all_11_11_0write_batch_to_part(matched, columns, sort_key, new_id)# original part stays on disk, untouched, until merge
Where this sits in Build a columnar OLAP store (ClickHouse / Druid style)
Scene 06 of 13, in the Write side act — Bulk inserts → immutable parts → background merge → too-many-parts cliff.. Each batched write lands as an immutable directory of column files plus an index — a part. A table is a stack of parts.
Up next. Many small parts pile up fast. If the count grows unbounded, every query has to interrogate every part — there must be a background process that fuses them, or the engine collapses.
All 13 scenes in Build a columnar OLAP store (ClickHouse / Druid style) · Every curriculum