Inverted index vs labels-only index
ELK tokenises every value in the line and stores each token in a per-term posting list — the inverted index; Loki hashes the label-set into a stream id and appends the raw line verbatim to that stream's chunk — so the same input becomes a heavy index plus small body on one side and a tiny index plus verbatim body on the other.
Same line, two vocabularies — fields in ELK, labels in Loki. Now watch what each system actually writes to disk when that one line lands.
Scene 05
Inverted index vs labels-only index
- Watch
- Try it
- Predict
- Capture
One source line, two disks. We want to find every line containing ERROR later — so we pre-build a lookup from word → list of lines containing it. That structure is called an inverted index, and it's what the LEFT (ELK) panel is filling, one token at a time. The RIGHT (Loki) panel does something different: it indexes ONLY the labels — {app=api,env=prod} — and appends the line VERBATIM to a per-stream blob called a chunk. Watch both sides land the same line.
Highlighted lines are the ones running in the diagram right now.
def index_doc(doc, segment):for field, value in doc.items():if mapping.fields_used >= 1000: # total_fields.limitraise MappingError('mapping explosion')mapping.ensure_field(field)for term in analyzer.tokenize(value):postings[term].append_sorted(doc.id)segment.store_source(doc) # original kept for _sourceif segment.size_bytes >= 5 * GB: # max_merged_segmentsegment.seal()
def write_line(labels, line, ts):stream_id = hash(canonicalize(labels))if stream_id not in streams:streams[stream_id] = Chunk(labels=labels)chunk = streams[stream_id]chunk.append(ts, line) # raw bytes, NOT parsedif chunk.compressed_size >= 1_500_000: # chunk_target_sizeflush_chunk(stream_id, chunk)
def flush_chunk(stream_id, chunk):# closes on size (1.5 MB compressed), age (2h), idle (30m)if chunk.idle_for() > 30 * MINUTES:chunk.seal()blob = chunk.encode() # compressed block-by-blockchunk_id = object_store.put(blob)# ONE index row — labels only, body never parsedindex.append(stream_id, chunk.time_range, chunk_id)del streams[stream_id]
Where this sits in Build a distributed logging stack (ELK / Loki)
Scene 05 of 12. ELK tokenises every value into a per-term posting list; Loki hashes the label-set into a stream id and appends the line verbatim — heavy index + small body vs tiny index + verbatim body.
Up next. We have segments full of inverted-index entries on one side and chunks plus a tiny labels-only index on the other. Same query — service=api ERROR last 1h — runs through two completely different machines.
All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum