Inverted index vs labels-only index

ELK tokenises every value in the line and stores each token in a per-term posting list — the inverted index; Loki hashes the label-set into a stream id and appends the raw line verbatim to that stream's chunk — so the same input becomes a heavy index plus small body on one side and a tiny index plus verbatim body on the other.

Previously

Same line, two vocabularies — fields in ELK, labels in Loki. Now watch what each system actually writes to disk when that one line lands.

Scene 05

Inverted index vs labels-only index

  1. Watch
  2. Try it
  3. Predict
  4. Capture
2026-05-09T12:00:01Z lvl=ERROR svc=api user_id=42 event=checkout_failedSOURCE · api pod · prodsame line, two index strategiesELK · Lucene segmentinverted indexmapping47 / 1000 fieldssegment 003 (immutable, max 5 GB)TERMPOSTING LIST (sorted doc-ids)api[]ERROR[]checkout_failed[]42[]user_id[]BYTES ON DISK1.30× raw1.0× rawLoki · chunkslabels-only indexchunk {app=api,env=prod}0 B / 1.43 MBINDEX (labels → chunk_id){app=api,env=prod}→ chunk-001BYTES ON DISK0.20× raw1.0× rawOne source line. Watch it fork: tokens fly into posting lists on the LEFT (inverted index); the SAME line is appended verbatim to a chunk on the RIGHT and ONE labels-only index entry is touched.
What to watch for

One source line, two disks. We want to find every line containing ERROR later — so we pre-build a lookup from word → list of lines containing it. That structure is called an inverted index, and it's what the LEFT (ELK) panel is filling, one token at a time. The RIGHT (Loki) panel does something different: it indexes ONLY the labels — {app=api,env=prod} — and appends the line VERBATIM to a per-stream blob called a chunk. Watch both sides land the same line.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Lucene.index_doc
ELK side: tokenise every field, grow a posting list per term
def index_doc(doc, segment):
for field, value in doc.items():
if mapping.fields_used >= 1000: # total_fields.limit
raise MappingError('mapping explosion')
mapping.ensure_field(field)
for term in analyzer.tokenize(value):
postings[term].append_sorted(doc.id)
segment.store_source(doc) # original kept for _source
if segment.size_bytes >= 5 * GB: # max_merged_segment
segment.seal()
Loki.write_line
Loki side: hash label-set → stream_id; append line verbatim
def write_line(labels, line, ts):
stream_id = hash(canonicalize(labels))
if stream_id not in streams:
streams[stream_id] = Chunk(labels=labels)
chunk = streams[stream_id]
chunk.append(ts, line) # raw bytes, NOT parsed
if chunk.compressed_size >= 1_500_000: # chunk_target_size
flush_chunk(stream_id, chunk)
Loki.flush_chunk
close on size or idle, upload, write ONE index entry
def flush_chunk(stream_id, chunk):
# closes on size (1.5 MB compressed), age (2h), idle (30m)
if chunk.idle_for() > 30 * MINUTES:
chunk.seal()
blob = chunk.encode() # compressed block-by-block
chunk_id = object_store.put(blob)
# ONE index row — labels only, body never parsed
index.append(stream_id, chunk.time_range, chunk_id)
del streams[stream_id]

Where this sits in Build a distributed logging stack (ELK / Loki)

Scene 05 of 12. ELK tokenises every value into a per-term posting list; Loki hashes the label-set into a stream id and appends the line verbatim — heavy index + small body vs tiny index + verbatim body.

Up next. We have segments full of inverted-index entries on one side and chunks plus a tiny labels-only index on the other. Same query — service=api ERROR last 1h — runs through two completely different machines.

All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum

Built with Arqly
Every scene in Build a distributed logging stack (ELK / Loki) builds on the one before it.All 12 Build a distributed logging stack (ELK / Loki) scenes