Retention vs deletion — the index has to forget — delete-by-query tombstones and manual force-merge

Retention is when the system stops promising you can read; deletion is when the bytes are physically gone — and the index that points at the bytes is a separate structure with its own lifecycle, which is where compliance bugs live.

Previously

We have a cost ladder for where bytes live as they age. Aging through the ladder is the ILM policy; but the policy decides when bytes MOVE, not when bytes DIE — and 'die' has its own lifecycle.

Scene 08

Retention vs deletion — the index has to forget

  1. Watch
  2. Try it
  3. Predict
  4. Capture
BASELINE · 30-day retentionDATA RECORDS · on-disk segmentscompactor last ran day 26 · lag 4dday 26r26EXPIREDday 27r27EXPIREDday 28r28EXPIREDday 29r29EXPIREDday 30r30day 31r31day 32r32day 33r33day 34r34cutoff · day 30INDEX ENTRIES · pointers into recordsages independently of the records they point atage 8d→ r26age 7d→ r27age 6d→ r28age 5d→ r29age 4d→ r30age 3d→ r31age 2d→ r32age 1d→ r33age 0d→ r34COMPACTORlast cycle · day 26 · lag 4d (expired bytes still on disk)Baseline retention: cluster default 30 days. The compactor cycles daily and trims expired records; the index shrinks in lockstep. Retention says "we don't promise reads past day 30"; deletion (bytes physically gone) only happens after the compactor runs.
Retention cut-off (day 30) — system stops promising reads
Expired but still on disk — compactor hasn't reached here yet
What to watch for

Watch the two timelines. The TOP row is the data on disk — each cell is a day of records. A retention cut-off line sits at day 30: anything to its left is retention-expired but still physically present. The BOTTOM row is the index — it has its own aging clock. The compactor ticks daily; only AFTER it runs do bytes actually leave disk and only AFTER it runs does the index forget. In steady state both shrink in lockstep — but they are NOT the same lifecycle.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Retention.expire
daily ILM pass — flips records to expired (still on disk)
# runs daily as part of the ILM delete action
def expire(now):
for index in cluster.indices:
# per-index policy; templates can override cluster default
window = index.template.retention or cluster.default_retention
for record in index.records:
age = now - record.ingest_day
if age > window:
record.retention_expired = True # promise dropped
# NOTE: bytes still on disk until Compactor.run()
Compactor.run
periodic — rewrite/segment-merge that actually frees bytes
# runs daily; this is when bytes physically LEAVE disk
def run():
for segment in index.segments:
if segment.size_gb > max_merged_segment: # 5 GB default
continue # SKIP — too big to auto-merge
kept = [r for r in segment.records
if not r.retention_expired
and not r.tombstoned]
new_segment = rewrite(kept) # bytes finally gone
replace(segment, new_segment)
compactor.last_run_day = today()
DeleteByQuery.execute
writes tombstones; bytes leave only on merge (or force_merge)
def execute(query):
# segments are immutable — no in-place removal
matches = index.search(query)
for doc in matches:
segment = doc.segment
segment.tombstones.add(doc.id) # marker, not eviction
# index now returns 0 hits for query (the promise)
# but bytes stay until segment.merge() rewrites it
# auto-merge SKIPS segments > max_merged_segment (5 GB)
return {deleted: len(matches), bytes_freed: 0}
def force_merge(index): # operator-invoked; IO-heavy, blocking
for segment in index.segments:
if segment.size_gb > max_merged_segment:
continue # default: still skips >5 GB
rewrite_without(segment, segment.tombstones)

Where this sits in Build a distributed logging stack (ELK / Loki)

Scene 08 of 12. Retention is when the system stops promising you can read; deletion is when bytes are physically gone — and the gap is where compliance bugs live.

Up next. Bytes die on a schedule and the index has to forget. But even with perfect retention, at scale we cannot keep everything we emit. The honest question is which lines we sacrifice — and the wrong answer is 'the only error of the day'.

All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum

Built with Arqly
Every scene in Build a distributed logging stack (ELK / Loki) builds on the one before it.All 12 Build a distributed logging stack (ELK / Loki) scenes