Hot, warm, cold, frozen — log storage tiers and ILM phases
Logs grow forever and nobody can afford 'all logs on NVMe forever' — the economics of two orders of magnitude between NVMe and object storage force log data through tiers. ELK has four ILM phases (HOT → WARM → COLD → FROZEN) that read from progressively cheaper storage; Loki collapses the layers because chunks live on S3 from day one. ILM only DEMOTES bytes here — when bytes DIE is scene 8.
High-cardinality data has to live in the body, and the body is BIG and most of it is OLD. NVMe is too expensive to hold months of body — economics force tiering.
Scene 07
Hot, warm, cold, frozen
- Watch
- Try it
- Predict
- Capture
Watch one index tile age. It's born today on HOT (NVMe, replicated, ms latency). The ILM clock advances 7 days — it slides to WARM (force-merged, fewer replicas). 30 days — it slides to COLD (a searchable snapshot in S3, fully mounted, ~50% disk savings, no replicas). 90 days — it slides to FROZEN (small NVMe cache, partially mounted from S3, up to 20× warm capacity). To the right, the Loki lane shows the same body parked on S3-Standard from day 0 — no movement, ever.
Highlighted lines are the ones running in the diagram right now.
# runs once per day on the master nodedef ilm_tick():for index in cluster.indices:age = now() - index.creation_datepolicy = index.ilm_policyif age >= policy.hot.max_age: # rolloverrollover(index)if age >= policy.warm.min_age:demote(index, phase='warm')if age >= policy.cold.min_age:demote(index, phase='cold')if age >= policy.frozen.min_age:demote(index, phase='frozen')
def demote(index, phase):if phase == 'warm':force_merge(index, max_num_segments=1)set_replicas(index, 0)allocation.require(data_warm)elif phase == 'cold':snap = searchable_snapshot(index, repo=s3)mount(snap, type='full_copy') # ~50% disk savingsdelete_local_index(index)elif phase == 'frozen':snap = searchable_snapshot(index, repo=s3)mount(snap, type='partial') # NVMe cache + S3
# chunks were written to S3 the moment they flushed —# the body never moves. Only the index is compacted.def compact_index():shards = list_index_shards(object_store)for day in shards.by_day():merged = merge_boltdb_shards(day) # → TSDBupload(merged, object_store)delete(day.original_shards)# chunks: untouched, still on S3-Standard from day 0
Where this sits in Build a distributed logging stack (ELK / Loki)
Scene 07 of 12. Two orders of magnitude in cost between NVMe and Deep Archive force tiering. ELK has four ILM phases; Loki collapses to S3 from day one.
Up next. We have a cost ladder from NVMe to Deep Archive — two orders of magnitude. Aging data through it is the policy engine; but the policy decides when bytes MOVE, not when bytes DIE.
All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum