Keep a year cheaply: blocks on object storage

Ingesters hold only the recent hours in memory and every two hours seal that window into an immutable block on object storage, where a compactor merges the three copies replication created and joins small blocks into day-sized ones — which turns long retention from a memory problem into a storage-cost and query-speed problem.

Previously

Now that exactly one stream per series lands in memory, the next question is how the cluster keeps months of it without holding it all in memory.

Scene 11

Keep a year cheaply: blocks on object storage

  1. Watch
  2. Try it
  3. Predict
  4. Capture
instrumentcollectstorehours in memory · years…queryalertnotifyingester-1recent ~2 h in memoryrecent ~2 h in memoryidleingester-2recent ~2 h in memoryrecent ~2 h in memoryidleingester-3recent ~2 h in memoryrecent ~2 h in memoryidlelocal copy kept 13 h, then it lives only in object storagelocal copy kept 13 h, then it lives only in object storageCOMPACTOROFFoff — nothing merges; every upload stays exactly as itlandedOBJECT STORAGEno blocks uploaded yetRETENTION15 daysmemory + local diskobject storagememory: pricey per GB · object storage: cheap per GBstored 3× — one copy per ingester, nothing merges them30-DAY QUERYopens 1,080 blocks — 360 windows × 3opens 1,080 blocks — 360 windows × 3DOWNSAMPLINGOFFoff — Mimir does not downsample; it recommends recording rules insteadThree ingesters, each holding Shopfront's most recent two hours of samples in memory. Nothing has reached object storage yet.
What to watch for

How does a metrics cluster keep months of data without holding it in memory? Watch one tenant's two hours travel. An ingester keeps only recent samples in memory, so every two hours it seals that window into a file and uploads it to object storage — the cheap, durable, write-once store build-s3 covered (s3-00). Then watch what the other two ingesters do with the same two hours.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Ingester.cut_block
seal the last two hours, upload it, keep it locally a while
def cut_block(tenant):
block = seal(tenant.head) # never edited again
upload(bucket, block)
keep_local(block, for_="13h")
tenant.head = fresh_head()
every("2h"):
for ingester in owners_of(tenant.series):
ingester.cut_block(tenant)
Compactor.run
the background job that rewrites what is already stored
def run():
if not compactor_enabled:
return # nothing is rewritten
for window in bucket.windows():
copies = bucket.blocks(window)
if len(copies) > 1:
merge_vertical(copies)
for group in adjacent(bucket):
merge_horizontal(group) # 2h -> 12h -> 24h
for block in bucket.blocks():
if block.age > blocks_retention_period:
delete(block) # unset = kept forever
Compactor.downsample
coarser copies of old blocks, kept beside the raw one
def downsample(block):
if not downsampling_enabled:
return # Mimir removed this path
if block.age > "40h":
coarse = aggregate(block, every="5m")
upload(bucket, coarse) # added, not swapped
if block.age > "10d":
coarse = aggregate(block, every="1h")
upload(bucket, coarse) # added, not swapped

Where this sits in Metrics / Monitoring System

Scene 11 of 18, in the Scale out act — Remote-write, ingesters, dedup, blocks, split queries.. Ingesters keep only recent hours in memory and cut immutable two-hour blocks to object storage, where a compactor merges the copies replication created — retention becomes a storage and query-speed problem.

Up next. Now that years of data sit in compacted blocks, the next question is how a 30-day query over them comes back fast.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes