Churn: yesterday's pods still cost you — active series versus series held in memory

Every rolling deploy gives pods new names, and because pod is a label each new name creates brand-new series while the old ones go quiet but stay in memory for hours and in every block index, so cost follows every series seen recently, not the count that is getting samples now.

Previously

Query cost follows the series a query touches; the dashboard's count of series that are getting samples can badly understate that number.

Scene 13

Churn: yesterday's pods still cost you

  1. Watch
  2. Try it
  3. Predict
  4. Capture
series churn: pods come and go, series pile upcheckout pod generations over time (newest at the bottom)gen 1 · 10k seriesactivequery range: 1 h≈ 10k series touchedactivequiet · in memoryflushed · blocks onlyquiet series leave memory only at the next ~2 h cuteach deploy: 10k brand-new seriesseries getting samples right now10kseries in memory10kseries seen today10kOne generation of checkout pods: 10k series, all of them getting samples. Nothing has been replaced yet.
What to watch for

Why can a flat count of live series still blow up memory and queries? Checkout runs 10k series per generation of pods, and Shopfront deploys once an hour. Watch the three counters on the right as each deploy lands: the top one, the one every dashboard shows, and the two below it.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Deploy.roll
every replacement gives a pod a brand-new name
def roll(deployment, every):
while True:
for pod in deployment.pods:
kill(pod)
name = f"{deployment}-{rand_suffix()}"
pod = start(pod.spec, name = name)
# pod and instance are labels
pod.labels["pod"] = name
pod.labels["instance"] = pod.ip
sleep(every) # 144 rolls/day = every 10 min
Head.append
a label set never seen before becomes a new series
def append(labels, value, t):
ref = series_by_labels.get(hash(labels))
if ref is None:
ref = head.create_series(labels)
head_series_created_total += 1
ref.last_sample = t
ref.chunk.append(t, value)
head_series = len(series_by_labels)
active_series = count(last_sample > t - 20m)
Head.compact
the only moment a quiet series leaves memory
def compact(): # ingesters check every 1 m
if head.max_time - head.min_time < 3h:
return
cut = head.min_time + 2h
blocks.write(head.samples_before(cut))
head.truncate(before = cut)
for ref in series_by_labels.values():
if ref.last_sample < cut:
series_by_labels.remove(ref)
Query.select
a range walks every block index it overlaps
def select(matchers, start, end):
series = {}
for block in blocks.overlapping(start, end):
# 2 h blocks: 24 h = 12 of them, 7 d = 84
for ref in block.index.matching(matchers):
series[ref.labels] = block.chunks(ref)
for ref in head.select(matchers, start, end):
series[ref.labels] = ref.chunks
return evaluate(matchers, series)

Where this sits in Metrics / Monitoring System

Scene 13 of 18, in the Cardinality act — Churn, defenses, aggregation, then design it.. Every rolling deploy gives pods new names and brand-new series, while the old ones go quiet but linger in memory and in every block index — cost follows series seen recently, not the active count.

Up next. Now that churn and bad labels create series faster than anyone expects, the next question is where an explosion can be stopped, and what each stopping point costs.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes