Churn: yesterday's pods still cost you — active series versus series held in memory
Every rolling deploy gives pods new names, and because pod is a label each new name creates brand-new series while the old ones go quiet but stay in memory for hours and in every block index, so cost follows every series seen recently, not the count that is getting samples now.
Query cost follows the series a query touches; the dashboard's count of series that are getting samples can badly understate that number.
Scene 13
Churn: yesterday's pods still cost you
- Watch
- Try it
- Predict
- Capture
Why can a flat count of live series still blow up memory and queries? Checkout runs 10k series per generation of pods, and Shopfront deploys once an hour. Watch the three counters on the right as each deploy lands: the top one, the one every dashboard shows, and the two below it.
Highlighted lines are the ones running in the diagram right now.
def roll(deployment, every):while True:for pod in deployment.pods:kill(pod)name = f"{deployment}-{rand_suffix()}"pod = start(pod.spec, name = name)# pod and instance are labelspod.labels["pod"] = namepod.labels["instance"] = pod.ipsleep(every) # 144 rolls/day = every 10 min
def append(labels, value, t):ref = series_by_labels.get(hash(labels))if ref is None:ref = head.create_series(labels)head_series_created_total += 1ref.last_sample = tref.chunk.append(t, value)head_series = len(series_by_labels)active_series = count(last_sample > t - 20m)
def compact(): # ingesters check every 1 mif head.max_time - head.min_time < 3h:returncut = head.min_time + 2hblocks.write(head.samples_before(cut))head.truncate(before = cut)for ref in series_by_labels.values():if ref.last_sample < cut:series_by_labels.remove(ref)
def select(matchers, start, end):series = {}for block in blocks.overlapping(start, end):# 2 h blocks: 24 h = 12 of them, 7 d = 84for ref in block.index.matching(matchers):series[ref.labels] = block.chunks(ref)for ref in head.select(matchers, start, end):series[ref.labels] = ref.chunksreturn evaluate(matchers, series)
Where this sits in Metrics / Monitoring System
Scene 13 of 18, in the Cardinality act — Churn, defenses, aggregation, then design it.. Every rolling deploy gives pods new names and brand-new series, while the old ones go quiet but linger in memory and in every block index — cost follows series seen recently, not the active count.
Up next. Now that churn and bad labels create series faster than anyone expects, the next question is where an explosion can be stopped, and what each stopping point costs.
All 18 scenes in Metrics / Monitoring System · Every curriculum