Keep the answer, drop the series — pre-aggregation on arrival and bucket exemplars

Summing away the labels nobody queries as samples arrive cuts stored series by orders of magnitude, while deleting every distinguishing label makes counters collide and rates lie. One saved example request keeps the path to a customer without a series per customer.

Previously

Every defense stops the explosion by dropping, blinding or refusing something; we want the answers without the series.

Scene 15

Keep the answer, drop the series

  1. Watch
  2. Try it
  3. Predict
  4. Capture
keep what dashboards ask forincoming raw labelsroutekeptstatuskeptlabel strategy▸ keep alldelete at scrapersum on arrivalevery label stored as scrapedseries stored vs queried (log scale)stored50 series stored (route × status)queried50 series queried by the dashboardscheckout request ratecheckout latency buckets (seconds)≤ 0.1≤ 0.5≤ 1≤ 5+InfREADOUTwhich customer saw the 5 s spike?no customer_id label yetTIME TO DETECT≈ 3 minexampleCheckout's dashboards only ever ask by route and status: 50 series, and every panel is drawn from those.
every dashboard panel is drawn from these 50 series
What to watch for

How do you keep people's questions answerable while storing far fewer series? Start by looking at the gap. Checkout's dashboards only ever query by route and status — 50 series in total. Watch the labels the scrape actually carries pile up beside them, and watch the stored bar leave the queried bar behind.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Scraper.drop_labels
deleting a label deletes what tells two writers apart
def drop_labels(samples, drop): # labeldrop
if not drop:
return samples # stored as scraped
for s in samples:
for name in drop: # ["pod", "instance"]
s.labels.pop(name)
s.key = series_key(s.labels) # route, status
return samples # 50 pods -> one key
Ingester.append
every sample for one key is written on one timeline
def append(key, t, value):
last = series[key].last_sample
if last and t == last.t:
return reject("duplicate sample")
if last and t < last.t:
return reject("out of order")
series[key].push(t, value)
Aggregator.on_sample
sums the label away as samples arrive, before storage
DROP = ["pod", "instance"] # stream aggregation
def on_sample(s):
key = series_key(without(s.labels, DROP))
src = (key, s.labels["pod"]) # one state per pod
delta = max(0, s.value - last.get(src, s.value))
last[src] = s.value
total[key] += delta
def flush(): # every interval
for key, v in total.items():
remote_write(key, v) # only the rollup
App.observe
pins one example request beside the bucket it landed in
def observe(seconds, trace_id):
for b in buckets: # 0.1, 0.5, 1, 5, +Inf
if seconds <= b.le:
b.count += 1
b.exemplar = Exemplar(trace_id, seconds)
break
def expose(b): # OpenMetrics text
line = f"{b.count}"
if b.exemplar: # one per bucket
line += f' # {{trace_id="{b.exemplar.id}"}}'
return line # labels <= 128 chars

Where this sits in Metrics / Monitoring System

Scene 15 of 18, in the Cardinality act — Churn, defenses, aggregation, then design it.. Summing away labels nobody queries as samples arrive cuts series by orders of magnitude, while deleting every distinguishing label makes counters collide — and an exemplar keeps one example customer.

Up next. Now that you can control what every series costs without losing the answers, the next question is which combination of all these choices a real workload actually needs.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes