Cost grows with labels, not traffic — label cardinality and the series count

The app reports one number per interval for each distinct combination of labels, so a metric's cost ignores traffic but multiplies with every label, and one label with unlimited values like customer_id costs more than a 100× traffic spike.

Previously

A rule needs numbers to check; before collecting them everywhere, we need to know what one of those numbers costs to keep.

Scene 02

Cost grows with labels, not traffic

  1. Watch
  2. Try it
  3. Predict
  4. Capture
what a metric costs: labels, not trafficcheckout traffic10 req/s (example)log byteslog bytes ×1series count1 serieslabels on checkout's request metricservice1 value= 1×route10 valuesnot added×status5 valuesnot added×pod50 valuesnot added×customer_id2M valuesnot addedrunning total = series, one per distinct combination of label valuesseries for this one metric1traffic: 10 req/s (example)well under the example budgetbudgetmonitor memory≈ 3 kB (estimate)a few KiB per series (measured)TIME TO DETECT≈ 30 sexampleEvery request adds 1 to a number the app already holds. The logs grow with every request; the number of series doesn't.
What to watch for

Why do metrics stay cheap as traffic grows, and what makes them expensive? Watch checkout's traffic climb from 10 to 10,000 requests per second: the log bytes climb with it, and the series count for checkout's request metric (a series is short for time series, one number reported every interval) stays put. Then the on-call engineer needs to know which route is failing, so the number gets tagged, and the tree in the middle starts to multiply.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

App.handle_request
every request writes a log line and bumps a count
def handle_request(req):
resp = checkout(req)
log.write(req, resp)
labels = {
"service": "checkout",
"route": req.route,
"status": resp.status,
"pod": POD_NAME,
"customer_id": req.customer_id,
}
request_count[labels] += 1
App.report
one value per label combination, per interval
every interval: # e.g. every 15 s
for labels, value in request_count.items():
send(labels, value, now())
Monitor.store
memory is held per series, not per request
def store(labels, value, ts):
series = active.get(labels)
if series is None:
series = new_series(labels) # ~3 kB (measured)
active[labels] = series
series.append(ts, value)
def check_memory():
used = len(active) * SERIES_BYTES
if used > memory_budget:
crash("out of memory")

Where this sits in Metrics / Monitoring System

Scene 02 of 18, in the Why measure act — Numbers beat complaints; labels, not traffic, set the bill.. The app reports one number per label combination per interval, so a metric's cost ignores traffic and multiplies with every label: one unbounded label like customer_id dwarfs a 100× traffic spike.

Up next. Now that each series is one number the app reports every interval, the next question is what that number should mean, a total so far or a count since the last read, so a missed or repeated read doesn't corrupt it.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes