Design your monitoring platform

A monitoring platform is a chain of decisions about collection, alerting, scaling out, retention and label budget, each right only for a workload whose series count, team count, retention and detection target justify it.

Previously

You can now control what every series costs without losing answers; the last step is choosing which of these pieces a given workload actually needs.

Scene 16

Design your monitoring platform

  1. Watch
  2. Try it
  3. Predict
  4. Capture
StartupPlatformPer-customerTIME TO DETECTnevertarget ≤ 2 minPlatformpreset · nothing placed yetTEAMS30PODS50kRETENTION13 monthsDETECT≤ 2 minactive series≈ 15M (estimate)ingester RAM≈ 375 GB (estimate)instrumentcollectstorequeryalertnotifycustomer_idexemplars onbucketsredundant pairdoors: relabel +limitssum away podthe pair keeps itremote-write ×3dedup: one streamcompactorfrontend + cacherule: wait 2 minrule: wait 1 minrouter +heartbeatINJECT FAILURESoffsteady state — no failures injectedCHECKERnothing to check yetchecking…Platform's numbers are on the card, derived from Shopfront's label math. The design is empty, so nothing would ever page.
What to watch for

Can you put every piece together for a real workload and defend each choice? Here is the workload: 30 teams, 50k pods, 13 months of history, and a page within 2 minutes of checkout breaking. The card's blue numbers are estimates derived from Shopfront's label math — roughly 300 series per pod, so ≈ 15M active series, and at 3 copies in memory that is ≈ 375 GB of ingester RAM. Grafana's own advice is to measure your own cluster rather than trust an estimate. The pipeline strip below is empty: nothing has been placed yet. Watch the checker grade that empty design.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Checker.size_the_fleet
every number it returns is an estimate, not a measurement
def size_the_fleet(workload, design):
# estimate from Shopfront's label math
active = workload.pods * 300
if design.customer_id_label:
active *= workload.customers # 2M values
if not design.cluster:
return active <= one_server_ceiling # ~10M
in_memory = active * copies # copies = 3
ram_gb = in_memory / 300_000 * 2.5 # estimate
return active, ram_gb
Checker.time_to_detect
the clock is a sum of four waits, checked against the target
def time_to_detect(workload, design):
if not design.rules:
return "never" # nothing evaluates
ttd = scrape_interval # 15 s
ttd += evaluation_interval # 15 s
ttd += design.for_duration # 1 min, or 2 min
ttd += group_wait # 30 s bundling
for f in design.failures:
if absorbed(f): # twin alerts, 2 of 3 confirm
continue # the clock does not move
return "never" # the monitor itself is gone
return ttd <= workload.target
Checker.verify
one row per rule this workload asks for, each naming its scene
RULES = [ # each row cites a scene
("fits memory", fits_memory, "metrics-09"),
("series budget", under_budget, "metrics-02"),
("13 months kept", has_compactor, "metrics-11"),
("which customer?", has_exemplars, "metrics-15"),
("complexity", is_justified, "metrics-09")]
def verify(workload, design):
rows = []
for check, rule, scene in RULES:
if rule.applies(workload):
rows += row(check, rule(workload, design), scene)
return rows

Where this sits in Metrics / Monitoring System

Scene 16 of 18, in the Cardinality act — Churn, defenses, aggregation, then design it.. A monitoring platform is a chain of decisions — collection, alerting, scaling out, retention and label budget — each right only for a workload whose series count, teams, retention and detection target justify it.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes