Design your monitoring platform
A monitoring platform is a chain of decisions about collection, alerting, scaling out, retention and label budget, each right only for a workload whose series count, team count, retention and detection target justify it.
You can now control what every series costs without losing answers; the last step is choosing which of these pieces a given workload actually needs.
Scene 16
Design your monitoring platform
- Watch
- Try it
- Predict
- Capture
Can you put every piece together for a real workload and defend each choice? Here is the workload: 30 teams, 50k pods, 13 months of history, and a page within 2 minutes of checkout breaking. The card's blue numbers are estimates derived from Shopfront's label math — roughly 300 series per pod, so ≈ 15M active series, and at 3 copies in memory that is ≈ 375 GB of ingester RAM. Grafana's own advice is to measure your own cluster rather than trust an estimate. The pipeline strip below is empty: nothing has been placed yet. Watch the checker grade that empty design.
Highlighted lines are the ones running in the diagram right now.
def size_the_fleet(workload, design):# estimate from Shopfront's label mathactive = workload.pods * 300if design.customer_id_label:active *= workload.customers # 2M valuesif not design.cluster:return active <= one_server_ceiling # ~10Min_memory = active * copies # copies = 3ram_gb = in_memory / 300_000 * 2.5 # estimatereturn active, ram_gb
def time_to_detect(workload, design):if not design.rules:return "never" # nothing evaluatesttd = scrape_interval # 15 sttd += evaluation_interval # 15 sttd += design.for_duration # 1 min, or 2 minttd += group_wait # 30 s bundlingfor f in design.failures:if absorbed(f): # twin alerts, 2 of 3 confirmcontinue # the clock does not movereturn "never" # the monitor itself is gonereturn ttd <= workload.target
RULES = [ # each row cites a scene("fits memory", fits_memory, "metrics-09"),("series budget", under_budget, "metrics-02"),("13 months kept", has_compactor, "metrics-11"),("which customer?", has_exemplars, "metrics-15"),("complexity", is_justified, "metrics-09")]def verify(workload, design):rows = []for check, rule, scene in RULES:if rule.applies(workload):rows += row(check, rule(workload, design), scene)return rows
Where this sits in Metrics / Monitoring System
Scene 16 of 18, in the Cardinality act — Churn, defenses, aggregation, then design it.. A monitoring platform is a chain of decisions — collection, alerting, scaling out, retention and label budget — each right only for a workload whose series count, teams, retention and detection target justify it.
All 18 scenes in Metrics / Monitoring System · Every curriculum