Design your logging stack
Every distributed-logging configuration — agent + buffer + structuring + index strategy + labels + tiers + retention + sampling + backend topology — is a deliberate trade against the workload's volume, query SLO, and compliance posture, and a verifier can trace each choice back to the scene that proved it.
Every choice from scenes 02 through 10 is now a knob. Pick a workload and ship a configuration the verifier won't reject.
Scene 11
Design your logging stack
- Watch
- Try it
- Predict
- Capture
Workload A is on the canvas — 50 hosts, 10 GB/day, query latency forgiving, 90-day retention. The default stack loads — Promtail with disk spill, JSON at the app, Loki labels-only, low-cardinality labels, S3 from day 0, no sampling, RF=3. Watch the verifier walk each knob and cite the scene that defends it.
Highlighted lines are the ones running in the diagram right now.
def check_workload(workload, stack):issues = []# dlog-05a: query SLO picks the index strategy.if workload.required_query_p99_ms < 1000:if stack.index_strategy != 'elk-inverted':issues.append(reject('dlog-05a'))# dlog-06: high-card identifiers stay out of labels.for label in stack.labels:if label in UNBOUNDED_KEYS: # trace_id, user_id…issues.append(reject('dlog-06'))# dlog-08: long retention demands per-index template + PII audit.if workload.required_retention_y >= 1 and not stack.pii_audit:issues.append(warn('dlog-08'))# dlog-03: never block the app on the agent buffer.if stack.buffer_policy != 'spill':issues.append(warn('dlog-03'))return issues
def serialize_config(stack):return {'agent': stack.agent, # dlog-02'buffer_policy': stack.buffer_policy, # dlog-03'structuring': stack.structuring, # dlog-04'index_strategy': stack.index_strategy, # dlog-05'labels': stack.labels, # dlog-06'tier_policy': stack.tier_policy, # dlog-07'retention': stack.retention, # dlog-08'pii_audit': stack.pii_audit, # dlog-08'sampling': stack.sampling, # dlog-09'replication_factor': stack.rf, # dlog-10}
def profile(card):return Workload(volume_per_day=card.gb_per_day,required_query_p99_ms=(forgiving if card.id == 'startup-A'else 500 if card.id == 'fintech-B'else rare # iot-C: queries are rare),required_retention_y=(0.25 if card.id == 'startup-A'else 7 if card.id == 'fintech-B'else 0.08 # iot-C: ~30 days),pii_required=(card.id == 'fintech-B'), # GDPRmust_include_trace_id=(card.id == 'iot-C'),)
Where this sits in Build a distributed logging stack (ELK / Loki)
Scene 11 of 12. Capstone: agent + buffer policy + structuring + index strategy + tiers + retention + sampling — the verifier traces every choice back to the scene that earned it.
All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum