What a million traces say together — the service dependency graph and span metrics
The parent-child edges of many traces draw the live dependency graph and counting spans by service and operation yields rates and error ratios for hops no dashboard covered — but both are only as honest as the sampling that fed them.
Reading one trace is a skill, and one trace can mislead. The same spans in bulk draw something nobody drew by hand.
Scene 15
What a million traces say together
- Watch
- Try it
- Predict
- Capture
What can a million traces tell you that one trace cannot? Two things, and neither of them needs a new agent. First: whenever one service's span is the parent of another service's span, that is a call across a boundary — do that match over ten thousand traces and you have drawn a map of who calls whom that nobody maintained by hand. Second: count those same spans by service, operation, kind and status and you get a request rate, an error ratio and a p99 for every hop, including hops no dashboard ever covered. You already know what a rate and a latency histogram are from the metrics curriculum; the question here is only where these ones came from. Watch the traces fold in, and keep an eye on what shows up late — and on what never shows up at all.
Highlighted lines are the ones running in the diagram right now.
def on_span(span, store):if span.kind == CLIENT:key = (span.trace_id, span.span_id)elif span.kind == SERVER and span.parent_span_id:key = (span.trace_id, span.parent_span_id)else:return # no parent id, no keypeer = store.pop(key) # keys expire at max_ageif peer is None:return store.put(key, span)client, server = order(span, peer)edges[(client.service, server.service)].calls += 1
def on_span(span):key = (span.resource['service.name'],span.name, # an id here = one key per requestspan.kind,span.status.code,)if is_new(key) and len(calls) >= cardinality_limit:return # aggregation_cardinality_limit: 0calls[key] += 1duration[key].observe(span.end - span.start)
def pipeline(span):service_graph.on_span(span)span_metrics.on_span(span)if tail_policy.decide(span.trace_id) == DROP:returnstorage.write(span)
Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)
Scene 15 of 17, in the Ship it act — What a million traces say, then design the stack.. The parent-child edges of many traces draw the live dependency graph, and counting spans by service and operation yields rates and error ratios for hops no dashboard covered — both only as honest as the sampling.
Up next. Maps and span metrics are only as honest as the sampling that fed them, and every mechanism in this curriculum now has a price. Can you pick a set for a real workload and defend each choice?
All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum