What a million traces say together — the service dependency graph and span metrics

The parent-child edges of many traces draw the live dependency graph and counting spans by service and operation yields rates and error ratios for hops no dashboard covered — but both are only as honest as the sampling that fed them.

Previously

Reading one trace is a skill, and one trace can mislead. The same spans in bulk draw something nobody drew by hand.

Scene 15

What a million traces say together

  1. Watch
  2. Try it
  3. Predict
  4. Capture
ten thousand traces, folded into one pictureevery span passes the collector, so every call the spans saw becomes an edgederived from every span at the collector1% keptcontext stops at the queueSERVICE MAP · EDGE WIDTH = CALLS/S???????gatewaycheckoutcartinventorypaymentsorders-workerfraud-checkledger? queue hopcheckout → orders-worker was never recorded0 edges thinned below the sample floor? = a call no span recordedDERIVED FROM THESE SPANSTRUTHedges on the map0 edges7 edgescheckout error rate0.4 %0.4 %PER-OPERATION, FROM THE SAME SPANSrate · errors · p99, counted from the same spansOPERATIONRATEERRP99GET /checkoutgateway1.2k/s0.4 %3.0 sPOST /checkoutcheckout1.2k/s0.4 %3.0 sGET /cartcart1.2k/s0.1 %24 msGET /stockinventory1.1k/s0.2 %38 msPOST /chargepayments900/s2.4 %410 msPOST /scorefraud-check900/s0.3 %2.4 sPOST /entryledger0.60/s0.4 %90 msprocess orderorders-worker900/s0.6 %150 msderived at the collector, before the sampler — the counts match the traffic
What to watch for

What can a million traces tell you that one trace cannot? Two things, and neither of them needs a new agent. First: whenever one service's span is the parent of another service's span, that is a call across a boundary — do that match over ten thousand traces and you have drawn a map of who calls whom that nobody maintained by hand. Second: count those same spans by service, operation, kind and status and you get a request rate, an error ratio and a p99 for every hop, including hops no dashboard ever covered. You already know what a rate and a latency histogram are from the metrics curriculum; the question here is only where these ones came from. Watch the traces fold in, and keep an eye on what shows up late — and on what never shows up at all.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

ServiceGraph.on_span
matches a client span to its server span, emits an edge
def on_span(span, store):
if span.kind == CLIENT:
key = (span.trace_id, span.span_id)
elif span.kind == SERVER and span.parent_span_id:
key = (span.trace_id, span.parent_span_id)
else:
return # no parent id, no key
peer = store.pop(key) # keys expire at max_age
if peer is None:
return store.put(key, span)
client, server = order(span, peer)
edges[(client.service, server.service)].calls += 1
SpanMetrics.on_span
counts spans by service, operation, kind and status
def on_span(span):
key = (
span.resource['service.name'],
span.name, # an id here = one key per request
span.kind,
span.status.code,
)
if is_new(key) and len(calls) >= cardinality_limit:
return # aggregation_cardinality_limit: 0
calls[key] += 1
duration[key].observe(span.end - span.start)
Collector.pipeline
both connectors run before anything is dropped
def pipeline(span):
service_graph.on_span(span)
span_metrics.on_span(span)
if tail_policy.decide(span.trace_id) == DROP:
return
storage.write(span)

Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)

Scene 15 of 17, in the Ship it act — What a million traces say, then design the stack.. The parent-child edges of many traces draw the live dependency graph, and counting spans by service and operation yields rates and error ratios for hops no dashboard covered — both only as honest as the sampling.

Up next. Maps and span metrics are only as honest as the sampling that fed them, and every mechanism in this curriculum now has a price. Can you pick a set for a real workload and defend each choice?

All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum

Built with Arqly
Every scene in Build a distributed tracing system (Jaeger / Zipkin style) builds on the one before it.All 17 Build a distributed tracing system (Jaeger / Zipkin style) scenes