One row per hop, and the id that links them — parent span ids and read-time assembly

Each hop writes one flat record carrying its own id, its parent's id and a shared trace id, and nothing is nested when it is written — the tree is rebuilt later by matching parent ids.

Previously

The recorded request became one picture because the records were tied together. This scene opens one record and shows exactly what does the tying.

Scene 02

One row per hop, and the id that links them

  1. Watch
  2. Try it
  3. Predict
  4. Capture
one span, and five of them becoming a treeTHE SPAN, FIELD BY FIELDcheckout · one spannameplace_orderstart_time12:04:07.112duration2.90 sspan_id00f067aa0ba902b7this record's own idparent_span_idb7ad6b7169203331the id of the record whose work caused ittrace_id4bf92f35…0e0e4736identical on every record of this one requestkindserverwhich side of the call wrote this rowstatusunsetsuccess or failure, without guessingSPANS AS THEY ARRIVED — NO ORDER, NO TREE0750ms1.5s2.3s3.0s3.00 s (example)Five machines, five separate writes. Nothing is nested when it is written.
a name and two times. Useful — and pointing at nothing else.
What to watch for

What actually turns five records, written on five different machines, into one picture? Watch one record fill in field by field: first the name and the times, then three ids, then two more fields. Then five records land on the rail in whatever order they finished — five flat rows with nothing nested inside anything. Watch what happens when the ids are used to link them.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Tracer.start_span
the fields that exist the moment a hop begins
def start_span(name, ctx, kind = INTERNAL):
span = Span()
span.name = name # the only required argument
span.start_time = now()
parent = ctx.current_span
span.trace_id = parent.trace_id if parent else new_id()
span.span_id = new_span_id()
span.parent_span_id = parent.span_id if parent else ""
span.kind = kind # SERVER, CLIENT, PRODUCER, …
span.status = UNSET # Ok is asserted, not inferred
return span
Span.end
one flat record, written and shipped on its own
def end(span):
span.end_time = now()
span.duration = span.end_time - span.start_time
# nothing is nested here: no child is attached,
# no parent is notified, no sibling is waited for
exporter.enqueue(span)
def on_batch(batch): # BatchSpanProcessor
for span in batch:
collector.send(span) # each record travels alone
Viewer.build_tree
the tree is assembled here, when someone reads it
def build_tree(trace_id):
rows = store.scan(trace_id = trace_id)
children = defaultdict(list)
root = None
for row in rows:
if row.parent_span_id == "": # empty => root
root = row
else:
children[row.parent_span_id].append(row)
# no timestamp, host or arrival index is ever read here
# a parent id naming a row that is absent matches nothing
return render(root, children)

Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)

Scene 02 of 17, in the Why trace act — Per-service records can't blame a hop; one id can.. Each hop writes one flat record carrying its own id, its parent's id and a shared trace id. Nothing is nested when it is written — the tree is rebuilt later, which is why a missing record leaves a plausible-looking hole.

Up next. The tree only assembles because every record already carried the same trace id and a pointer to its parent. So how does the next service learn that id in the first place?

All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum

Built with Arqly
Every scene in Build a distributed tracing system (Jaeger / Zipkin style) builds on the one before it.All 17 Build a distributed tracing system (Jaeger / Zipkin style) scenes