Never make a customer wait for telemetry

A finished span goes into a bounded in-memory queue that a background worker drains in batches, and when that queue is full the library drops spans and counts the drops — because blocking the request instead would turn a backend slowdown into an application outage.

Previously

We asked every span of a trace to travel to a shared decision point. The first leg of that journey is out of the application itself, and the application is serving a customer.

Scene 10

Never make a customer wait for telemetry

  1. Watch
  2. Try it
  3. Predict
  4. Capture
export: a bounded queue between the request and the networkinstrument → export → collect → sample → store → find → readinstrumentexportbatch 512 · 5 scollectsamplestorefindreadrequest path (hot)handlerdoes the workrecord spanappend to queuerespondcustomer servedhand off — no waitrequest latency200ms budget180ms of work+0.2ms exportbatched: the request returns before a byte leaves the boxbackground exporterburst ×1bounded queue → flush in batches0 / 8 slotsbackendbatches0spans dropped in a 30 s burst (example)0 dropped this burstwhat the viewer shows afterwardsdashed = never arrived · no error anywhere0750ms1.5s2.3s3.0sspanPOST /checkout3.0scheckout2.9sGET /cartcharge2.6sPOST /score2.4spost_entryTIME TO CULPRIT6 minexampleOne bounded enqueue on the request path; batches leave on a background worker. This queue is in app memory, not Shopfront's rose queue.
What to watch for

A span is finished — so how does it get out of the application and over to wherever traces are kept, without the customer who is still waiting for a reply paying for that trip? Watch the two lanes. On the left, the handler records the span and replies immediately. On the right, the finished spans pile into a fixed-size holding area that a background worker drains in batches. Then traffic bursts, the worker falls behind, and the holding area runs out of room. Watch what the human sees at the bottom afterwards.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

SimpleSpanProcessor.on_end
one network call per span, inside the request
def handle_request(req):
spans = do_work(req) # 4 spans per request
for span in spans:
processor.on_end(span)
return response # +4 x export_rtt
def on_end(span): # SimpleSpanProcessor
resp = exporter.export([span]) # one OTLP call
block_until(resp)
BatchSpanProcessor.on_end
the handler appends once and returns
MAX_QUEUE_SIZE = 2048 # spans, in app memory
def on_end(span): # BatchSpanProcessor
if len(queue) >= MAX_QUEUE_SIZE:
spans_dropped.add(
1, {"error.type": "queue_full"},
)
return
queue.append(span)
BatchSpanProcessor.worker
a background thread drains the queue in batches
def worker(): # background thread
while running:
wait(scheduled_delay_ms=5000)
while queue:
batch = queue.take(max_export_batch=512)
resp = exporter.export(batch) # OTLP
wait(resp, timeout_ms=30000)

Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)

Scene 10 of 17, in the Pipeline act — Off the hot path, then a tier that isn't your app.. A finished span goes into a bounded queue that a background worker drains in batches. When it fills, the library drops spans and counts them — blocking the request would turn a backend slowdown into an outage.

Up next. The application now hands batches to something else and forgets them. What is that something else, and what else should it be doing?

All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum

Built with Arqly
Every scene in Build a distributed tracing system (Jaeger / Zipkin style) builds on the one before it.All 17 Build a distributed tracing system (Jaeger / Zipkin style) scenes