Never make a customer wait for telemetry
A finished span goes into a bounded in-memory queue that a background worker drains in batches, and when that queue is full the library drops spans and counts the drops — because blocking the request instead would turn a backend slowdown into an application outage.
We asked every span of a trace to travel to a shared decision point. The first leg of that journey is out of the application itself, and the application is serving a customer.
Scene 10
Never make a customer wait for telemetry
- Watch
- Try it
- Predict
- Capture
A span is finished — so how does it get out of the application and over to wherever traces are kept, without the customer who is still waiting for a reply paying for that trip? Watch the two lanes. On the left, the handler records the span and replies immediately. On the right, the finished spans pile into a fixed-size holding area that a background worker drains in batches. Then traffic bursts, the worker falls behind, and the holding area runs out of room. Watch what the human sees at the bottom afterwards.
Highlighted lines are the ones running in the diagram right now.
def handle_request(req):spans = do_work(req) # 4 spans per requestfor span in spans:processor.on_end(span)return response # +4 x export_rttdef on_end(span): # SimpleSpanProcessorresp = exporter.export([span]) # one OTLP callblock_until(resp)
MAX_QUEUE_SIZE = 2048 # spans, in app memorydef on_end(span): # BatchSpanProcessorif len(queue) >= MAX_QUEUE_SIZE:spans_dropped.add(1, {"error.type": "queue_full"},)returnqueue.append(span)
def worker(): # background threadwhile running:wait(scheduled_delay_ms=5000)while queue:batch = queue.take(max_export_batch=512)resp = exporter.export(batch) # OTLPwait(resp, timeout_ms=30000)
Where this sits in Build a distributed tracing system (Jaeger / Zipkin style)
Scene 10 of 17, in the Pipeline act — Off the hot path, then a tier that isn't your app.. A finished span goes into a bounded queue that a background worker drains in batches. When it fills, the library drops spans and counts them — blocking the request would turn a backend slowdown into an outage.
Up next. The application now hands batches to something else and forgets them. What is that something else, and what else should it be doing?
All 17 scenes in Build a distributed tracing system (Jaeger / Zipkin style) · Every curriculum