One monitor hits a wall — splitting targets, federation and remote-write

One server's memory grows with its series, so a growing fleet has three exits, each giving something up: split targets so fleet-wide questions must ask every server, federate pre-summed series and lose per-pod detail, or remote-write every sample to a cluster you must now run.

Previously

Now that an outside heartbeat tells us when the monitor dies, here is the most common reason it dies at scale: its memory must hold every series, and the fleet keeps growing.

Scene 09

One monitor hits a wall

  1. Watch
  2. Try it
  3. Predict
  4. Capture
instrumentcollectstorequeryalertnotifyTIME TO DETECT≈ 3 minexample1k podsone monitorbudget30%every series, in memoryway outfan-out ↓QUESTIONwhich pod, in any region, is returning5xx?1k pods. One monitor holds every series it is currently collecting, and sits well under the budget of the machine it runs on.
the bar is series held in memory — it tracks how many pods exist, not how busy they are
What to watch for

What are your options when one monitoring server can't hold every series? Watch the left side. Shopfront's fleet grows from 1k pods to 50k, and the monitor's memory bar climbs with it, because a monitor keeps every series it is currently collecting in memory — build-tsdb showed why (tsdb-09). At the top of the slider the bar breaks through the budget line of the biggest machine you can rent. This is the wall build-tsdb handed over as somebody else's problem (tsdb-10a). Nothing is fixed yet; just watch the bar.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Shard.assign_targets
each server scrapes one bucket; fleet questions ask them all
def assign_targets(all_pods, shard_id):
n_shards = ceil(series_of(all_pods) / ram_budget)
mine = []
for pod in all_pods: # from service discovery
if hash(pod.labels) % n_shards == shard_id:
mine.append(pod) # hashmod relabeling
return mine
def fleet_wide(query):
replies = [srv.ask(query) for srv in servers]
return merge(replies) # one round trip each
GlobalServer.federate
pulls /federate on an interval and stores what comes back
def federate():
for leaf in regional_servers:
resp = leaf.get("/federate", match=RULE_SUMS)
for series in resp:
store(series.labels, series.value)
sleep(scrape_interval)
def answer(query):
return evaluate(query, over=store)
Sender.remote_write
streams the write-ahead log out, batch by batch
def remote_write():
while True:
batch = wal.read(max_samples_per_send)
shard = queue.pick(max_shards) # ~25% more RAM
if shard.send(batch, to=receiver):
wal.advance(len(batch))
else:
shard.retry(batch)
if wal.oldest() > 2h:
wal.truncate() # unsent data lost after ~2 h

Where this sits in Metrics / Monitoring System

Scene 09 of 18, in the Scale out act — Remote-write, ingesters, dedup, blocks, split queries.. One server's memory holds every active series, so a growing fleet has three exits: split targets and query them all, federate sums and lose per-pod detail, or remote-write to a cluster you must now run.

Up next. Now that remote-write sends every sample to a cluster, the next question is what that cluster does with a sample so no single machine is overloaded — or loses it.

All 18 scenes in Metrics / Monitoring System · Every curriculum

Built with Arqly
Every scene in Metrics / Monitoring System builds on the one before it.All 18 Metrics / Monitoring System scenes