One monitor hits a wall — splitting targets, federation and remote-write
One server's memory grows with its series, so a growing fleet has three exits, each giving something up: split targets so fleet-wide questions must ask every server, federate pre-summed series and lose per-pod detail, or remote-write every sample to a cluster you must now run.
Now that an outside heartbeat tells us when the monitor dies, here is the most common reason it dies at scale: its memory must hold every series, and the fleet keeps growing.
Scene 09
One monitor hits a wall
- Watch
- Try it
- Predict
- Capture
What are your options when one monitoring server can't hold every series? Watch the left side. Shopfront's fleet grows from 1k pods to 50k, and the monitor's memory bar climbs with it, because a monitor keeps every series it is currently collecting in memory — build-tsdb showed why (tsdb-09). At the top of the slider the bar breaks through the budget line of the biggest machine you can rent. This is the wall build-tsdb handed over as somebody else's problem (tsdb-10a). Nothing is fixed yet; just watch the bar.
Highlighted lines are the ones running in the diagram right now.
def assign_targets(all_pods, shard_id):n_shards = ceil(series_of(all_pods) / ram_budget)mine = []for pod in all_pods: # from service discoveryif hash(pod.labels) % n_shards == shard_id:mine.append(pod) # hashmod relabelingreturn minedef fleet_wide(query):replies = [srv.ask(query) for srv in servers]return merge(replies) # one round trip each
def federate():for leaf in regional_servers:resp = leaf.get("/federate", match=RULE_SUMS)for series in resp:store(series.labels, series.value)sleep(scrape_interval)def answer(query):return evaluate(query, over=store)
def remote_write():while True:batch = wal.read(max_samples_per_send)shard = queue.pick(max_shards) # ~25% more RAMif shard.send(batch, to=receiver):wal.advance(len(batch))else:shard.retry(batch)if wal.oldest() > 2h:wal.truncate() # unsent data lost after ~2 h
Where this sits in Metrics / Monitoring System
Scene 09 of 18, in the Scale out act — Remote-write, ingesters, dedup, blocks, split queries.. One server's memory holds every active series, so a growing fleet has three exits: split targets and query them all, federate sums and lose per-pod detail, or remote-write to a cluster you must now run.
Up next. Now that remote-write sends every sample to a cluster, the next question is what that cluster does with a sample so no single machine is overloaded — or loses it.
All 18 scenes in Metrics / Monitoring System · Every curriculum