grep + ssh stops working — centralized log collection across a host fleet

Once your fleet has more than a handful of hosts, the only way to answer 'which host saw this error' is to ship the log lines somewhere central — local files plus ssh is an O(N) dead end.

Scene 01

grep + ssh stops working

  1. Watch
  2. Try it
  3. Predict
  4. Capture
DEV LAPTOP$ for h in $hosts; do ssh $h \\ grep 'ERROR' /var/log/app.logdone# find ERRORs in last hourRESPONSEssh fan-out: ~0.4shosts queried: 3FLEET · 3 hostshost-001last write 12:04:21/var/log/app.logapp.log.1app.log.2.gzhost-002last write 12:04:19/var/log/app.logapp.log.1app.log.2.gzhost-003last write 12:04:22/var/log/app.logapp.log.1app.log.2.gzTIME TO ANSWER0ms10s60s3 hosts. ssh + grep answers in seconds — slow, but it works.
What to watch for

Three hosts. You suspect a connection error so you ssh to each one in turn and grep for 'connection refused'. The time-to-answer meter at the bottom is the wall clock you're spending. Notice the rotated files stacked under each app.log — yesterday's data is already gzipped.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Engineer.find_error
the on-call's serial ssh-grep walk across the fleet
# you are the loop body
for host in hosts: # O(N) in fleet size
out = ssh(host,
'grep -h ERROR /var/log/app.log'
' /var/log/app.log.1'
' 2>/dev/null')
if ssh.exit_code != 0:
print(host, 'unreachable') # debugging the debugger
continue
for line in out.splitlines():
print(host, line) # may be empty: file rotated
Logrotate.rotate
the cron-driven cycle that evicts yesterday's file
# /etc/logrotate.d/app, run nightly by cron
def rotate(path='/var/log/app.log', keep=3):
# shift the chain: app.log.2.gz -> app.log.3.gz, etc.
for i in range(keep, 0, -1):
mv(f'{path}.{i}.gz', f'{path}.{i+1}.gz')
# yesterday's file becomes app.log.1, then gets gzipped
mv(path, f'{path}.1')
gzip(f'{path}.1') # -> app.log.1.gz
truncate(path) # app starts fresh
# anything past `keep` is deleted from local disk
rm(f'{path}.{keep+1}.gz') # silent eviction
signal(app_pid, SIGHUP) # reopen log file
Host.serve_log
how one line ends up in app.log and later in app.log.N.gz
# the application is the only writer; logrotate is the only mover
def emit(line):
open('/var/log/app.log', 'a').write(line + '\n')
# on-disk layout after a few days of rotation:
# /var/log/app.log <- today, being appended
# /var/log/app.log.1.gz <- yesterday, compressed
# /var/log/app.log.2.gz <- 2 days ago
# /var/log/app.log.3.gz <- 3 days ago (last kept)
# anything older was rm'd by Logrotate.rotate — gone from disk.

Where this sits in Build a distributed logging stack (ELK / Loki)

Scene 01 of 12. Once your fleet has more than a handful of hosts, the only way to answer 'which host saw this error' is to ship lines centrally — local files plus ssh is an O(N) dead end.

Up next. The fix has to be central — but the moment we ship lines off-host, we put a network between the application and its own log file, and something on each host has to keep that pipe full.

All 12 scenes in Build a distributed logging stack (ELK / Loki) · Every curriculum

Built with Arqly
Every scene in Build a distributed logging stack (ELK / Loki) builds on the one before it.All 12 Build a distributed logging stack (ELK / Loki) scenes