The DB already has a log — reading the Postgres WAL and MySQL binlog

Every replicated database keeps a durable, ordered, post-commit log of every row change — Postgres calls it the WAL, MySQL the binlog — and CDC reads that log the same way a replica does, instead of querying tables.

Previously

Polling cannot be the answer — it loses deletes and collapses fast flips — so we need a source of truth that records every change in order, ideally one the database already keeps. It does: open the hood and look at the log strip.

Scene 03

The DB already has a log

  1. Watch
  2. Try it
  3. Predict
  4. Capture
DB · Postgrestablesid=42 │ status=pending │ qty=1id=43 │ status=pending │ qty=2id=44 │ status=pending │ qty=1append-only log · WALtail →replicatails WAL(idle)The WAL is empty. Fire a write to see one record appended.
What to watch for

Every replicated database keeps a durable, ordered record of every change. In Postgres it's called the WAL; in MySQL it's the binlog. Watch one INSERT, one UPDATE, one DELETE land in the tables region — and watch exactly one record per change get appended to the log strip below.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Database.applyDml
write-ahead: log the change first, then mutate the table
def applyDml(stmt):
xid = txn.current_id()
before = stmt.read_pre_image()
after = stmt.compute_post_image()
# 1. Append to the WAL (Postgres) / binlog (MySQL).
# Durable, ordered, monotonic LSN.
log.append(WalRecord(
xid, stmt.op, before, after,
))
log.fsync() # before any table page is dirtied
# 2. Only now mutate the heap / index pages.
table.apply(stmt.op, before, after)
Replica.fetchAndApply
the replica's standard fetch loop — tails the same log strip
loop forever:
# walreceiver (PG) / IO thread (MySQL):
records = primary.stream_from(
startLsn = self.confirmed_lsn,
)
for rec in records:
self.log.append(rec)
self.apply(rec.op, rec.before, rec.after)
self.confirmed_lsn = rec.lsn
primary.ack(self.confirmed_lsn)
primary.stream_from(startLsn)
the protocol replicas use — and CDC will use it the same way
# Postgres: START_REPLICATION SLOT <name> LOGICAL <lsn>
# MySQL: COM_BINLOG_DUMP_GTID <server_id> <gtid_set>
def stream_from(startLsn):
cursor = log.open_at(startLsn)
while True:
rec = cursor.next() # blocks until new WAL appended
yield rec # one row change, in commit order

Where this sits in Build a CDC pipeline (Debezium + outbox)

Scene 03 of 12. Postgres WAL, MySQL binlog — the database already keeps a durable, ordered log of every change for replication. CDC reads this log instead of the tables.

Up next. The DB already has a durable, ordered log of every change — so we need a process that tails it and turns each record into a downstream event.

All 12 scenes in Build a CDC pipeline (Debezium + outbox) · Every curriculum

Built with Arqly
Every scene in Build a CDC pipeline (Debezium + outbox) builds on the one before it.All 12 Build a CDC pipeline (Debezium + outbox) scenes