Ordering is per key, never global — hash partitioning by primary key across Kafka partitions

Kafka preserves order within a partition; Debezium hashes the row's PK (or outbox aggregate_id) to a partition, so events for one aggregate stay ordered while events across aggregates may interleave arbitrarily.

Previously

Cleanup is sorted; the outbox emits one event per business action — but those events are about to be sharded across Kafka partitions, and the question of which events stay in order matters as soon as there is more than one partition.

Scene 10

Ordering is per key, never global

  1. Watch
  2. Try it
  3. Predict
  4. Capture
Debezium c…aggregate=o…Debezium c…aggregate=o…Debezium c…aggregate=o…Debezium c…aggregate=o…hash(aggregate_id)mod 3PARTITIONS · 3P0→ ConsumerP1→ ConsumerP2→ ConsumerCONSUMER GROUP · 1 CONSUMERreads all partitions; per-partition order preservedC0owns P0, P1, P2
What to watch for

Watch the connector emit change events from the outbox. Each event's aggregate_id is hashed to a partition — events for the same aggregate (same color) always land on the same lane, and within that lane they keep commit order.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

KafkaProducer.partitionFor(key)
Hash the key, mod numPartitions — that's it.
def partition_for(key, num_partitions):
if key is None:
return sticky_round_robin() # no key = no order promise
# murmur2 in real Kafka; any stable hash works
h = murmur2(serialize(key))
# aggregate_id -> one lane per order (correct)
# tenant_id -> one lane per tenant (collapse)
# random/none -> any lane every time (kill order)
return (h & 0x7fffffff) % num_partitions
Connector.recordToKey(record, mode)
Outbox SMT picks the field that becomes the message key.
def record_to_key(record, mode):
# record came from the outbox row Debezium just read
if mode == 'aggregate_id':
return record.aggregate_id # unit of consistency
if mode == 'tenant_id':
return record.tenant_id # coarser than aggregate
if mode == 'random':
return uuid4() # different key every event
raise ValueError('unknown key mode')
Consumer.canAssumeOrder(a, b)
Per-key ordering is the only contract you may rely on.
def can_assume_order(a, b):
# 'a happened-before b' is sound iff they share a key.
if key_of(a) != key_of(b):
return False # different lanes, or hash collision
# same key -> same partition -> commit-order preserved
return True

Where this sits in Build a CDC pipeline (Debezium + outbox)

Scene 10 of 12. Kafka guarantees order within a partition; partition by aggregate_id keeps per-aggregate events ordered while accepting that cross-aggregate order is never preserved.

Up next. Within-key ordering is preserved end-to-end — but Debezium's delivery guarantee is at-least-once, so the consumer will see the same ordered event more than once on retries and restarts unless it is built to deduplicate.

All 12 scenes in Build a CDC pipeline (Debezium + outbox) · Every curriculum

Built with Arqly
Every scene in Build a CDC pipeline (Debezium + outbox) builds on the one before it.All 12 Build a CDC pipeline (Debezium + outbox) scenes