Partitions — splitting the log
A partition is the unit of parallelism. Same key → same partition (per-key order preserved); ordering is sharded across partitions. Within a group, M > N consumers leaves some idle; N > M means some consumers own multiple. Picking partition count is picking your maximum future parallelism — rebalancing later is expensive.
One log preserves order. But one log is also one machine's throughput. Splitting it into N partitions is how Kafka scales — at the cost of giving up strict GLOBAL order in exchange for per-key order.
Scene 03
Partitions — splitting the log
- Watch
- Try it
- Predict
- Capture
Watch each producer's key get hashed into a specific partition. Same key always lands in the same partition — that's how per-key order is preserved. The 2 consumers below belong to the same consumer group; they split the partitions between them.
Highlighted lines are the ones running in the diagram right now.
def partitionFor(record, partitionCount):if record.key is None:# round-robin or sticky batchreturn next_round_robin(partitionCount)h = murmur2(serialize(record.key))# toPositive masks the sign bitreturn toPositive(h) % partitionCount
def assign(members, partitions):owners = {m: [] for m in members}for j, p in enumerate(partitions):owner = members[j % len(members)]owners[owner].append(p)# members with [] never receive a fetchreturn owners
def onJoinGroup(groupId, member):g = groups[groupId] # per-group stateg.members.add(member)g.generationId += 1 # fences old assignmentsplan = assignor.assign(list(g.members), topic.partitions,)for m, owned in plan.items():m.send(SyncGroup(g.generationId, owned))
Where this sits in Build Kafka
Scene 03 of 13, in the Write side act — Partitioning, replication, and durability knobs.. Parallelism by sharding ordering. Hot partitions, key skew.
Up next. Replication — a partition lives on more than one broker. ISR (in-sync replicas) is what decides when a write is safely committed, and it's NOT a quorum vote.
Designs that use this
- Twitter / X TimelinePush or pull? Both. The canonical fanout problem.
- Uber / Lyft — Match Drivers and RidersMatch a rider to the closest acceptable driver in under 3 s. Geohash, S2, surge.
- Slack / DiscordChannels and history. Push or pull — and how a hot-channel fanout doesn't melt the gateway.
- WhatsApp / MessengerHundreds of millions of long-lived sockets, sub-second 1:1 + group delivery, E2E-encrypted, multi-device, multi-region active-active.