Distributed Lock — a worked solution
Redlock controversy, fencing tokens, lease vs lock.
Try it yourself first.
You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.
Open the Distributed Lock workspaceThe problem
Build a service that lets a fleet of workers acquire a named lock and guarantee only one holder runs the protected critical section at a time. Sounds like a one-day project — until you read the Kleppmann/antirez exchange and realize the whole problem is whether the lock survives clock skew, GC pauses, and network partitions. The disagreement is the lesson.
The reference architecture
Stage by stage
The same 10 stages the workspace walks, answered.
01Clarifications
What would you ask before drawing a single box?
Worth pinning down:
- Mutual exclusion or coordination? True mutex (only one) vs leader election (one preferred, others ready) — different solutions.
- Fairness? FIFO, or any holder ok?
- Re-entrant? Same client can re-acquire?
- Critical section duration. 50 ms vs 5 minutes vs hours flips the design.
- What's protected? A pure CPU operation? A side effect on an external system? The latter is where naive locks fail.
- Failure tolerance. Is "lock held by a dead client until lease expires" acceptable?
Assumptions:
- Use case: ensure only one worker processes a particular order, batch, or scheduled job.
- Critical sections: 1s–10min.
- Acceptable to wait briefly (seconds) on contention.
- External systems exist (DBs, payment APIs) that need protection.
02Functional reqs
What must this system actually do?
- Acquire a lock by name with a TTL (lease).
- Renew the lease while the holder is still alive.
- Release the lock cleanly when done.
- Lock acquisition is atomic: only one client succeeds when N clients race for the same name.
- Operations using the lock can detect when their lease has expired (so they don't keep doing dangerous work).
03Non-functional
What must it promise about speed, uptime and correctness?
- Safety (the hard one): Two clients should never simultaneously believe they hold the lock and act on it.
- Liveness: A dead lock holder must not block forever — TTL guarantees release.
- Latency: Acquire P99 < 50 ms uncontested.
- Throughput: ~10K acquires/s globally is plenty for most use cases.
- Availability: Lock service down means business operations queue up. Often not as critical as it sounds.
04Capacity estimation
How much load and data does this have to hold?
- Lock state size: ~200 bytes per lock × ~1M concurrent locks = 200 MB. Fits trivially.
- Acquire/renew QPS: 10K acquires/s + renewals every TTL/3 (=10s) for 1M active leases = 10K + 1M/10 = ~110K writes/s. That renewal load dominates and exceeds what a single etcd cluster comfortably sustains — push back via fewer-but-longer leases (raise TTL), shard the keyspace across clusters, or move to Redis if the resource enforces fencing.
- Watchers: Worker waiting on a lock subscribes via watch (etcd) or polling (Redis). Negligible.
The whole problem is not throughput. It's correctness under failure.
05API design
What does the outside world call, and what comes back?
POST /api/v1/locks/acquire
{ "name": "process_order:42", "ttlMs": 30000, "clientId": "worker-7-uuid" }
200 { "acquired": true, "lockId": "lock-abc123", "fencingToken": 4711, "expiresAt": "..." }
or
200 { "acquired": false, "expiresAt": "..." }
POST /api/v1/locks/renew
{ "lockId": "lock-abc123", "fencingToken": 4711, "ttlMs": 30000 }
200 { "renewed": true, "expiresAt": "..." }
or
409 { "renewed": false, "reason": "expired" | "stolen" }
POST /api/v1/locks/release
{ "lockId": "lock-abc123", "fencingToken": 4711 }
204
The fencing token is the critical part. Returned on acquire, monotonically increasing per lock name. The protected resource (e.g., the storage system) must check it on every operation.
06Data model
What gets stored, and what is it looked up by?
Backend choice — two real options:
A. Redis (Redlock or single-instance with persistence).
- Pro: fast, simple, ubiquitous.
- Con: Redlock relies on bounded clock drift, which is not a safe assumption. Kleppmann's critique: under GC pause + clock jump, two clients can both "hold" the lock — and his fix is fencing tokens. antirez's rebuttal: that same GC/clock-jump risk breaks any TTL-based lock (etcd/ZK included), not just Redlock, and safety can come from a unique-token compare-and-delete rather than only a monotonic fencing token.
- Verdict: Use Redis only if (a) the protected resource enforces fencing tokens, and (b) you accept the risk model.
B. etcd / ZooKeeper (consensus-based).
- Pro: Linearizable. Built on Raft (etcd) or Zab (ZK). Fencing tokens come naturally from the modification revision number.
- Con: Higher latency (~10ms uncontested), operationally heavier.
- Verdict: This is the correct answer for actual safety guarantees. Use it.
Recommended: etcd, with a session lease (TTL is server-side, renewed by client keepalives). The lock key is created with PutIfAbsent; the modification revision becomes the fencing token. On client crash, the session expires and the lock is freed automatically.
Schema (etcd):
/locks/{name}→{ ownerClientId, leaseId }, with a TTL'd lease.- Modification revision serves as the fencing token.
The fencing token must propagate. It's not the lock service's job to prevent misuse — that's the protected resource's job.
07High-level design
Which components handle a request, and in what order?
Architecture:
- Client (worker). Calls
Acquire(name, ttl). On success, gets a lockId + fencing token. - Lock Service (stateless). Talks to etcd. Translates HTTP to etcd primitives.
- etcd cluster. 3 or 5 nodes, Raft consensus, the actual source of truth.
- Protected resource (e.g., orders DB). On every write, requires the fencing token to be ≥ the highest token previously seen for this resource. Reject if stale.
- Renewal loop. Client runs a background renew every TTL/3. If renewal fails (expired or stolen), the client must STOP doing protected work — and treat any in-flight writes as suspect.
The fencing dance:
Client A: acquire(order_42) → token=10, ttl=30s, starts work
Client A: GC pauses for 60s
Client A's lease expires after 30s
Client B: acquire(order_42) → token=11
Client B: writes (token=11) to DB. DB updates highest_seen=11.
Client A: wakes up, writes (token=10). DB rejects: 10 < 11.
Without fencing, Client A's stale write succeeds and corrupts the order.
Why a Lock Service in front of etcd? Auth, rate-limit, observability, abstraction so callers don't need etcd client libraries directly. It must remain stateless — the source of truth is etcd.
08Deep dives
Which part breaks first, and what do you do about it?
1. Why Redlock is controversial. Kleppmann's argument: distributed locks built on TTL alone require a bounded-drift assumption between clocks on different machines. Real OSes can pause processes (GC, swap, VM migration) for arbitrary durations and clocks can jump. With these, two clients can both believe they hold the lock. Kleppmann's fix for correctness is fencing tokens — a monotonic number the protected resource checks — and he notes that once you have them the lock algorithm barely matters. antirez's rebuttal pushes back on both points: the GC-pause / clock-jump scenario breaks every TTL lock (ZooKeeper and etcd included), so it isn't Redlock-specific; and he's skeptical of the fencing framing, noting the resource often can't supply a monotonic token and that a unique random token checked via compare-and-delete provides safety another way. The lesson: TTL-based locks are unsafe for correctness on their own; if you need correctness, fencing (a monotonic token the resource enforces) is the robust fix — and it makes the underlying lock service almost interchangeable.
2. Lease vs lock — what's the real difference? A lock is forever-until-released. A lease is forever-until-released-or-TTL. In distributed systems we always use leases — there's no other way to handle a dead holder. So "distributed lock" is really "distributed lease with monotonic fencing tokens."
3. What if the resource doesn't support fencing? Then your lock has no safety guarantee. Either:
- Wrap the resource in your own layer that does fencing (writes to an intermediate store with a fencing-checked unique constraint).
- Accept the risk and use locks only for performance hints, not correctness.
4. Reentrancy. A worker might call into a function that also wants the same lock. Solutions:
- Track holder identity at the service layer; same client_id = re-entrant ok.
- Or design code to not need reentrancy (preferred — re-entrant locks have always been a code smell).
5. Lock acquisition fairness under contention. etcd lock recipe gives you fair queue ordering using sequenced keys (each waiter creates a unique sequenced key, watches the predecessor). FIFO is mostly free. Redis-based locks need a separate fair-queue layer (e.g., Lua script + sorted set).
09Trade-offs
What did this design cost, and what breaks at 10×?
Comparing the options:
| Aspect | Redis (Redlock) | etcd / ZK | Postgres advisory lock |
|---|---|---|---|
| Latency | ~1 ms | ~10 ms | ~5 ms |
| Safety (with fence) | Acceptable | Strong | Strong (single DB) |
| Multi-DC | Hard (clocks) | Native | Hard |
| Operational weight | Light | Medium | Light (if you have PG) |
| Fencing token | Manual | Native | Manual |
Failure stories:
- etcd leader election during acquire: Acquire blocks for a few hundred ms then succeeds. Acceptable.
- Network partition: Quorum side keeps working; minority side can't acquire/renew. Locks held on minority side will lapse via lease expiry. Correct behavior.
- Client GC pause longer than TTL: Lease expires; another client acquires. Original client's writes still carry the stale fencing token, and the protected resource must reject them. If the resource doesn't enforce fencing, you have a bug — that's the lesson of the Kleppmann/antirez exchange.
- Lock service down: Workers can't acquire new locks — and because renewals also flow through the service, existing holders fail to renew and their leases expire within one TTL (30s). Holders must stop protected work when renewal fails, so an outage shorter than the renewal margin (~TTL/3) passes unnoticed; anything longer aborts every in-flight critical section. Mitigation: the service is stateless — run several replicas behind a LB. Catastrophic only if the lock is on the read path of high-traffic operations — and it shouldn't be.
Primary sources
- Kleppmann: How to do distributed locking
- antirez response
Now defend it
Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.
Work Distributed Lock yourselfMore in Consensus, Coordination & Durable Execution
Getting N machines to agree, and getting one job to happen exactly once: Raft, coordination services, CRDTs, locks, leader election, schedulers and durable workflows.
- Build Build Raft — consensus you can defendReplicate a deterministic state machine across N servers with safety as a theorem and liveness under partial synchrony. Build the protocol from term to commit to safety proof to reads, and feel why etcd, Cockroach, and TiKV ship slightly different Rafts.
- Build Build a workflow engine (Temporal / Airflow / Cadence style)A function that survives crashes, restarts, and re-deploys — and still finishes. Build a durable execution engine where workflow code is replayed deterministically from an event history, activities retry with exponential backoff, sagas compensate on failure, and the same workflow definition runs identically a year later. Internalize why 'just retry the cron job' breaks at the second step.
- Build Build a production AI agent (from one API call up)A model API is a function from text to text — it remembers nothing, does nothing and fails without saying so. Everything an agent actually does lives in the program around it: the loop, the tools, the window, the budgets, the gates, the sandbox, the run log. Build that program, one mechanism at a time, until it can resolve a support ticket and issue a refund exactly once.
- AI Agent PlatformLong-running multi-step LLM agents — durable workflow + sandboxed execution + LLM gateway. Brain / hands / state are independently replaceable. The agent run is a workflow, not a request.