Distributed Unique ID Generator — a worked solution
Generate globally unique, monotonic-ish IDs at scale.
Try it yourself first.
You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.
Open the Distributed Unique ID Generator workspaceThe problem
Build a globally-unique, k-sortable ID generator that every write in the platform depends on. The hard part isn't picking a bit layout — Twitter Snowflake's 41/10/12 has been right since 2010 — it's that ID generation is the single highest-leverage piece of state in the entire infrastructure. A 99.9% ID service drags every dependent service to 99.9%. A worker-id collision silently corrupts data at rest until a downstream consumer notices a primary-key violation hours later. A leap-second smear bug rewinds a host's clock and the next ID issued duplicates one already in storage.
The canonical here serves three caller classes side-by-side because that is what production teams actually run: (1) UUIDv7 in-process for ~85% of services that can take a 128-bit ID — zero network calls, ~2 µs per ID, immune to coordinator outages; (2) Snowflake RPC fallback for legacy callers that need 64-bit BIGINTs, sortable cursors, or polyglot interop; (3) Leaf-segment range allocator for human-facing sequential IDs (invoice numbers, order numbers) where customers expect short monotonic integers. The architecture's job is to make all three correct under the failure modes that historically cause silent dup-ID incidents — clock rewind, worker-id collision, segment-failover burn — and to detect any dup that does slip through within the next audit window.
This is a single-region active-active design with cross-region DR for the segment table. Multi-region active-active for ID generation is doable (per-region DC-id bits in Snowflake, per-region etcd cluster) but adds coordination cost without buying meaningful availability over the in-process default for the 85% of callers who don't need RPC at all.
The reference architecture
What each component is for
- Service · Feature FlagGo 1.22 + Postgres + Redis edge cache
Internal feature-flag service (LaunchDarkly-like, but homegrown over Postgres + edge cache). The SDK polls every 30 s for two flags: idgen.rpc_enabled (per-region) and idgen.dr_mode (per-region). On dr_mode=us-west-2, the SDK switches RPC fallback target to the DR region; on rpc_enabled=false, the SDK refuses RPC fallback and forces UUIDv7-only mode for the caller. Also the kill-switch surface for the deploy-controller's auto-rollback action.
Why it exists. Two scenarios need a dynamic-config surface that doesn't require code or config-file deploys: (1) a bad SDK release ships and we need to disable RPC fallback platform-wide within seconds (the kill-switch); (2) regional DR step #1 ("freeze callers") needs to flip atomically across all callers without a coordinated rolling restart. We considered baking the flag into the gateway's response headers, rejected because the SDK boundary is where the policy lives — the gateway can't disable a UUIDv7 caller because the gateway never sees that traffic.
When it fails. Flag service down → SDK uses last-known cached flag values; no per-call impact for up to 30 s × N (cache survives across polls until pod restart). Catastrophic outage > 30 min → DR-mode flips can't propagate, runbook is to bypass the flag service via a coordinated SDK config-deploy (slower, but safe). Flag misconfiguration (operator types dr_mode=us-west-3 instead of us-west-2) → callers route to nonexistent region; mitigated by enum-validated flag types and a 1% canary rollout on flag changes.
- Client (caller service + SDK)polyglot SDK (Go / Java / Python / Rust)
Internal service that needs an ID for a write. Wraps two strategies in one polyglot SDK: (1) generate a UUIDv7 in-process via RFC 9562 — zero network calls, ~2 µs/ID, the default for ~85% of callers; (2) RPC fallback to the Snowflake service for cross-language services or callers that need a 64-bit BIGINT instead of a 128-bit UUID. The SDK also wraps the Leaf-segment 'allocate-range' API for callers that need short, sequential, human-readable IDs (invoice numbers, support tickets).
Why it exists. Modeled as a node because the SDK is the single load-bearing decision surface — 'do I take a network hop on every write?' The cheapest correct answer is 'no, generate UUIDv7 in-process' (Shopify's payment system reported 50% INSERT improvement vs UUIDv4). We considered making this an external library only and not drawing it on the canonical, but on-call needs to see the SDK as the first hop because the most common ID outage is a bad SDK release shipping with a clock-rewind bug — that's the node that pages.
When it fails. Bad SDK release ships a regression in UUIDv7 monotonicity (per-backend counter forgotten on a clock_gettime branch) → duplicate IDs across processes inside the same ms. Detected by the downstream Uniqueness Audit Worker within ~24 h, by which point ~10 K dups can be in storage. Mitigation: progressive delivery on platform-libs (1% → 10% → 100%) gated on a synthetic write-then-uniqueness probe per canary; SDK includes a self-check that hashes the last 1 M IDs and asserts uniqueness in-process at debug builds.
- API GatewayEnvoy 1.30 + SPIRE ext-auth
Internal-only API surface for the ID platform. Terminates mTLS from caller services, validates SPIFFE SVIDs via ext-auth, enforces per-caller rate limits (50 KQps default, override per service), routes /snowflake/* to the Snowflake service and /segment/* to the Segment Allocator. NOT a public surface — there is no internet edge in front of this; ID generation is an internal infrastructure dependency.
Why it exists. Single tenancy / authn enforcement point so the Snowflake and Segment services can stay focused on the ID-issuance hot path. Without a gateway, every backend grows its own SVID validation, its own rate limiter, and its own observability — and we lose the per-caller blast-radius cap. Considered exposing services directly via the mesh sidecar (no central gateway), rejected because cross-service rate limits and audit logging needed a single chokepoint when one caller starts misbehaving.
When it fails. Gateway down → the RPC and segment paths fail (~15% of platform writes) but UUIDv7-in-process callers (~85%) keep working — graceful degradation. Multi-AZ deploy + LB health-checks ensure single-AZ loss is automatic. Whole gateway tier dies → mesh sidecars fail open to a static fallback list of Snowflake-svc IPs (last-resort fast path, no rate limit) for 60 s while gateway recovers; auditable via mesh access logs.
- Auth / Identity (SPIRE)SPIRE Server 1.9 (HA on embedded etcd)
Issues short-lived (1 h) SVIDs to every workload in the platform via the SPIRE agent on each node. Gateway's ext-auth filter calls SPIRE Server to validate inbound SVIDs on every request. Also rotates the Snowflake/Segment service certs continuously — no static service-account credentials anywhere in the platform.
Why it exists. ID generation is the most-called internal API in the platform — every write depends on it — so a stolen mTLS cert or shared API key would be the highest-leverage credential in the system. Short-lived SVIDs cap the blast radius of a leaked cert at the rotation window. The alternative (long-lived static service tokens) was rejected because cert-rotation incidents were paging on-call quarterly.
When it fails. SPIRE Server outage → existing SVIDs continue to validate via the agent's cached trust bundle for the SVID TTL (1 h). New workloads (autoscale, deploys) can't get SVIDs and won't pass auth → blocks rolling deploys but doesn't break the steady-state ID hot path. Hard outage > 1 h → all SVIDs expire and the entire platform stops accepting RPCs; on-call action is restore SPIRE from etcd snapshot, no data loss.
- Snowflake ID ServiceGo 1.22 gRPC, SPIFFE mTLS, chrony local
Issues 64-bit Snowflake-shape IDs over gRPC: (1 sign | 41 ts-ms | 10 worker | 12 seq) — Twitter 2010 layout, custom epoch 2020-01-01. Each pod holds an etcd-leased worker-id and generates IDs locally with no per-call coordination. Reads CLOCK_MONOTONIC for sequence-delta within a ms; reads CLOCK_REALTIME (chrony-disciplined) only for the high-bit timestamp. Halt-on-rewind: refuses to mint if the wall-clock goes backward by more than 5 ms within a single pod's session.
Why it exists. Some callers need 64-bit IDs — legacy schemas with BIGINT FKs, sortable database cursors, ProtoBuf int64 fields where switching to a 128-bit type is a multi-quarter migration. UUIDv7 in-process is the right default but doesn't help these callers. Considered making them all migrate to UUIDv7, rejected because the migration cost (schema + downstream consumers + analytics queries) is months and the RPC fallback is operationally cheap once the bit-budget is sized. Twitter's published ~10K/sec/process realizes ~30K/sec in modern Go gRPC without GC pauses — we size the fleet on the realized number, not the bit-budget ceiling.
When it fails. Single pod with bad clock (NTP drift > 5 ms) → halt-on-rewind kicks in, pod refuses to mint, gRPC returns UNAVAILABLE → mesh load-balancer routes around it within one health check (~5 s). Whole fleet sees correlated NTP drift (bad upstream stratum-1) → all pods halt-on-rewind together → ID RPCs fail platform-wide, BUT UUIDv7-in-process callers (85%) keep working. On-call signal: idgen_clock_rewind_total > 0 over 1 min (Cloudflare-2017 leap-second class). Bad deploy that breaks the lease-renewal loop → leases expire after 10 min, pods refuse to mint, platform write QPS for legacy callers drops to zero — auto-rollback gated on synthetic write-then-read uniqueness probe.
- Segment Allocator (Leaf-segment)Go 1.22, pgx, pre-fetch goroutine
Hands out ID ranges via Meituan Leaf-segment double-buffer pattern: each pod holds an active in-memory segment of step=1000 IDs and a pre-fetched 'next' segment. When the active segment crosses 90% used, a background goroutine UPDATEs the (biz_tag, max_id) row in Postgres and atomically swaps the next-buffer in. Per-call latency is sub-ms (in-memory counter bump + bounds check); a DB hit happens once per 1000 IDs.
Why it exists. Two caller classes need this: (1) human-facing sequential IDs — invoice numbers, order numbers, support ticket IDs — where customers expect a roughly monotonic, short integer; (2) Postgres B-tree FKs that prefer dense bigint over sparse Snowflake IDs (~2× tighter B-tree leaves on dense vs Snowflake-spaced inserts). UUIDv7 doesn't satisfy 'short' or 'sequential'; Snowflake doesn't satisfy 'dense'. Considered a bare DB SERIAL per biz-tag, rejected because the per-call DB roundtrip caps throughput at ~500 IDs/sec/biz-tag (Postgres synchronous_commit=on, 5-15 ms tx) and we need ~50K IDs/sec aggregate.
When it fails. Segment-table primary down → pods serve from active + pre-fetched buffers (~9 min worth of headroom at peak QPS for the typical biz-tag) before refusing further allocates. Post-failover, dual-master odd/even on the segment index (each master advances max_id by 2 × step, owning alternate step-blocks — the Flickr pattern generalized to ranges) means the surviving master keeps issuing without waiting on promotion. Segment burn on pod restart is bounded to step (1000 IDs/biz-tag/restart) — leaves gaps in the ID sequence that downstream analytics must tolerate. Hot biz-tag (one tenant burning segments 100× faster than others) → DB write contention on that one row; mitigated by per-biz-tag rate-limit at the gateway and an alert on segment_fetches_per_biz_tag_per_min.
- Lease Manager (Worker-ID Reconciler)Go 1.22, etcd clientv3
Background controller that grants and renews worker-id leases. On Snowflake pod boot, the pod calls the Lease Manager's local sidecar which acquires an EPHEMERAL_SEQUENTIAL slot in etcd (/snowflake/{region}/workers/). The Manager renews leases every 1 min; on session-loss it FAIL-STOPs the pod (forces a restart and re-acquire) rather than silently re-claiming the same slot — Discord's published lesson on dup-IDs from silent re-acquire.
Why it exists. Worker-id assignment is the single highest-leverage piece of state in the system: two pods with the same worker-id silently corrupt the entire ID space. Concentrating the lease lifecycle in a dedicated reconciler (vs each pod talking to etcd directly) lets us enforce the fail-stop policy uniformly, instrument lease churn centrally, and gate the cold-start thundering herd at the SDK boundary. Considered putting this logic in a library inside the Snowflake service, rejected because the policy change ('stop minting if your lease has expired') needs to be enforceable from outside the pod when the pod's clock has rewound.
When it fails. Lease Manager down → existing Snowflake pods continue minting from cached leases (fail-static, up to lease TTL = 10 min). New pods can't boot — autoscale freezes, but steady-state ID hot path keeps working. Manager fleet AND etcd both down → hard fail at lease-TTL expiry; on-call action is restore etcd from snapshot then resume Manager. Bug in the renewal loop (e.g. silently retrying a session-expired lease) is the worst-case — manifests as worker-id collision detected by downstream Uniqueness Audit; runbook is freeze deploys, identify the bad pod via count(distinct host) by (worker_id), fail-stop it, then root-cause.
- Lease Cache (Redis)Redis 7 (cluster mode)
Read-through cache for worker-id leases — the Snowflake service consults this on every Nth ID generation (default: every 1024 IDs) to verify its lease is still valid before issuing the next batch. Holds (worker_id → {host, lease_expires_at, dc_id}) keyed by worker-id, plus a reverse (host → worker_id) index for diagnostics. TTL set to lease-expires-at; eviction policy ttl-only so we never serve a logically expired lease.
Why it exists. Without a cache, the verify-lease check would have to hit etcd on every ID call (or every 1024) — adds 5-10 ms of network latency to a sub-ms operation, defeats the point of in-pod ID generation. We considered just caching the lease in pod memory (no shared cache), rejected because the Lease Manager's emergency 'revoke' command needs a place to publish that the pod sees within milliseconds — the cache is the broadcast surface. Hot path = cache hit (~99.9%); cold path = cache miss → etcd fetch → cache write.
When it fails. Cache cold (post-restart, all keys evicted) → first ID call per pod cache-misses, falls through to etcd → 12 pods × 1 lookup = 12 etcd reads, trivially absorbed. Cache cluster fails completely → Snowflake pods fall back to a 1-second pod-local TTL on their cached lease (degraded mode). The dup-window math: 1 s × 30 K IDs/sec/pod = 30 K worst-case silent dup IDs per concurrently-misissued slot, IF (a) an emergency Lease Manager revoke is in flight AND (b) the freed slot is granted to another pod within that 1 s window. The streaming overlap probe is sized to catch this within seconds. Accepted because (1) emergency revokes happen < 1×/quarter, (2) the bounded blast radius is observable rather than open-ended. On-call signal: lease_cache_hit_rate < 0.99 for 5 min → ticket; lease_cache_unavailable_total > 0 for 1 min → page (because of the dup-window math above, not because of latency).
- Coordinator (etcd)etcd 3.5 (5-node, 3-AZ)
Source of truth for (1) Snowflake worker-id leases (/snowflake/{region}/workers/{seq} ephemeral keys), (2) Lease Manager leader election, (3) Segment Allocator service leader for the hot biz-tag pre-fetch coordinator. Watches let the Lease Manager and Lease Cache react to lease changes within the etcd --election-timeout budget (default 1 s).
Why it exists. Worker-id leasing requires linearizable coordination — two pods cannot agree to share a slot, ever. Raft gives us linearizable writes with provable safety under network partitions. Considered ZooKeeper (Twitter's original choice) — rejected because etcd's HTTP/2 + gRPC is friendlier to operate, and the SPIFFE story is cleaner. Considered no coordinator at all (worker-id from k8s downward-API: hash of pod-name) — rejected because pod-name collisions across regions would silently allocate the same worker-id, and we'd lose the centralized fail-stop enforcement.
When it fails. Quorum loss (2 of 5 nodes alive) → all writes blocked, watches stale → no new leases granted. Existing Snowflake pods keep minting from cached leases for up to 10 min (lease TTL); steady-state ID hot path unaffected. Manager fail-stops pods whose cached leases expire if it can't renew → progressive write outage for legacy callers, with pace = lease-TTL / (number of pods at lease-expiry concurrence). On-call signal: etcd_server_has_leader == 0 for 30 s → page; etcd_server_proposals_failed_total rising → ticket. Recovery = restore the lost-AZ etcd nodes and let the cluster re-converge.
- SQL DB · Segment PrimaryPostgres 16 (dual-master + DR replica)
Stores the (biz_tag, max_id, step, version, updated_at) row per business type that uses the segment allocator. Every segment fetch is UPDATE leaf_segment SET max_id = max_id + step, version = version + 1 WHERE biz_tag = ? RETURNING max_id. Single-row hot, ~50 segment-fetches/sec at peak (50K IDs/sec / step 1000). Also stores the audit_event table (segment-issued events) written in the same transaction as the segment UPDATE — that's the load-bearing outbox.
Why it exists. Centralized source of truth so every Segment Allocator pod sees the same monotonically-increasing max_id per biz-tag. Considered a fully decentralized Sonyflake-style allocation (each pod owns a worker-id and embeds it in the ID), rejected because the human-facing 'short, sequential' requirement is incompatible with worker-id bits in the ID. The dual-master pattern (Flickr odd/even applied to the segment index, so each master advances its cursor by 2 × step and owns alternate step-blocks) survives a single primary loss without write outage at the cost of accepting non-strict monotonicity across the two masters.
When it fails. Primary down → segment fetches fail; Allocator pods serve from buffer (~9 min headroom at peak). Dual-master mode means the surviving master keeps issuing without waiting on Patroni-style promotion. Replication lag spike → DR replica stale; failover would issue dups within the lag window — mitigated by SAFETY_BURN. Hot biz-tag DOSes the row with sub-100 ms UPDATE contention → row-lock waits show in pg_locks. Explicit knobs: gateway per-biz-tag token-bucket = 200 segment-fetches/min default (override per-tag); alert at segment_fetches_per_biz_tag_per_min > 1000 (ticket, SRE-platform owner) and > 5000 (page, on-call SRE-platform). Auto-step-bump (1000 → 10000) fires on > 5000 fetches/min sustained 5 min — flips that biz-tag from 1 fetch/sec to 1 fetch/10s at the cost of larger restart-burn.
- SQL DB · Segment DR ReplicaPostgres 16 logical replica
Async logical replica of the Segment Primary, hosted in a different AZ from both masters. Two jobs: (1) DR target — promoted on a regional Primary loss (RTO ~5 min, manual gate because dual-master already covers AZ loss); (2) read-only target for the nightly Uniqueness Audit Worker, which scans the audit_event table for duplicate IDs without putting query load on the production primaries.
Why it exists. Dual-master gives us AZ-loss survival but not a clean snapshot for the audit. The nightly segment range-overlap self-join over segment.allocated rows is heavy (sequential scan over 24 h of issuance events) and would saturate the primaries' buffer cache mid-day (the worker-id lease-overlap check runs against the audit-sink Parquet, not here — lease events never land in Postgres). Rejected: running the audit against the primaries with a low-priority resource group — Postgres doesn't have a robust IO priority story, and the audit's worst-case latency is hours.
When it fails. Replica down → audit job has no read target → uniqueness audit lapses → silent dup IDs could accumulate undetected for days. Detection: audit_job_last_success_age_hours > 25. Mitigation: secondary audit job runs against the surviving primary at 04:00 UTC (degraded mode, with rate-limited reads). DR scenario: replica behind primaries by > 5 min lag triggers a page (RPO breach).
- External · NTP / Chrony Poolchrony 4.x + GPS stratum-1
Out-of-process time source. Each Snowflake pod runs chrony as a host daemon syncing to two stratum-1 sources (cellular-grade GPS-disciplined oscillators on bare-metal NTP servers, plus Cloudflare's time.cloudflare.com as backup). The Snowflake process itself reads clock_gettime(CLOCK_REALTIME) (vDSO, ~25 ns) for the high-bit timestamp and CLOCK_MONOTONIC for the intra-ms sequence delta — chrony only adjusts the wall clock, never the monotonic.
Why it exists. Snowflake's correctness invariant is lastTimestamp ≤ now() — a clock rewind breaks it. Bad NTP is the #1 documented cause of distributed-ID dups (Cloudflare 2017 leap-second; Meta migrated from ntpd to chrony to drop drift from 10 ms to 100 µs). We model NTP as an explicit external dependency rather than pretending it's free because the dependency contract — '|drift| ≤ 5 ms over 1 min' — is what halt-on-rewind enforces against. Considered relying on cloud-provider NTP only (AWS Time Sync), rejected because correlated drift across an entire AZ is a real failure mode and a second independent source is cheap insurance.
When it fails. Single NTP source goes bad → chrony picks the other; no impact. Both NTP sources go bad → chrony enters 'unsynchronized' state, drift grows monotonically. Snowflake pods detect via monotonic_vs_wall_drift_ms > 5 and refuse to mint. Leap-second style event (clock jumps backward) → halt-on-rewind triggers within microseconds of the rewind; gRPC returns UNAVAILABLE; mesh routes around the affected pod. Detection: page on idgen_clock_rewind_total > 0 over 30 s. Catastrophic: bad NTP source feeds a leap-second-style smear bug into all pods simultaneously → entire fleet halts; UUIDv7 callers (85% of platform) keep working.
- Stream · Audit EventsKafka 3.7 (3 brokers across 3 AZs)
Kafka topic carrying the structural issuance events: snowflake.lease.issued (worker-id grants, renewals, revokes) and segment.allocated (each Leaf-segment range [from_id, to_id]). It deliberately does NOT carry per-ID events — ~13 B individual IDs/day would be prohibitive — so a duplicate is detected structurally: two overlapping segment ranges, or two overlapping worker-id leases. Consumers are (a) the Audit Sink that archives to S3, (b) the Uniqueness Audit Worker. Producers are the Snowflake and Segment services; for the Segment side this is the outbox half of the segment-allocation transaction.
Why it exists. Without a centralized audit log, after-the-fact 'why did we issue ID X' forensics is impossible — you'd have to scrape logs across 18 pods plus the segment-table version history. The audit stream is also the substrate for the uniqueness probe: the consumer can spot dups in an issuance window without querying every downstream consumer table. Considered writing audit rows to Postgres directly (synchronous), rejected because the per-call latency budget for ID issuance is < 1 ms and we cannot afford a synchronous fanout to a separate store.
When it fails. Kafka cluster down → audit producers buffer in pod-local disk-backed outbox (60 s capacity); under prolonged outage, ID-issuance still works (we don't block on audit) but the audit log gets a gap. Detection: audit_outbox_lag_seconds > 60 → page. Hot partition (one biz-tag dominates) → consumer lag grows on that partition; mitigation is partition reassignment via the partition-key hash including a sub-key. Poison message → DLQ topic + alert on first DLQ message.
- Worker · Audit Outbox DrainerGo 1.22 + pgx + Sarama Kafka producer
Background worker that polls the audit_event table on the Segment Primary every 200 ms for rows with published=false, publishes them to the audit-stream Kafka topic, and on Kafka-publish-success marks the row published=true (it does NOT delete it — audit_event doubles as the retained audit trail the nightly uniqueness scan reads over a 25 h window; space is reclaimed by dropping day-partitions after 90 d, not per-row). Idempotent on the Kafka side (enable.idempotence=true + event_id as the dedup key on consumers). This worker — not the Segment Allocator — is what makes the segment+audit outbox actually drain to Kafka.
Why it exists. The segment UPDATE and the audit_event INSERT are written in the SAME Postgres transaction (single-store outbox); the actual Kafka publish has to happen out-of-band, otherwise the segment service would block on Kafka availability and we'd lose the at-least-once delivery contract. Considered making the segment service itself drain its own outbox in a goroutine, rejected because (a) the goroutine's failure modes are invisible to on-call (no separate alert surface), (b) restart of the segment service would orphan the in-flight publishes. A dedicated drainer with its own SLI (audit_outbox_drainer_lag_seconds) and its own deploy unit is operationally correct.
When it fails. Drainer down → audit_event rows accumulate in Postgres; segment-allocator hot path unaffected (it doesn't depend on the drainer). Detection: audit_outbox_drainer_lag_seconds > 60 ticket, > 600 (10 min) page. Worst case: drainer fleet down for hours → audit_event table grows; mitigation is the table is partitioned by issued_at-day, capacity 100 M rows/day fits comfortably. Drainer publishes a row but crashes before marking it published → Kafka has the event, Postgres still shows it unpublished, next poll re-publishes → at-least-once delivery; consumers dedup on event_id.
- Object Store · Audit ArchiveS3 + Glacier IR + Kafka Connect S3 sink
Long-term cold storage of every ID-issuance audit event. A Kafka Connect S3 sink consumer batches events from audit-stream into 5-minute Parquet files partitioned by (date, svc). Holds the full historical record of which worker_id was issued to which host at which lease_start — the substrate for any after-the-fact compliance or debugging query.
Why it exists. Kafka retention is 30 days; audits, security investigations, and data-lineage queries can require years of retrospect (SOX 7 years for financial systems, GDPR-defensible retention windows for compliance). S3 + Glacier is the cheapest durable substrate. Considered a wide-column DB (Cassandra) for queryable audit history, rejected because we have no real-time query pattern on years-old audit data — the queries are batch-analytics in a warehouse, and Parquet-on-S3 is the warehouse-native format.
When it fails. S3 PUT errors → Kafka Connect retries with exp-backoff, holds offset commit until success; under prolonged S3 outage, Kafka retention (30 d) is the buffer. Sink fails to keep up → consumer lag grows; alert on s3_sink_lag_minutes > 30. Worst case: silent corruption of a Parquet file — mitigated by daily checksum job that re-reads a sample of files and verifies against the source Kafka offsets.
- Worker · Uniqueness AuditGo 1.22, in-memory interval index (RoaringBitmap for range sets)
Two pipelines in one worker fleet, both structural (the stream carries ranges + leases, not per-ID rows). (1) Streaming probe — consumes audit-stream and maintains, per biz_tag, an interval index of allocated [from_id, to_id] ranges; any new segment.allocated that overlaps an existing range is a dup candidate, and for snowflake.lease.issued any two live leases for the same worker_id with overlapping validity are a dup candidate. Overlap detection is exact, so a hit is a confirmed dup — the worker pulls the two conflicting records from the audit-sink Parquet for the on-call. (2) Nightly batch probe — at 03:00 UTC over the last 25 h: a range-overlap self-join on segment.allocated against the Segment DR Replica (a.biz_tag = b.biz_tag AND a.from_id <= b.to_id AND b.from_id <= a.to_id AND a.event_id <> b.event_id — segment rows are txn-guaranteed in Postgres), plus a worker-id lease-overlap pass over snowflake.lease.issued read from the audit-sink Parquet (lease events are published straight to Kafka and archived to Parquet — they are never written to Postgres, so the lease check cannot run against the DR replica). Any overlap fires a P0 page.
Why it exists. ID-uniqueness is the platform's single hardest-to-test invariant — it can be silently wrong for hours before any caller notices, and by then dups are durable in storage across many tenant systems. The downstream uniqueness probe is the load-bearing SLI; without it, the failure modes that produce silent dup IDs (worker-id collision, clock rewind on one pod, segment-failover-burn) are undetectable. Considered making each consumer system run its own dup-check, rejected because that's N teams owning the same control and they will skip it; a centralized probe per audit-event source is correct and observable.
When it fails. Streaming probe consumer lags → window grows beyond 24 h → an overlap against a range/lease already evicted from the in-memory window could be missed by the streaming path (the nightly scan + the audit-sink Parquet cross-check backstop it). Detection: audit_consumer_lag_seconds > 300 page. On restart the interval index is rebuilt from the audit-sink. Worst: the worker silently fails AND the nightly scan also lapses → silent dup window of days; mitigated by audit_job_last_success_age_hours > 25 paging at 25 h regardless of cause.
- External · KMS (audit + SVID keys)AWS KMS (per-region, HSM-backed)
AWS KMS (or equivalent HSM-backed key store) holding two key categories: (1) the SSE-KMS data key for the audit-sink S3 bucket — every Parquet write does an AWS:Encrypt call to wrap a per-object data key; (2) the intermediate CA signing key the SPIRE Server uses to issue SVIDs. KMS calls are out of the per-call ID hot path — only audit writers and the SPIRE Server hit KMS, and SPIRE caches signed material aggressively.
Why it exists. Modeled explicitly because KMS outage is a real correlated-failure mode: it disables both audit immutability (S3 sink can't encrypt new objects) and SVID rotation (existing SVIDs survive their TTL but no new ones can be minted). On-call must see the KMS dependency on the canonical to understand why a regional KMS incident pages the ID platform — without this node, KMS shows up as a phantom dependency in post-mortems.
When it fails. KMS regional outage → S3 audit-sink writes fail with KMS:Encrypt error → Kafka Connect S3 sink retries with exp-backoff and pauses offset commits → consumer lag grows on audit-stream. Detection: s3_sink_kms_error_rate_per_min > 0 for 5 min → page. SPIRE Server: existing SVIDs (1 h TTL) keep validating, but no new workloads can boot — autoscale freezes. KMS keypair rotation incident (unlikely but real): old key disabled before all consumers can re-key → audit replay can't decrypt historical Parquet; mitigated by a 30 d key-disable grace window before deletion.
- Monitoring (Prometheus + Alertmanager)Prometheus 2.50 + Alertmanager
Scrapes metrics from every node in the platform every 15 s; evaluates burn-rate alerts on the four load-bearing SLIs: (1) ID-issuance availability (99.999% SLO), (2) duplicate-ID rate (target = 0; alert on first occurrence), (3) lease-renewal success (any failure = ticket; sustained = page), (4) NTP sync (chrony fell out of sync = page within 5 min). Pages via PagerDuty; tickets via the standard issue tracker.
Why it exists. ID generation is the platform's single highest-leverage dependency — a 99.9% ID service degrades the entire platform to 99.9%; we need explicit observability to claim 99.999%. Without burn-rate alerts on the four SLIs, on-call learns about ID outages from downstream consumer pages (which take longer to fire and have wider blast radius). Considered using cloud-provider managed monitoring (CloudWatch), rejected because the dup-ID burn-rate alert needs custom recording rules and per-SLO budget math that we want explicitly authored.
When it fails. Prometheus down → alerts don't fire; the platform keeps working but on-call is blind. Mitigated by a heartbeat alert (Prometheus pushes a 'I am alive' metric to a dead-man's-switch external service every 30 s; on absence > 2 min, the external service pages). Alertmanager misroute → a dup-ID page goes to the wrong on-call; mitigated by routing-config-as-code with PR review.
Stage by stage
The same 10 stages the workspace walks, answered.
01Clarifications
What would you ask before drawing a single box?
Typical clarifications to surface:
- Do callers need 64-bit IDs or are 128-bit OK? Drives UUIDv7 vs Snowflake choice. Postgres/MySQL BIGINT FKs need 64-bit; new services on UUID-native schemas can take 128-bit.
- Are IDs ever shown to humans? If yes (invoice numbers, support tickets), Snowflake's timestamp prefix is unfriendly; segment allocator's short sequential integers are correct.
- Is monotonicity required across processes / regions, or per-process only? Snowflake gives per-worker monotonic; UUIDv7 gives per-process monotonic; segment allocator gives roughly monotonic per biz-tag (small inversions across the two masters).
- What's the duplicate-ID tolerance? For most platforms, the answer is "zero, ever" — which forces the downstream uniqueness probe and the fail-stop policies on coordinator/clock anomalies.
- Cross-region active-active? If yes, partition the worker-id space by DC bits (Discord 5/5 split). If no, single-region with cross-region DR replica is enough.
- Audit retention? Compliance (SOX, GDPR, sector-specific) drives the audit-sink retention story. Default 90 d hot + 7 y cold.
Assumptions stated:
- 1 M total platform writes/sec at peak. 85% via in-process UUIDv7 (no RPC). 10% via Snowflake RPC = 100 K Snowflake QPS. 5% via segment allocator = 50 K segment-IDs/sec ≈ 50 segment-fetches/sec.
- Single AWS region with three AZs; one cross-region DR replica for the segment table.
- Custom Snowflake epoch 2020-01-01 (wraps 2089-09 — tracked).
- 10-bit worker-id space = 1024 slots; 12 active Snowflake pods → 99% headroom.
- No ID is a secret — auth lives separately. IDs may leak timestamps and rough rate (Snowflake) or be effectively opaque (UUIDv7 random 74 bits).
02Functional reqs
What must this system actually do?
- Issue Snowflake-shape 64-bit ID via RPC for legacy / polyglot callers.
- Allocate ID range via Leaf-segment for callers that need sequential, short, human-facing IDs.
- Generate UUIDv7 in-process via the SDK for the default 85% of callers (no network call).
- Audit every issued lease + range to a Kafka stream and S3 archive.
- Detect duplicate IDs via a streaming range/lease-overlap probe + nightly DB scan.
- Lease worker-IDs via etcd with EPHEMERAL_SEQUENTIAL slots; renew or fail-stop.
- Halt-on-rewind if the wall clock goes backward by more than 5 ms.
- Survive single-AZ loss with zero ID-issuance impact (multi-AZ etcd, dual-master segment, multi-AZ Snowflake fleet).
- Survive single coordinator outage for up to 10 min (lease TTL) with no ID-issuance impact (fail-static on cached lease).
03Non-functional
What must it promise about speed, uptime and correctness?
- Availability: 99.999% on the ID-issuance API (5.26 min annual budget). Why so high?
A_platform = A_id × A_db × A_app— a 99.9% ID layer caps the platform at 99.9% no matter how good downstream is. The 99.999% target is achievable because (a) the in-process UUIDv7 path is the default, with effectively 100% availability, (b) the RPC and segment paths are fail-static off cached state, surviving most coordinator/cache outages without user-visible impact. - Duplicate-ID rate: 0 over any 24 h window. There is no acceptable burn rate on duplicate IDs — they corrupt data at rest. Detection SLI is the streaming range/lease-overlap probe + nightly DB scan; a single confirmed dup = P0 page.
- Latency:
- UUIDv7 in-process: p99 < 5 µs (CSPRNG draw + clock_gettime via vDSO).
- Snowflake RPC: p99 < 30 ms intra-DC (within Twitter's published p99 envelope; Discord's Go gRPC achieves < 5 ms).
- Segment allocator: p99 < 5 ms warm path (in-memory counter); p99 < 200 ms cold path (DB UPDATE).
- Durability: Worker-ID leases (etcd quorum), Segment max_id (Postgres semi-sync + DR replica). Audit events (Kafka acks=all + S3).
- Consistency: Linearizable on lease grants (etcd) and segment fetches (Postgres semi-sync). Sequential on the audit stream (Kafka per-partition order). Eventual on the audit-sink read path (S3).
- Security: All inter-service edges mTLS via SPIFFE SVIDs (1 h TTL). No long-lived static service-account credentials anywhere. Audit log is append-only with S3 object-lock.
- Multi-region / DR: Single-region active-active for the issuance hot path. Segment DB has a cross-region async DR replica with RTO ~5 min on regional loss; UUIDv7 callers are unaffected by any region loss because they hold no platform state.
04Capacity estimation
How much load and data does this have to hold?
All numbers derived from the inputs in the simulator; alter to re-derive.
Total platform writes (peak): 1,000,000 qps
via UUIDv7 in-process (85%): 850,000 qps → no platform load
via Snowflake RPC (10%): 100,000 qps → snowflake-svc fleet
via segment allocator (5%): 50,000 qps → 50 segment-fetches/sec
Snowflake fleet sizing. Twitter quoted ~10K IDs/sec/process for Snowflake circa 2010 (Java + Thrift + JVM GC). Modern Go gRPC without GC pause penalties realizes ~30K/sec/pod. We size on the realized number, not the bit-budget ceiling.
Bit-budget ceiling per pod = 2^12 × 1000 ms/sec = 4,096,000 IDs/sec (theoretical)
Realized per pod (Go gRPC, p99 30 ms RTT) = 30,000 IDs/sec
Fleet capacity (12 replicas × 30 K) = 360,000 IDs/sec
Peak demand (10% of 1M) = 100,000 IDs/sec
Fleet headroom = 3.6× peak
Headroom on single-AZ loss (8 of 12 pods) = 2.4× peak ✓
Segment-table sizing. Each segment fetch is one Postgres UPDATE ... RETURNING round-trip. At synchronous_commit=on, the typical tx is 5-15 ms (WAL fsync dominates).
Segment fetches/sec at peak = 50,000 IDs/sec / step 1000 = 50 fetches/sec
Postgres primary capacity = 1 / 10 ms × 1 connection ≈ 100 fetches/sec (single-row contention)
Headroom = 2× peak
If we needed 10× headroom (a hot biz-tag burning segments faster), we'd raise step to 10,000 (1 fetch / 10 sec) — gives 100× the headroom for that biz-tag at the cost of larger gaps on pod restart (10 K IDs lost per crash vs 1 K).
Snowflake bit-budget exhaustion math.
Epoch wrap = 2^41 ms / (1000 × 86400 × 365.25) = 69.73 years
Custom epoch 2020-01-01 → wraps 2089-09 (tracked on capacity dashboard)
Worker-ID slots = 2^10 = 1024
Active workers / region = 12 → 99% headroom
Sequence per worker per ms = 2^12 = 4096 (4.096M IDs/sec/worker ceiling)
If we exceed 1024 workers per region (Discord did, when they expanded to dozens of DCs), we re-allocate bits to 5-bit DC-id + 5-bit worker-id (32 DCs × 32 workers = same 1024 but DC-routable). Tracked as a planned migration, not an emergency.
Storage.
- Lease table (etcd): ~120 B/lease × 12 leases per region = ~1.5 KB. Trivial.
- Segment table (Postgres): ~80 B/row × 1 K active biz-tags = ~80 KB. Trivial; WAL volume dominated by checkpoint cadence.
- Audit stream (Kafka, 30 d retention): ~150 events/sec × 80 B/event × 30 d = 30 GB. One Kafka broker absorbs this comfortably.
- Audit sink (S3, 7 y retention): Same data Snappy-compressed in Parquet ≈ 2 TB over 7 y. ~$45/month at S3 Glacier Instant Retrieval.
- Lease cache (Redis): ~120 B/lease × 12 leases = trivial. Even at 100× scale, well under one redis shard.
Why these timeouts (and the cascade math).
- Caller → gateway 80 ms (e1, exp-backoff): the per-call SLO budget for the RPC fallback path. Sized to fit two Snowflake attempts (e3 = 30 ms × 2 = 60 ms) plus 20 ms slack for retry-jitter and gateway overhead. Earlier draft used 50 ms but that left no room for retry; widened to 80 ms after the cascade-math review.
- Gateway → Snowflake 30 ms (e3, exp-backoff): gateway holds long-lived gRPC streams; Twitter published p99 = 2 ms, modern Go gRPC realizes < 5 ms typical. 30 ms is 6× typical p99 to absorb burst. Two attempts fit within the e1 = 80 ms parent.
- Snowflake → lease-cache 5 ms (e4, fixed-retries=1): same-AZ Redis hop; fail-fast, use cached state on miss.
- Lease-cache → etcd 20 ms (e5, fixed-retries=1): cold-fetch path; etcd Raft commit p99 ~10 ms across 3 AZs.
- Gateway → segment 200 ms (e7, fixed-retries=2): Postgres WAL fsync dominates (~5-15 ms typical, up to ~100 ms p99 under load); two retries × 100 ms = 200 ms, exactly fills the parent budget. The segment-row UPDATE is not idempotent —
max_id = max_id + stepadvances on every execution, so a retried allocate burns a fresh range and skips the lost one — but it is safe to retry because it can never mint a duplicate ID; the cost is a tolerated gap. (For true idempotency-on-retry, add a request/idempotency key so a retried allocate returns the previously-issued range instead of advancing again.) - Segment → seg-pri 100 ms (e8, fixed-retries=2): the segment service explicitly sets
SET LOCAL statement_timeout='80ms'per attempt because pgx does NOT propagate gRPC deadlines — seesegment.keyChoices. Two attempts × 80 ms statement-timeout + 40 ms slack fits in e7 = 200 ms. - No retry storm under partial failure: the cascading chain converges because at every level callee timeout × max-retries ≤ caller per-attempt timeout — enforced per route, since the SDK applies a route-specific caller deadline (a single edge can only show one number, so the canonical draws e1 at the tighter RPC value). On the RPC route: e3 = 30 ms × 2 = 60 ms ≤ e1 = 80 ms caller budget. On the segment route the SDK uses a wider 500 ms caller deadline (the cold-path p99 budget), which covers e7 = 200 ms × 2 = 400 ms, and inside e7 the DB attempts fit too (e8
statement_timeout80 ms × 2 = 160 ms + 40 ms slack ≤ e7 = 200 ms). Because each nested budget bounds the one below it, the outermost caller deadline — not a sum of the inner timeouts — bounds the whole chain, so a caller never abandons work a callee is still retrying.
05API design
What does the outside world call, and what comes back?
POST /api/v1/snowflake/next-id
Authorization: SVID
Request: { "count": 1, "caller": "order-svc" }
// `caller` is metadata only (audit/metrics attribution) — a Snowflake ID is
// (1 sign | 41 ts-ms | 10 worker | 12 seq) minted locally from the pod's clock +
// leased worker-id; there is no per-caller/per-biz_tag dimension in the issued ID.
Response 200:
{
"ids": [1234567890123456789],
"format": "snowflake",
"issued_at": "2026-05-07T12:34:56.789Z",
"worker_id": 47
}
POST /api/v1/segment/allocate-range
Authorization: SVID
Request: { "biz_tag": "invoice", "count": 1000 }
Response 200:
{
"biz_tag": "invoice",
"from_id": 8472000,
"to_id": 8472999,
"step": 1000,
"issued_at": "2026-05-07T12:34:56.789Z"
}
GET /api/v1/health
200 { "snowflake_ready": true, "segment_ready": true, "worker_id_lease_seconds_remaining": 412 }
Why no GET /id endpoint. We deliberately don't expose a per-ID GET endpoint. The SDK's UUIDv7 path is the default; the RPC path issues a batch (default 1, but callers should request larger batches when minting many IDs at once). A single-ID GET would invite per-row foreach loops that scale linearly with ID count and create N round-trips per write batch.
06Data model
What gets stored, and what is it looked up by?
Worker-ID lease (etcd): /snowflake/{region}/workers/{seq:010d} → {"host": "snowflake-pod-xyz", "leased_at": "...", "lease_expires_at": "...", "dc_id": 1}. Ephemeral with TTL = lease TTL. Sequential numbering ensures no two pods get the same key.
Segment table (Postgres leaf_segment):
| field | type | notes |
|---|---|---|
| biz_tag | varchar(64) | PK (single row per biz_tag); both masters advance it via Flickr odd/even step-blocks |
| max_id | bigint | next ID to issue for this biz_tag |
| step | int | range size to issue per fetch (default 1000; 10000 for hot biz-tags) |
| version | int | bumped on every UPDATE to detect stale-replica writes |
| updated_at | timestamp |
Single-row hot — the contention is intentional and bounded. Each fetch is UPDATE leaf_segment SET max_id = max_id + step, version = version + 1, updated_at = NOW() WHERE biz_tag = ? RETURNING max_id. This is not idempotent — every execution advances max_id, so a retried allocate consumes a brand-new range and permanently skips the lost one — but it is safe to retry: it never issues a duplicate ID, and the skipped range is just a tolerated gap. (To make it idempotent-on-retry, carry a request/idempotency key so a retry returns the same previously-issued range rather than advancing again.)
Audit event table (Postgres audit_event) — the segment outbox, written in the same txn as the segment UPDATE. It persists segment.allocated rows ONLY; snowflake.lease.issued events are produced straight to the audit-stream (Kafka) and archived to Parquet, never written here (the event_type / worker_id columns describe the shared audit-stream event shape, so the nightly Postgres scan does segment range-overlap while the lease-overlap check reads Parquet):
| field | type | |
|---|---|---|
| event_id | uuidv7 | |
| event_type | enum | segment.allocated / snowflake.lease.issued |
| svc | text | |
| host | text | |
| worker_id | int? | for snowflake events |
| biz_tag | text? | for segment events |
| from_id | bigint? | |
| to_id | bigint? | |
| issued_at | timestamp |
A worker tails this table and publishes to Kafka (the Outbox drainer); on success it marks the row published=true (not deleted — the table is also the retained audit trail; day-partitions are dropped after 90 d). Eventual consistency between local commit and Kafka publish is bounded by the drainer lag, alerted on outbox_drainer_lag_seconds > 60.
Database choice — recommended.
Postgres 16 in dual-master mode (Flickr odd/even applied to the segment index max_id/step, each master advancing by 2 × step so their step-blocks never overlap) for the segment table. Why: (1) we need linearizable single-row UPDATEs for monotonic max_id; (2) we need the same-txn outbox for audit-event durability; (3) Postgres semi-sync replication gives RPO 0 cross-AZ at acceptable cost. Considered MySQL with native auto_increment_increment=2 (the original Flickr design), rejected because Postgres' WAL-based logical replication is cleaner for the cross-region DR and the pg_export_snapshot() story is better for the audit job.
etcd 3.5 (5-node, 3-AZ) for worker-id leasing. ZooKeeper would also work (Twitter's original choice) but etcd's gRPC + HTTP/2 is friendlier for our SPIFFE story and the operational tooling in the platform is etcd-native. We do NOT use Consul — the multi-DC story is irrelevant for ID-leasing (each region owns its slot space) and the additional features add operational surface we don't need.
S3 for the audit archive — uncontroversial, the standard pattern for compliance-grade append-only logs.
07High-level design
Which components handle a request, and in what order?
Architecture summary — three lanes plus the audit substrate:
- In-process UUIDv7 (default for ~85% of callers). The SDK draws a CSPRNG sample + reads
clock_gettime(CLOCK_REALTIME)to assemble a 128-bit RFC 9562 v7 UUID — never leaves the process. Per-backend monotonicity (brandur's pattern) prevents same-ms collisions on the same connection. This lane has no platform components on the canonical because there are none — it's a library. - Snowflake RPC (~10% of callers). Caller SDK → Gateway → Snowflake service → (lease-cache verify, mostly cache hit). Each Snowflake pod holds an etcd-leased worker-id and mints IDs locally — no per-call coordination. Halt-on-rewind via
CLOCK_MONOTONICfor sequence delta. Audit event published async to Kafka per lease grant/renewal. - Segment allocator (~5% of callers). Caller SDK → Gateway → Segment Allocator → (in-memory counter bump or DB UPDATE for next chunk). Outbox-coupled with the audit event so segment + audit are atomic. Dual-master Postgres for AZ-loss survival without write outage.
- Audit substrate. Every issuance flows through Kafka → (Audit Sink for long-term archival, Uniqueness Audit Worker for dup detection). The uniqueness probe is the single most load-bearing piece of observability in the system.
Why three lanes coexist. Production teams that scale beyond ~10K writes/sec inevitably need all three: UUIDv7 for the new services, Snowflake RPC for the legacy bigint schemas, segment for the human-facing IDs. Forcing one strategy on every caller creates worse problems than running three (forced UUIDv7 → BIGINT FK migrations stall for years; forced Snowflake → every write takes a network hop; forced segment → human-facing IDs are exposed to scraping). Discord and Meituan both run multiple lanes side-by-side for exactly these reasons.
Why the lease cache (Redis) sits between Snowflake and etcd. Verifying the lease on every Nth ID call (default every 1024) is the correctness-load-bearing check — without it, an expired lease could continue minting IDs whose worker-id slot has already been re-issued. The Redis lease-cache absorbs that check at sub-ms latency; without it, every 1024th ID would take an etcd round-trip (5-20 ms). The cache is shared across the Snowflake fleet so a Lease Manager's emergency revoke write propagates platform-wide on the next periodic check.
Why halt-on-rewind, not 'wait it out'. The Cloudflare 2017 leap-second outage is the textbook case: Go's time.Now() returned a value less than a prior reading; rand.Int63n got a negative argument and panicked. The naive 'wait it out' Snowflake variants (issue an ID with the same timestamp as the last) preserve liveness at the cost of correctness — an ID issued just before the rewind, then issued again with the same (worker_id, ts, seq), is a silent dup. Halt-on-rewind trades availability (one pod refuses to mint for the rewind window) for guaranteed uniqueness — which is what the platform contract requires. The 5 ms threshold is the tolerance for normal NTP corrections; anything bigger is a real anomaly worth blocking on.
Why the dual-master segment DB instead of single primary + automatic failover. Failover under a write spike is when segment-allocator dups historically happen — the binlog-lag window between primary loss and replica promotion is exactly when the promoted replica issues IDs the old primary already issued (Jepsen MySQL 8.0.34 cautioned that semi-sync isn't even guaranteed to be the freshest binlog). Dual-master with Flickr odd/even — applied to the segment index (max_id/step), each master advancing by 2 × step so A owns even step-blocks and B odd, disjoint — sidesteps the failover question: both masters are always live, and a single master loss just halves capacity rather than triggering a promotion gap. The trade-off is that monotonicity is no longer strict across the two masters (a faster master's later block can be handed out before a slower master's earlier block) — accepted because the consumers of segment IDs (invoice numbers etc.) tolerate small ordering inversions.
Why outbox, not 2PC, on segment + audit. The segment UPDATE and the audit-event INSERT are written in a single Postgres transaction (both in the same database — no cross-store coordination needed). A separate worker tails the audit_event table and publishes to Kafka; on Kafka publish success it marks the row published=true (retention is by day-partition drop, so rows remain available to the nightly uniqueness scan). This gives at-least-once delivery of audit events with zero blocking on the issuance hot path. 2PC across Postgres + Kafka would add ~10 ms of coordination per ID range allocation and introduce the classic 2PC-coordinator-down liveness failure (covered in deep dive #5). Saga is wrong here because there is no compensating action for "audit event already fired" — the outbox at-least-once + idempotent consumer (de-dup on event_id) is the correct pattern.
Why the Uniqueness Audit Worker is its own component. Every other failure mode in the system either fails loudly (RPC error, lease-renewal failure) or fails into the audit log. Worker-id collision and clock rewind into the same ms can produce silent dup IDs that don't surface until a downstream consumer hits a primary-key violation hours later. The streaming probe (per-biz_tag interval index of allocated ranges + live worker-id leases) catches overlaps near-real-time; the nightly scan (segment range-overlap on the Postgres DR replica + worker-id lease-overlap on the audit-sink Parquet) is defense-in-depth against any service that skipped publishing to the audit stream. This is the single piece of observability that makes "duplicate-ID rate = 0" testable rather than aspirational.
Why we model NTP / Chrony as an external dependency on the diagram. Bad NTP is the documented #1 root cause of distributed-ID dup incidents (Cloudflare 2017, plus countless internal post-mortems). Drawing it as an explicit edge Snowflake → NTP/Chrony makes the dependency visible in failure-mode reviews — when on-call sees the canonical, they understand that an NTP outage is an ID-platform incident, not just a host-os incident. The Meta migration paper from ntpd → chrony for the same reason (10 ms → 100 µs precision) is the operational baseline.
Multi-region posture. This canonical is single-region (active-active across 3 AZs) with a cross-region async DR replica for the segment DB only. Snowflake leases are region-scoped (DC bits in the etcd key prevent cross-region collision). UUIDv7 callers are inherently multi-region — they hold no platform state. The cross-region DR replica is for catastrophic regional loss only; on a normal AZ loss, dual-master + multi-AZ etcd handles it without operator action. Going to active-active multi-region for ID issuance would require partitioning the worker-id space by DC (Discord 5/5) and per-region etcd clusters — defer until SLO violations or regulatory data-residency forces it.
Primary sources
- Twitter Engineering — Announcing Snowflake (2010)
- Discord — How Discord Stores Trillions of Messages (snowflake layout)
- Instagram Engineering — Sharding & IDs at Instagram (PG PL/pgSQL next_id)
- Flickr Code — Ticket Servers: Distributed Unique Primary Keys on the Cheap
- Meituan Tech — Leaf: open source ID-gen (segment + snowflake, 双 buffer)
- RFC 9562 — Universally Unique IDentifiers (incl. UUIDv7)
- Sony — Sonyflake (39-bit 10ms-tick + 16-bit machine-id)
- Cloudflare — How and why the leap second affected Cloudflare DNS (2017)
- Meta Engineering — NTP service migration (chrony, 100µs precision)
- Jepsen — MySQL 8.0.34 (semi-sync replication and binlog freshness)
- PlanetScale — MySQL semi-sync: durability, consistency, split-brains
- Shopify Engineering — Building Resilient Payment Systems (ULID vs UUIDv4)
- Stripe Blog — Designing robust and predictable APIs with idempotency
- Google SRE Workbook — Ch. 22 Addressing Cascading Failures
- AWS Builders' Library — Timeouts, retries, and backoff with jitter
Now defend it
Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.
Work Distributed Unique ID Generator yourselfMore in System Design Fundamentals
The four primitives every later problem assumes — unique IDs, rate limits, caching a read-heavy endpoint, and making a retry safe.
- URL ShortenerShorten a long URL. Read-heavy. Don't collide.
- PastebinStore text/code blobs with TTL and access control.
- Distributed Rate LimiterEnforce a per-key request limit across a fleet of enforcers — accurately, in under a millisecond, without becoming the outage.
- Submit Order (Prevent Double-Charge)Idempotency keys, dedup window, retry storms.