Pastebin
Worked solution

Pastebin — a worked solution

Store text/code blobs with TTL and access control.

Try it yourself first.

You will remember almost none of this if you read it cold. The workspace walks the same 10 stages and runs the architecture you draw through a simulator, so you find out where your version breaks before you see ours.

Open the Pastebin workspace

The problem

Build a Pastebin-class service: users submit a text / code blob (a "paste"), receive a short URL, share it. Anyone with the URL can read the blob until it expires or is taken down. Visibility ranges from public (search-engine-indexable, edge-cacheable) through unlisted (URL is the secret) to private (owner-only, JWT-authed) and password-protected (server-side check). The system spans an edge CDN, a metadata DB, an object store for bytes, an async abuse pipeline, and a TTL sweeper.

The interview-tier 6-node sketch is wrong here in three load-bearing ways: (1) it under-models the cache-key story for private content; (2) it treats S3 lifecycle as the source of truth for TTL; (3) it puts blob storage on the same Postgres as metadata. We unwind all three.

The reference architecture

Reference architecture for Pastebin: 20 components — Client, CDN · Public Edge, API Gateway, KV Store · Rate Limiter, Auth / Identity, Service · Read, Service · Write, Cache · Hot Pastes, SQL DB · Metadata (sharded), Object Store · Blobs, Stream · Paste Events, Worker · Abuse Scanner, Scheduler / Cron · TTL Sweeper, External · Abuse / Phish API, Coordinator, Worker · CDC Outbox Relay, Worker · Orphan Blob GC, Queue · Manual Purge DLQ, Log Shipper, Monitoring — connected by 35 flows.ClientclientCDN · Public EdgeCloudflare (Tiered Cach…API GatewayEnvoy + Cloudflare WAFKV Store · Rate Limit…Redis 7 cluster (sharde…Auth / IdentityEnvoy ext_authz + JWKS-…Service · ReadGo 1.22, fasthttpService · WriteGo 1.22, fasthttpCache · Hot PastesRedis 7 Cluster (16 sha…SQL DB · Metadata (sh…Postgres 16 + Citus (8 …Object Store · BlobsS3 (Standard → IA@30d →…Stream · Paste EventsKafka 3.7 (RF=3, acks=a…Worker · Abuse ScannerPython + Faust consumer…Scheduler / Cron · TT…Kubernetes CronJob + wo…External · Abuse / Ph…VirusTotal + Google Saf…Coordinatoretcd 3.5 (3-node Raft)Worker · CDC Outbox R…Debezium / Postgres log…Worker · Orphan Blob …Python + boto3 + pagina…Queue · Manual Purge …Kafka (separate cluster…Log ShipperVector (DaemonSet) → S3…MonitoringPrometheus + Alertmanag…
20 components, 35 flows. A dashed line is an asynchronous hop. This is the reference design, not the only one that works.

What each component is for

Client

Issues GET /:paste_id for reads (browser tab, curl pipe, bot), POST /api/v1/pastes for writes, and (rarely) DELETE for owner-initiated takedown. Handles paste TTL UX (relative-time display, '410 Gone' rendering for expired pastes). On private/password pastes, presents the bearer token / paste-key cookie that the server checks server-side before returning content.

Why it exists. All entry points are clients — there's no other ingress. The interesting design choice is the SPLIT: public reads go to cdn.paste.example (cacheable), private/password reads + writes go to api.paste.example (uncached, auth required). Sharing the same hostname with Vary: Authorization was rejected because it destroys edge hit ratio on public pastes (every signed-in viewer = unique cache entry).

When it fails. Client-side caches that ignore Cache-Control: private (older mobile WebViews) can leak private pastes via shared device profiles; mitigation is Cache-Control: private, no-store plus Pragma: no-cache belt-and-braces, AND routing private content through a separate hostname so even a misbehaving cache can't serve it cross-tenant.

CDN · Public EdgeCloudflare (Tiered Cache + Argo)

Serves cached GET /:paste_id responses for public pastes. Two-tier: edge POPs (~300 worldwide) ↦ regional shield (~12 datacenters) ↦ origin. The shield collapses concurrent misses for the same paste_id into a single origin fetch — single-flight built into the CDN. Strips Cookie and Authorization headers before computing the cache key so all anonymous requesters share entries. Receives async cache-purge calls from the abuse worker and the TTL sweeper when a paste flips to disabled / expired.

Why it exists. 90%+ of read QPS is public pastes that fit a Zipfian distribution — the CDN absorbs that read fan-out so origin only sees ~3,000 reads/sec at peak instead of 28,000. Tiered cache (over basic edge cache) is non-negotiable: at 50k req/s on one viral paste, ~300 edge POPs each independently fetching origin would crush the read service. Tiered cache collapses that to one shield-region fetch. We considered routing public pastes through Cloudflare R2 + Workers (origin-less) but kept S3 origin to retain control over TTL semantics and lifecycle.

When it fails. Cloudflare 2025-06-12 KV-backend outage (third-party storage failed for 2h28m, 91% KV error rate) is the load-bearing precedent — when the CDN control plane degrades, edge nodes serve stale-but-warm objects until their last-good cache expires, and origin gets hit by every miss. Our mitigation: extended stale-while-revalidate (24h) on public pastes so edge-served stale wins over origin-overload during a CDN brownout, and a shed-rule that returns 503 Retry-After: 30 for non-cached writes if the gateway can't reach origin in 2s.

API GatewayEnvoy + Cloudflare WAF

Terminates TLS for api.paste.example and origin-pull from CDN. Runs WAF rules (OWASP CRS + custom paste-spam patterns). Normalizes headers (strips X-Forwarded-Authorization that some clients leak). Routes POST /api/v1/pastes → write service, GET /api/v1/pastes/:id → read service, DELETE /api/v1/pastes/:id (owner takedown) → write service, DELETE /api/v1/admin/... → admin path. Calls Auth + Rate Limiter inline (ext_authz) before forwarding upstream — the upstream services trust the gateway's X-User-Id header and don't re-verify.

Why it exists. Single chokepoint for security policy (WAF, mTLS upstream, header normalization) and cost control (rate-limit before we even hit the read service). We considered making each service do its own WAF + auth — rejected because (a) WAF rule deployments would need N rollouts; (b) you can't trust a service that hasn't authed the caller; (c) rate-limit decisions are global, not per-service. Envoy was picked over an in-house gateway because we want the same thing handling north-south + east-west via a single control plane.

When it fails. Gateway pool exhaustion on a viral paste — when origin slows down (e.g. cache miss to DB takes 100ms instead of 30ms), each gateway instance's outbound connection pool fills up and requests queue, eventually 503ing. Mitigation: outbound pool sized 4× steady-state demand; per-route circuit breaker; Retry-After: 5 on shed responses so clients don't retry-storm; Envoy's overload manager sheds at 90% memory rather than crashing.

KV Store · Rate LimiterRedis 7 cluster (sharded, AOF every 1s)

Atomic INCR + EXPIRE via Lua script per (ip, route, window). Three buckets: per-IP (300 reads/min anonymous), per-user (5,000 writes/day authenticated), per-paste-id (10 password attempts/hour to defeat timing attacks on protected pastes). Returns (allowed: bool, remaining: int, retry_after_s: int) in <1ms.

Why it exists. Pastebin-class abuse is overwhelmingly automated: scraper bots (post-2020 Pastebin paid-API kill prompted aggressive scraping), credential stuffing on login, dictionary attacks on password-protected paste IDs. Rate-limiting at the gateway BEFORE we hit the read service stops the cheap attacks at line rate. We considered in-memory token-bucket per-gateway-instance — rejected because horizontal-scaling means each replica sees only 1/6 of an attacker's traffic, and a 1k-RPS attacker would slip under per-instance limits while crushing globally.

When it fails. Failover gap = 15s. Mitigation has two halves: on the read-path the gateway fails OPEN (we accept brief abuse over dropping legitimate paste reads); on the write-path the gateway falls back to a per-replica in-memory token bucket — but ONLY because the front-door Envoy LB pins clients via consistent-hash on client_ip so an attacker lands on one replica (without ip-pinning the bucket-per-replica is broken: an attacker fan-outs across all 6 replicas and gets 6× the intended limit). Per-replica budget = global ÷ replicas ÷ 4 (= 12.5/min for the 300/min anonymous limit) — quarter-of-fair-share absorbs short-tail variance. Post-failover we accept 5s of RPO on counter state; a sustained attack during the gap pages on-call within 30s on origin RPS spike.

Auth / IdentityEnvoy ext_authz + JWKS-cached JWT verifier

Stateless JWT verifier called inline by the gateway via Envoy ext_authz. Pulls JWKS once on startup + every 10 min, caches in-process. Verifies signature, expiry, and issuer; extracts sub (user id) and scope claims; returns them to the gateway as headers. For anonymous reads, returns principal=anon without lookup. Does NOT do the paste-level ACL check — that lives in the read service against the metadata row, because access depends on paste visibility (public/unlisted/password/private) which only metadata knows.

Why it exists. Centralized JWT verification so each upstream service trusts a normalized X-User-Id rather than re-implementing JOSE crypto. We considered embedding verification in each service; rejected — JWKS rotation lag, library divergence, and the fact that misconfiguring iss / aud validation is a known CVE class. Auth is a separate process so a misbehaving JOSE library can't OOM the gateway.

When it fails. Auth pool exhaustion — if JWKS endpoint slows down on the 10-min refresh, in-flight verifications can starve. Mitigation: stale-JWKS-OK for up to 1 hour after the 10-min refresh fails; circuit breaker on the JWKS pull; reads degrade to anon-only mode and we lose private-paste reads but public reads keep flowing. Alarm threshold: any minute where >0.1% of writes are rejected because Auth said unverified.

Service · ReadGo 1.22, fasthttp

Receives GET /api/v1/pastes/:id from gateway. Looks up metadata in Hot Cache, falls through to Metadata DB replica on miss. Enforces paste-level ACL: public → serve; unlisted → serve (URL is the secret); password → constant-time argon2id verify, 401 on mismatch; private → 403 if X-User-Id ≠ owner; expired/disabled → 410. For burn-after-read: atomic CAS UPDATE pastes SET burn_consumed=true WHERE paste_id=$1 AND burn_consumed=false RETURNING blob_id BEFORE blob fetch — the CAS winner is the only reader that gets blob_id. Loser sees 410. CAS-then-fetch ordering means a winner whose S3 GET fails sees 410 on retry; we accept this trade for view-once correctness over availability (legitimate retry of view-once is rare and the contract is explicit). For small blobs (<10KB) returns inline. Larger blobs split by cacheability: public/unlisted responses proxy-stream the bytes so the CDN caches the body itself (a cached redirect would outlive its 30s signature under the 60s edge TTL and serve SignatureExpired 403s, and a viral paste must be absorbed at the edge, not by S3); private/password reads on the uncached api hostname get a 30-second presigned S3 GET redirect so we don't pay proxy egress on multi-MB pastes. Cross-region S3 fallback proxy-streams through the read service from a secondary-region replica.

Why it exists. Stateless serving tier exists separately from Write because their failure modes diverge: a buggy write service deploy can be rolled back without touching reads; reads can scale 10× independently for a viral burst. Mashing them into one service means autoscaler signals collide (read p99 is dominated by cache hit ratio, write p99 by DB primary CPU). We considered pushing the ACL check into the cache layer — rejected because cache is shared across visibilities and ACL evaluation depends on per-request principal.

When it fails. When Hot Cache is cold (post-restart, cluster failover) every read falls through to the DB replica — DB capacity is sized for 10% miss rate, not 100%. Mitigation: in-process LRU; single-flight in the read service (10 concurrent misses for the same paste_id collapse to 1 DB query); shed to 503 with Retry-After: 1 if DB pool fills. Pages on-call when read-service in-flight queue depth p95 > 50.

Service · WriteGo 1.22, fasthttp

Receives POST /api/v1/pastes (and DELETE /api/v1/pastes/:id for owner takedown — soft-delete + paste.disabled outbox row in one txn). Generates paste_id — a 12-char base62 token from a CSPRNG (~71 bits of entropy). This is load-bearing for security: unlisted pastes use the URL itself as the secret, so the id must be unguessable. A Snowflake/ULID-style id (timestamp + worker + monotonic sequence) is *enumerable* — an attacker who knows roughly when a paste was created could walk the sequence — so it must never be the secret. Hash-prefixes the blob key ({rand4}/{paste_id}) so writes spread across 65k S3 prefixes. Two create shapes by size: (a) SMALL text (≤256KB) — the body rides the POST, the service PUTs it to S3 with Content-MD5 + SSE-KMS itself, and commits the row as upload_state=committed; (b) LARGE / file — no body on the POST; the service mints a short-TTL presigned PUT (size-bound by the declared contentSize), commits the row as upload_state=pending, and returns the uploadUrl so the CLIENT streams bytes straight to the object store (the app tier never touches the megabytes). Either way, inside one DB transaction: INSERT metadata row + INSERT outbox row; populate Hot Cache write-through (only when committed). The CDC relay publishes the outbox row to Kafka asynchronously — atomicity stays in the DB; Kafka publish is at-least-once outside the txn.

Why it exists. Writes need stronger guarantees (durability, idempotency, abuse hooks) than reads, and a separate service lets us deploy rate-limited rollouts independently. Most importantly: the outbox pattern lives here. We considered 2PC across DB and Kafka — rejected because (a) Kafka is not a true 2PC participant; (b) coordinator failures stall both systems (see chaos pastebin:abuse-scan-third-party-out); (c) outbox gives us at-least-once into the abuse pipeline with simpler liveness.

When it fails. S3 PUT throttle (3,500/sec/prefix) is the canonical case — when a script-kiddie uploads 5k pastes/sec all to the same default prefix, S3 returns 503 SlowDown. We hash-prefix to defuse it, but a deployment that regressed prefix entropy would page within 60s on s3_503_slowdown_total. Mitigation: prefix-entropy assertion in CI; exponential backoff on the writer; circuit-break to 503 to client after 3 retries (better to fail fast than queue forever).

Cache · Hot PastesRedis 7 Cluster (16 shards, RF=2)

Stores paste_id → {visibility, owner_id, password_hash, blob_id, blob_size, blob_etag, expires_at, created_at, abuse_disabled} for the hot working set (~20 GB in RAM across 16 shards — the 300s entry TTL keeps only the recently-touched Zipf-hot slice resident, not the full 21 TB in-TTL blob corpus). Small public blobs (<10KB) cached inline as a separate blob:{paste_id} key. Write-through: write service populates on insert; abuse worker + TTL sweeper invalidate on disable/expire. TTL on entries = 300s (the lookup result, not the paste TTL) so a paste edited in one region eventually replicates via cache miss.

Why it exists. Origin DB peak is 1.4k creates/s + 3k cache-miss reads/s = 4.4k qps. Without cache, miss rate is 100% and DB takes 28k reads/s — 6× the budget. Cache absorbs the read fan-out for private/password pastes that can't sit at the CDN edge, plus public/unlisted misses that fall through the edge. We considered cache-aside (read-through) — rejected because write-through gives us read-after-write for the creator within the same region (see deep-dive read-your-write style failure mode).

When it fails. Cluster cold restart (deploy / rolling upgrade with TTL > restart window): every paste_id is a miss, DB stampedes. Mitigation: in-process LRU on the read service catches the first 60s; single-flight collapses concurrent misses to 1 DB query per key; pre-warm script reads the top-1k pastes from the DB into cache during deploy. Page on-call if cache hit ratio < 80% for >2 minutes.

SQL DB · Metadata (sharded)Postgres 16 + Citus (8 shards, sync repl)

Authoritative store for paste metadata: (paste_id PK, owner_id, blob_id, blob_size, visibility, password_hash, expires_at, created_at, soft_deleted, abuse_disabled, version). Plus the outbox(id, event_type, payload, created_at, published_at) table — outbox txn pattern for the abuse + TTL fanout. Sharded by hash(paste_id) mod 8 via Citus. Each shard has 1 sync follower (for RPO=0 failover) + 1 async follower (for read scale). Partial index on expires_at WHERE expires_at IS NOT NULL for the TTL sweeper. ACID guarantees: writes are linearizable per-shard; the outbox row commits in the same transaction as the metadata insert, so async fanout is exactly once at the source.

Why it exists. Source of truth for ACL, expiry, and abuse state. Cache is fast but ephemeral; S3 stores bytes but doesn't have the relational story for (owner, expires_at, visibility) queries we need on read + sweep. We considered DynamoDB — rejected because the 'list pastes by owner_id ordered by created_at' query is awkward without a GSI we'd then have to maintain consistency with, and Citus + Postgres gives us the SQL flexibility for incident-time ad-hoc queries (see deep-dive abuse-takedown). 8 shards sized for 5y growth (~50TB row volume).

When it fails. Primary loss with sync replication = 30s RTO write outage; reads survive via the async follower (eventual). The async-follower-only-mode is detected by the gateway's circuit breaker (DB writes 5xx for 30s), and writes get 503 Retry-After: 30 while reads keep flowing. We chose sync replication after observing the GitLab 2017-01-31 incident class — async replication's RPO>0 and the 'rm -rf primary' pattern is a deal breaker for paste durability.

Object Store · BlobsS3 (Standard → IA@30d → Glacier@180d via lifecycle)

Stores the paste body as {rand4}/{paste_id} objects. rand4 is the first 4 base62 chars of sha256(paste_id) — gives 65,536 prefixes so writes hit S3's 3,500 PUT/sec/prefix limit only if a single prefix sees >3,500/sec (would require >220M creates/sec across the fleet, never). Versioning enabled — soft-delete on tombstone leaves the prior version recoverable for 30 days, which we lean on for ops mistakes (see chaos pastebin:ttl-sweeper-shard-wipe). SSE-KMS encryption with one CMK per data classification (public / private / password). Lifecycle: Standard 30d → Standard-IA → Glacier Deep Archive 180d → expiry per metadata-DB TTL.

Why it exists. Postgres is wrong for 50KB-1MB blobs at 30M/day — TOAST overhead, WAL bloat, replication-lag amplification. Object store is the right primitive: bytes-by-key, 11×9s durability, lifecycle handles cold-tier costs. We considered storing small blobs (<4KB) inline in Postgres for fewer round trips — rejected because the 4KB threshold is fragile and the bimodal storage path complicates the read service. Better to put EVERYTHING in S3 and put 10KB-or-less in Redis when hot.

When it fails. us-east-1 S3 outage (the 2017-02-28 incident — typo nuked the index subsystem) is the canonical infra-class failure. Read path: presigned URLs fail; we fall back to proxy-streaming through the read service from a cross-region S3 bucket (replicated via S3 Cross-Region Replication, RPO ~15min). Write path: writers retry to a secondary region and reconcile the metadata blob_region field. Page on-call when origin S3 5xx rate >0.5% sustained for 60s.

Stream · Paste EventsKafka 3.7 (RF=3, acks=all, ISR=2)

Kafka topic paste.events carrying the outbox-relayed events from the metadata DB: paste.created (drives abuse scan), paste.disabled (drives cache + CDN purge), paste.expired (drives cache purge). Consumed by the abuse worker and a small fleet of housekeeping consumers. 64 partitions keyed by paste_id so abuse scan is per-paste-ordered. Retention 7 days for replay debugging. The outbox relay (Debezium-style sidecar tailing the metadata DB's logical replication stream) is what publishes here — the write service NEVER calls Kafka directly, so an at-least-once Kafka outage doesn't fail user writes.

Why it exists. Decouples user-facing writes from the abuse / cache-purge / CDN-purge fan-out. If we did all of that synchronously on POST, p99 write latency would be dominated by the slowest external scan API. Outbox txn (atomically commit the row + the event in the same DB txn) gives us at-least-once semantics for fan-out without 2PC. We considered eventing direct from the write service — rejected because a Kafka brownout would block writes, and consumer at-least-once + idempotent handlers is simpler than 2PC liveness.

When it fails. Consumer falls behind (50× lag spike from a poison message — see Cloudflare 2019-07-02 regex-CPU class) means new pastes go un-scanned for hours while malware payloads are read. Mitigation: poison-message DLQ with dead-letter queue and manual ops review; per-message timeout 30s; consumer parallelism 4× per partition so one slow record doesn't block all. Page on-call when consumer-group lag > 10k records or > 5 minutes.

Worker · Abuse ScannerPython + Faust consumer pool

Consumes paste.created events from Kafka. For each new paste: (1) hash the body and check a local Bloom filter of known-bad SHA-256s (DMCA list, malware IOC list) — a miss is definitive (no false negatives), so ~99% of pastes clear sub-ms; a hit is only probabilistic and is confirmed against the exact hash store before any action, because a Bloom false positive must never disable a legitimate paste; (2) submit heuristic-flagged suspicious bodies to the External Abuse API for deeper scan; (3) on a confirmed verdict, UPDATE metadata.abuse_disabled=true, INVALIDATE Hot Cache, PURGE CDN edge cache for the paste's URL. On paste.disabled events (owner or admin takedown), runs the same cache+CDN purge. Maintains its own local-state offset so a restart resumes mid-stream.

Why it exists. Pastebin-class services live or die by their abuse posture — Pastebin.com's 2014-2020 era is the canonical case study (Bladabindi RAT, Emotet IOC dumps via base64, the 2020 scraping-API kill in response). We built this to run async, never blocking the write path, but inside the SAME outbox event stream as the user-write so we can never miss a paste. We considered inline scan on POST — rejected because (a) external scan APIs have p99 of 2-5s; (b) a scan-API outage would block all creates; (c) most abuse vectors are server-detectable (hash match) without network calls.

When it fails. External API outage = backlog grows; bad pastes serve until the queue drains. Mitigation: Bloom filter catches known bad without external calls; explicit 'fail-open' policy with audit tag so ops can re-scan post-recovery; if backlog > 10k, autoscale workers + raise the 'unscanned-window' SLO alert. The really bad case is CDN purge failing silently — that needs the per-paste purge confirmation we wired in.

Scheduler / Cron · TTL SweeperKubernetes CronJob + worker pool

Runs every 5 minutes. Each replica owns 2 of the 8 metadata shards (range-partitioned). Per-shard: SELECT paste_id, blob_id FROM pastes WHERE expires_at < now() AND NOT soft_deleted ORDER BY expires_at LIMIT 10000 (uses the partial index). For each row: UPDATE soft_deleted=true (the read path now 410s); insert outbox row paste.expired. Cap per-run at 10k rows per shard so a sudden expiry wave doesn't take 30 minutes. S3 lifecycle handles the actual byte deletion (rounded to UTC midnight, 24h SLA).

Why it exists. Lazy expiry on read alone is not enough at 30M/day churn — old rows accumulate in the index and the partial index gets bloated. Active sweeper keeps the working set tight. We considered a TRIGGER ON SELECT lazy delete — rejected because (a) it amplifies hot-key reads with writes; (b) cold abandoned pastes never get hit and would never be cleaned. We considered S3 lifecycle as the only sweeper — rejected because S3 lifecycle's 24h+UTC-midnight semantics make the API contract 'expires in 1 hour' a lie.

When it fails. Sweeper bug deletes too many rows (the GitLab 2017-01-31 class incident). Mitigation: per-run cap (10k); dry-run mode in canary; hourly anomaly alert if delete rate > 3× baseline; soft-delete + 7-day grace + S3 versioning = 3-layer recovery. The opposite failure mode (sweeper falls behind) is detected by expired_unswept_count going up — page on-call if > 100k rows are past expiry but not soft-deleted.

External · Abuse / Phish APIVirusTotal + Google Safe Browsing + Cloudflare Phish

External APIs the abuse worker consults for unknown-hash pastes. VirusTotal for malware byte-pattern signatures, Google Safe Browsing for URL-in-content phishing checks, Cloudflare Phish for emerging campaigns. Rate-limited (VirusTotal Public: 4 req/min; Premium: 1k/day; Safe Browsing: 10k/day per project). Synchronous calls with 5s timeout per provider; we OR the verdicts (any ‘malicious’ → disable).

Why it exists. We can't build our own malware DB at this scale — the cited research (Pastebin malware abuse 2014-2020 era, Bladabindi RAT) is exactly the threat the providers are paid to track. Inline-on-write was rejected (see Abuse Worker whyItExists). We bought N providers rather than 1 because (a) provider-level outages are not rare (Heroku 2022 cert/DNS class); (b) malware variants ride one provider's blind spot for hours; (c) cost is bounded by the ~1% heuristic-flagged suspicious path.

When it fails. All three providers degrade simultaneously (rare but happened during widespread DNS issues). Bloom filter catches known-bad; new bad pastes enter the 'manual-review' queue and the unscanned window SLO pages on-call within 5 min. Trade-off: we accept a 5-15 min vulnerability window during a third-party outage rather than blocking writes (which would be a self-inflicted DoS).

Coordinatoretcd 3.5 (3-node Raft)

3-node etcd cluster providing distributed lock + leader-election primitives. Patroni uses it to maintain Postgres primary identity and gate failover (only one primary at a time per shard); the TTL Sweeper uses it for per-shard ownership leases (only one of 4 sweeper replicas owns a shard's expire-sweep at a time). The CDC Outbox Relay uses it for active-passive election so two relays don't double-publish from the same logical replication slot. Stores ~50 keys total: per-shard primary identity, per-shard sweep-owner, per-relay leadership.

Why it exists. metadb.failover.mode=automatic with rtoSeconds=30 requires a coordinated failover primitive — without quorum on 'who is the primary now', a partition produces split-brain. The TTL Sweeper's 'leader election only one runs per shard at a time' and the CDC Relay's 'active-passive deployment' both depend on the same primitive. We considered using Postgres advisory locks for sweeper coord — rejected because it would couple sweeper liveness to DB liveness AND because Patroni already needs etcd, so we'd be running two coordination systems. Single etcd cluster reused across all three concerns.

When it fails. Quorum loss (2 of 3 nodes down) → no new failover decisions can be made; existing leases continue until expiry, then the sweeper / relay halt. Mitigation: cross-AZ deployment so a single AZ outage leaves 2 nodes alive; alert on quorum-loss within 30s; auto-restart per-node on disk pressure (etcd's WAL is small but a full disk wedges it). The really bad case is a network partition between etcd nodes — Raft handles it, but the minority side wedges; we rely on Raft's correctness rather than trying to second-guess it.

Worker · CDC Outbox RelayDebezium / Postgres logical replication

Active-passive Debezium-style sidecar tailing the metadb's logical replication slot. Streams the WAL of the outbox table; for each new row, publishes a Kafka message to paste.events with key=paste_id and the same payload. UPDATEs outbox.published_at after Kafka acknowledges. Uses etcd leader-election so only one replica reads the slot — duplicate readers would replay rows. On restart, resumes from the last published offset stored in etcd. At-least-once semantics into Kafka; downstream consumers (abuse worker) are idempotent.

Why it exists. The outbox pattern decouples user-facing writes from Kafka availability — write service commits paste row + outbox row in the same DB txn, then the relay asynchronously publishes. We considered making the writer call Kafka directly — rejected because (a) Kafka outage would block writes; (b) atomic 'INSERT row + Kafka send' requires 2PC which has the liveness failure we explicitly avoid (see chaos pastebin:abuse-scan-third-party-out's commentary). Without a dedicated relay node, the metadb→outbox edge would be a fiction (the DB doesn't call Kafka). We considered embedding the relay in the writesvc — rejected because relay must be active-passive across the whole cluster, not per-replica.

When it fails. Relay falls behind → Postgres logical replication slot accumulates WAL → primary's pg_wal partition fills → writes block (the Postgres-OOMs-from-disk-pressure class). Mitigation: alert on slot lag > 5min (page-on-call well before pg_wal fill); auto-restart relay on lag spike; relay state is stored in etcd so a fresh process resumes cleanly. The really bad case is a poison row Debezium can't decode — we deploy with errors.tolerance=all + dead-letter-topic for unparseable rows so a single bad row doesn't wedge the relay.

Worker · Orphan Blob GCPython + boto3 + paginated S3 LIST

Daily cron job (active-passive via etcd). Reconciles S3 against metadb in BOTH directions: (1) lists S3 prefixes by hash-shard and tombstones any blob >24h old with no referencing row (the small-text writer-crash orphan); (2) scans the upload_state='pending' partial index for reservations whose presigned PUT never completed past the upload TTL — flips them to soft_deleted and tombstones any partial bytes. Non-referenced blobs move to a tombstone/ prefix for 7 days, then drop.

Why it exists. Two create shapes create two orphan classes. SMALL text: service PUTs the blob THEN commits the row in one txn — a crash after PUT but before commit leaves an S3 orphan with no row (~0.1% of 1,400 creates/sec ≈ 120k/day). LARGE / file: the row is committed FIRST as pending and the client uploads via presigned PUT — a client that abandons the upload leaves a pending row with no bytes. We make reservation-then-fulfill safe (rather than reject it) with the upload_state flag: a read of a pending paste returns 409 Conflict (with Retry-After), NOT 502, and this GC reaps abandoned reservations off the partial index. S3 lifecycle won't catch either class — they look like normal abandoned objects.

When it fails. GC bug deletes a NOT-orphan (race with a slow metadata commit, e.g., a cross-shard transaction that took >24h). Mitigation: 24h grace window + soft-delete + 7-day tombstone = 3-layer recovery. Alert on orphan_delete_rate > 5× baseline. The opposite failure (GC falls behind, orphans accumulate) shows up as S3 cost growth — caught by monthly cost-anomaly review.

Queue · Manual Purge DLQKafka (separate cluster from outbox)

Kafka topic holding events from the abuse worker when an automated CDN purge failed after 3 retries. Each entry contains paste_id, target POPs, error code, retry-history. PagerDuty integration fires on any new entry. Runbook for on-call: re-attempt purge via Cloudflare emergency API; verify with a synthetic GET against several POPs; ack the message only after confirmation. Retains 30 days for forensics. Steady state is empty; one entry/day is YELLOW; ten/day is an incident.

Why it exists. The abuse worker's failureMode says missed CDN purge is a regulatory issue, not a silent drop. We need a destination for failed purges that pages on-call rather than logging into the void. We considered routing failed purges back into the main paste.events topic — rejected because mixing user-data fan-out with ops-only events bloats consumer-group complexity and confuses the abuse SLO. Separate cluster from outbox so a Kafka outage on outbox doesn't take this DLQ down with it.

When it fails. DLQ piles up faster than human ops can clear (e.g., during a malware-spam wave with thousands of disables). Mitigation: backstop is the CDN cache TTL (60s) — even a missed manual purge converges within 60s naturally, so the worst window is the edge-TTL we already accept. The DLQ is more about audit trail than functional recovery.

Log ShipperVector (DaemonSet) → S3 + ClickHouse

Sidecar Vector instance per host. Tails service logs (gateway access logs, abuse worker DMCA decisions, sweeper soft-delete events, write service errors) and ships to S3 (raw JSONL) + ClickHouse (indexed audit table). WAF events, abuse-decision events, and sweeper-delete events get tagged with priority HIGH and retained 7 years for compliance. INFO-level access logs retained 30 days.

Why it exists. Pastebin-class regulatory ask is 'show me what happened to paste X between disable and CDN purge': required for DMCA compliance, GDPR right-to-erasure verification, post-incident forensics. Without a durable audit trail, we can't answer those queries. We considered shipping straight to ClickHouse via JSON inserts — rejected because the unstable-partition load during a traffic spike would lose log rows; Vector's buffered-disk-then-batch is the standard pattern that survives backpressure.

When it fails. Buffer fills → spillover starts dropping INFO logs. Mitigation: alert on disk-pressure within 5min; scale-out automatic via DaemonSet. The wedge case is a Vector bug that drops the wrong tag class — caught by daily audit-log diff vs DB diff (count of soft-deleted rows in DB should equal count of sweeper-delete log entries — discrepancy = log-loss bug).

MonitoringPrometheus + Alertmanager + Grafana

Pull-based metrics scrape every 15s from gw, readsvc, writesvc, cache, metadb, outbox, abuse, ttlsweeper, cdcrelay, orphangc. Records cache_hit_ratio, expired_unswept_count, consumer_group_lag, s3_503_slowdown_total, replication_lag_ms, p99_latency_ms, dlq_depth, orphan_count. Alertmanager fires PagerDuty pages on threshold breaches (e.g. cache_hit_ratio<80% for 2min, lag>10k for 5min, orphan_count > 200k). Long-term storage in Mimir; 30d Prometheus retention then ship to Mimir for 1y.

Why it exists. Every failureMode entry across this design names a specific page-on-call signal — without an alerting plane those pages have nowhere to fire. We picked Prometheus + Alertmanager over a managed product (DataDog, NewRelic) because (a) cost at this scale (28k QPS × paste_id label cardinality), (b) SLO definitions live as code in PromQL we version with the service, (c) Alertmanager's silencing + grouping primitives match the on-call workflow.

When it fails. Alertmanager itself fails — pages silently disappear. Mitigation: dead-man's-switch from an external Pingdom probe; PagerDuty also has its own dead-man's-switch (it pages if it stops getting heartbeats). The 2018 GitHub Octopus / Orchestrator class incident showed why this matters — alerts didn't fire because the alerting plane was offline. Cardinality explosion (a regression that adds paste_id as a metric label) is detected by a per-job series_per_target alarm.

Stage by stage

The same 10 stages the workspace walks, answered.

01Clarifications

What would you ask before drawing a single box?

Surface the constraints up front:

  • Visibility model? Public (cacheable), unlisted (URL is the secret), private (owner-only, authed), password-protected (server-side check). Drives the cache-key + auth design.
  • TTL? Yes — required. 1h / 1d / 7d / 30d / never. Drives sweeper + lifecycle.
  • Burn-after-read? Optional 'view-once' mode. Drives atomic CAS on read.
  • Blob size? 50KB typical, 1MB p99, 10MB hard cap. Bandwidth-bound, not just QPS-bound.
  • Read/write ratio? ~20:1 average; viral pastes >1M:1 (a single key).
  • Authenticated vs anonymous? Both. Anonymous can post (rate-limited per IP); auth gives owner control + private/longer TTL.
  • Latency? p99 read public CDN-hit < 60ms; p99 read miss < 350ms; p99 create < 250ms.
  • Abuse? Yes — DMCA, malware, phishing, IOC dumps. Real, regulatory, and ongoing (Pastebin.com 2014-2020 era is the case study).

Assumptions to state:

  • 30M new pastes/day → ~350 creates/sec avg, ~1,400 peak.
  • 600M reads/day → ~7k reads/sec avg, ~28k peak; viral can push 50k on a single key.
  • Avg 50KB blob, 1MB p99 → ~1.5 TB/day raw ingest; ~21 TB hot working set (14d mean retention).
  • 75% of pastes are unlisted/private, 25% public. Public + unlisted are edge-cacheable (the unguessable URL is the cache key); private/password are never cached at edge — drives cache-strategy bimodality.

02Functional reqs

What must this system actually do?

  • Create a paste (POST a blob + visibility + TTL); return a paste_id and short URL.
  • Read a paste (GET /:paste_id) — public is cacheable at edge, private is auth-gated, password is challenged.
  • Expire a paste (TTL elapses) — read returns 410 Gone.
  • Disable a paste (DMCA / abuse / owner action) — same effect: 410 Gone, plus CDN purge.
  • (Optional) Burn-after-read — first GET succeeds, subsequent GETs 410.
  • (Optional) List own pastes (auth: owner_id partition).

03Non-functional

What must it promise about speed, uptime and correctness?

  • Availability: 99.99% on the public read path (CDN-absorbed, multi-AZ origin). 99.9% on writes.
  • Latency: p99 < 60ms public-CDN-hit; < 350ms public-miss; < 250ms private read; < 250ms create.
  • Durability: Writes must not be lost. Sync replication on metadata; S3 11×9s for blobs.
  • Consistency: Read-after-write for the creator; eventual for others is fine. Abuse-disable visible globally within 60s (matches CDN edge TTL).
  • Scalability: Horizontal on every tier. No single shard pinned > 10% of total load.
  • Security: Private content NEVER cached at a shared edge. Password check is constant-time. Abuse pipeline is async + at-least-once (outbox).

04Capacity estimation

How much load and data does this have to hold?

Inputs: 30M creates/day, 20 reads/paste avg, 25% public, 50KB avg blob, 14d mean retention, 4× peak factor, 92% edge hit rate.

  • Creates: 30M / 86,400 ≈ 350 QPS avg, 1,400 QPS peak.
  • Reads: 600M / 86,400 ≈ 7,000 QPS avg, 28,000 QPS peak. Viral spike on a single key: +50k/s.
  • Origin reads after edge: 28k × (1 − 0.92 × 0.25) ≈ 22k/s in the worst case (much lower in practice because private content is the small minority).
  • Hot working set: 30M × 50KB × 14d ≈ 21 TB. 5y storage with 1.5 TB/day raw → ~3 PB cold, but lifecycle to IA → Glacier collapses cost by ~85%.
  • Origin egress: 600M × 50KB × ~10% (edge pass-through) ≈ 3 TB/day through the read service (bytes that aren't presigned-URL-redirected).
  • Bandwidth: Peak viral burst on 1MB paste = 50,000 × 1 MB = 50 GB/s — absorbed entirely at the CDN; origin sees ~1 MB/s (the regional shield collapses misses to ≤1 revalidation/s per 60s edge TTL).
  • S3 PUT/sec/prefix limit: 3,500. With 65k prefixes via 4-char hash, headroom is 220M/sec — never the bottleneck.
  • Cache footprint: ~20 GB Redis hot working set (top 0.1% of pastes serve 50% of traffic per Zipfian skew).

05API design

What does the outside world call, and what comes back?

Create is size-bimodal — two shapes for one concept. Avg paste is 50KB text; p99 is 1MB; the cap is 10MB. Proxying a 10MB body through the write service wastes app-tier ingress and memory, so large pastes go directly to the object store via a presigned PUT — mirroring the private read path, which already serves >10KB blobs via presigned GET rather than proxy-streaming. Small text takes a one-round-trip inline fast-path.

Fast path — small text (≤ inline threshold, e.g. 256KB): one round trip; the write service holds the bytes and PUTs them itself.

POST /api/v1/pastes
Content-Type: application/json
Authorization: Bearer <token>     # optional; anon also allowed (rate-limited)
Idempotency-Key: <uuid>           # client-side retry-safe

{
  "body": "...inline text...",
  "visibility": "public" | "unlisted" | "private" | "password",
  "password": "...",               # required iff visibility=password
  "ttlSeconds": 3600,              # 0 / null = never
  "burnAfterRead": false
}

201 Created
{
  "id": "p_K3vHt2",
  "url": "https://paste.example/p_K3vHt2",
  "expiresAt": "2026-05-09T03:00:00Z"
}

Upload path — large / file pastes (no body in the POST): two steps. The POST reserves metadata and hands back a short-lived presigned PUT; the client streams the bytes straight to the object store, bypassing the app tier entirely.

POST /api/v1/pastes
Content-Type: application/json
Authorization: Bearer <token>
Idempotency-Key: <uuid>

{
  "contentSize": 8388608,          # declared up-front; rejected if > 10MB cap
  "contentType": "application/zip",
  "filename": "dump.zip",          # optional
  "visibility": "private",
  "ttlSeconds": 86400,
  "burnAfterRead": false
}

201 Created
{
  "id": "p_K3vHt2",
  "url": "https://paste.example/p_K3vHt2",
  "uploadUrl": "https://blobs.paste.example/aF3x/p_K3vHt2?X-Amz-Signature=...",
  "uploadExpiresAt": "2026-05-08T03:05:00Z",   # presigned PUT TTL, ~5 min
  "expiresAt": "2026-05-09T03:00:00Z"
}
PUT <uploadUrl>                    # direct to object store; NOT through the gateway
Content-Type: application/zip
Content-Length: 8388608

<raw bytes>
→ 200 OK   (object store responds; the signature pre-authorizes exactly this key + size)

The metadata row is born in upload_state = pending. The client flips it to committed by calling POST /api/v1/pastes/:id/finalize once the PUT returns — that finalize call is the sole authoritative commit path. A read of a still-pending paste returns 409 Conflict (with Retry-After), never a 502 — treating the reservation as existing-but-not-finalized is what makes "metadata first, bytes second" safe (see deep-dive 6). Reservations that are never finalized are reaped by the Orphan Blob GC off the WHERE upload_state='pending' partial index. Presigned URLs are short-TTL (~5 min) and scoped to exactly one object key — S3 offers no single-use semantics, so a leaked URL can be replayed within the window, but a replay only rewrites that same key while the row is still pending; the declared contentSize is enforced by the signature so a client can't upload a 1GB blob against a 1MB reservation.

GET /:paste_id
# Public: served by CDN with `Cache-Control: public, max-age=60, s-maxage=60`.
# Unlisted: same, but URL is unguessable (12-char base62).
# Password: 401 + WWW-Authenticate; client retries with X-Paste-Password header.
# Private: 401 if no/invalid bearer; 403 if bearer ≠ owner.
# Pending upload (bytes not yet finalized): 409 Conflict + Retry-After.
# Expired / abuse_disabled: 410 Gone.

200 OK
Content-Type: text/plain; charset=utf-8
Cache-Control: public, max-age=60, s-maxage=60       # public/unlisted
Cache-Control: private, no-store                      # private/password

<blob body>
GET /:paste_id
# Public: served by CDN with `Cache-Control: public, max-age=60, s-maxage=60`.
# Unlisted: same, but URL is unguessable (12-char base62).
# Password: 401 + WWW-Authenticate; client retries with X-Paste-Password header.
# Private: 401 if no/invalid bearer; 403 if bearer ≠ owner.
# Expired / abuse_disabled: 410 Gone.

200 OK
Content-Type: text/plain; charset=utf-8
Cache-Control: public, max-age=60, s-maxage=60       # public/unlisted
Cache-Control: private, no-store                      # private/password

<blob body>

Owner takedown — the "owner action" half of the disable requirement (DMCA/abuse disables ride the admin path):

DELETE /api/v1/pastes/:id
Authorization: Bearer <token>      # owner-only; 401 without bearer, 403 if bearer ≠ owner

204 No Content                     # idempotent — repeating the DELETE also returns 204
# Sets soft_deleted + a paste.disabled outbox row in the same txn; the abuse
# worker's paste.disabled consumer runs the Hot Cache invalidate + CDN purge,
# so the next reader sees 410 Gone within the 60s edge-TTL window.

The optional "list own pastes" endpoint (GET /api/v1/pastes?cursor=&limit=, keyset-paginated over the (owner_id, created_at DESC) index) is deliberately left unspecified here — it's a straightforward authed read with no new design surface.

06Data model

What gets stored, and what is it looked up by?

pastes table (sharded by hash(paste_id), 8 Citus shards):

fieldtypenotes
paste_idvarchar(12)PK; base62 from Snowflake-style id
owner_idbigint?nullable (anonymous)
blob_idvarchar(40){rand4}/{paste_id} S3 key
blob_sizeintdeclared at create; enforced by presigned sig
upload_stateenumpending → committed; read of pending = 409
visibilityenumpublic/unlisted/private/password
password_hashvarbinary(32)?argon2id; null unless visibility=password
burn_after_readbool
burn_consumedboolCAS target for burn-after-read
created_attimestamp
expires_attimestamp?partial index target
soft_deletedboolsweeper sets; 7-day grace before S3 lifecycle
abuse_disabledboolabuse worker sets
versionintoptimistic-concurrency for ACL/visibility flips

Indexes:

  • PK on paste_id.
  • Partial index on (expires_at) WHERE expires_at IS NOT NULL AND NOT soft_deleted — sweeper.
  • (owner_id, created_at DESC) — list-mine (auth path).
  • (abuse_disabled, created_at) — admin abuse review.
  • Partial index on (created_at) WHERE upload_state = 'pending' — the orphan-GC sweep for reservations whose bytes never arrived.

outbox table (same shards, same txn):

fieldtype
event_idbigserial
event_typevarchar(32)"paste.created" / "paste.disabled" / "paste.expired"
paste_idvarchar(12)
payload_jsonjsonb
created_attimestamp
published_attimestamp?NULL until relay confirms publish

Why this split:

  • Metadata in Postgres + Citus because we need transactional (insert paste, insert outbox) and ad-hoc admin queries (DMCA review).
  • Blobs in S3 because Postgres is the wrong storage for 50KB-1MB byte-payloads at 30M/day.
  • Hot cache in Redis because Postgres + replicas can't sustain 28k reads/sec per shard.

07High-level design

Which components handle a request, and in what order?

Ingest split — public vs private hostnames:

cdn.paste.example     → CDN edge (cacheable, anonymous)
api.paste.example     → API gateway (auth, never cached)

This is the load-bearing decision. Putting public + private on the same host with Vary: Authorization would make every signed-in viewer a unique cache entry — destroys hit ratio. Splitting hostnames means the CDN is configured at edge to NEVER cache api.paste.example regardless of upstream Cache-Control.

Read flow — public: Client → CDN edge (POP) → (hit: serve cached 200, ~20ms) | (miss: tiered-cache shield → CDN origin pull → API gateway → Read Service → Hot Cache → (hit: serve, populate edge) | (miss: → Metadata DB replica → S3 GET → proxy-stream the body back so the edge caches it)).

Read flow — private/password: Client → API gateway (mTLS, JWT) → (Rate Limiter + Auth ext_authz) → Read Service → Hot Cache (metadata) → ACL evaluation against principal → (allow: → S3 fetch with private cache-control) | (deny: 403). NEVER cached at edge.

Write flow: Client → API gateway → (Rate Limiter + Auth) → Write Service → S3 PUT (blob) → Metadata DB INSERT (paste row + outbox row, single txn) → Hot Cache write-through populate → 201 Created. Outbox relay tails Postgres logical replication, publishes paste.created to Kafka asynchronously.

Async fan-out: Outbox → Kafka → Abuse Worker → (Bloom hit confirmed against the exact hash store / external API verdict) → if bad: Metadata DB UPDATE abuse_disabled + Hot Cache invalidate + CDN edge purge.

TTL sweep: Cron every 5min → TTL Sweeper reads partial index → soft-delete metadata + emit paste.expired outbox events + invalidate cache + (S3 lifecycle handles bytes 24h later, rounded to UTC midnight). Metadata DB is truth for the API contract; S3 lifecycle is only a janitor for byte cost.

08Deep dives

Which part breaks first, and what do you do about it?

1. The cache-key for private content. The bug to avoid is Vary: Cookie on public pastes — destroys hit ratio. The fix is hostname-split (public vs api), Cache-Control: private, no-store on private responses, and stripping Cookie + Authorization at the edge before computing the cache key for public requests. We deliberately picked NOT to use signed URLs on the read path (CloudFront-style) for private pastes because (a) signed URLs expose paste_id in logs; (b) revocation is impossible until expiry; (c) the auth-on-every-read is cheap when Auth + Rate Limiter are <10ms inline.

2. Burn-after-read race. Two readers hit a view-once paste 1ms apart from different POPs. The atomic primitive is UPDATE pastes SET burn_consumed=true WHERE paste_id=$1 AND burn_consumed=false RETURNING blob_id — single-row CAS on the metadata DB primary. Only the winner gets blob_id; the loser sees RETURNING empty rows = 410. We do NOT cache burn-after-read pastes (Cache-Control: no-store from the read service); we always hit the DB primary, never a follower; we invalidate cache + CDN as part of the same critical section.

3. Viral paste hot key. A paste hits the front page of HN. 50k req/s on one paste_id. CDN tiered cache absorbs ~99.9% — origin sees ≤1 req/s on that paste because the regional shield collapses concurrent misses. If origin DOES start to feel pressure (cache shard saturated), the Redis cluster's hash-slot for that paste is one of 16 — 1/16 of cluster capacity. In-process LRU on the read service catches the next layer. In the worst case (entirely new paste, cache cold), single-flight in the read service collapses N concurrent misses to 1 DB query. Mitigation order: edge tiered cache > in-process LRU > single-flight > Redis shard > DB replica.

4. TTL honesty. Our API says 'expires in 1 hour'. S3 lifecycle says 'within 24h, rounded to UTC midnight'. The honest contract is: the metadata DB is the truth — the read path returns 410 the moment now() > expires_at, regardless of whether S3 still has the bytes. CDN edge TTL is bounded at 60s so a stale entry can't outlive the metadata flip. S3 lifecycle is a janitor for byte cost, not a correctness mechanism.

5. Abuse takedown latency. A malware payload was just pasted. It's been read 2,000× in 60s. Ops hits 'disable'. Latency to next-reader-sees-410: metadata flip (~10ms within the abuse worker) + cache invalidate (~50ms) + CDN purge (60-300ms across 300 POPs). Worst case is a CDN POP that was serving cached 200s — its TTL is 60s, so the absolute worst case is 60s. We accept that 60s window deliberately because tightening edge TTL below 60s tanks hit ratio for non-abuse cases.

6. One-step inline vs two-step presigned create. Create is size-bimodal, so it has two shapes. The naive single endpoint — POST the body inline, write service PUTs it to S3 — is correct only for small text (avg 50KB). For a 10MB file it forces the megabytes through the app tier twice (client→writesvc ingress, writesvc→S3 egress), burning gateway memory and origin bandwidth for nothing. The fix mirrors the private read path, which already hands back presigned GETs for >10KB blobs: for large pastes, POST reserves metadata and returns a presigned PUT, and the client streams bytes straight to the object store (edge e-client-blob, never touching the gateway or write service).

That inverts the durability ordering, which is the subtle part. The inline path PUTs-then-commits, so a half-write is invisible (no row → 404). The presigned path commits the row first (upload_state=pending) and uploads second — the "reservation-then-fulfill" pattern. We earlier rejected reservation-then-fulfill because a metadata row pointing at absent bytes returns a confusing 502. The thing that makes it safe is the explicit state flag: a read of a pending paste returns 409 Conflict (reservation exists, not finalized), not 502, and the client flips pending → committed by calling POST /api/v1/pastes/:id/finalize once the PUT returns — that finalize is the sole authoritative commit. Abandoned reservations (client uploaded nothing, or never finalized) are reaped by the Orphan Blob GC off the WHERE upload_state='pending' partial index — the same component that catches inline-path orphans, now sweeping both directions. Presigned URLs are size-bound by the declared contentSize and short-TTL (~5 min); S3 can't make them single-use, so a leaked URL can be replayed until the signature expires — the bound is that a replay only rewrites that one key at that one size, never a 1GB blob against a 1MB reservation.

09Trade-offs

What did this design cost, and what breaks at 10×?

What we accepted:

  • 60s CDN edge TTL = up to 60s of stale-cache after abuse_disabled flips. Trade was hit ratio on the 99.99% non-abuse traffic vs the 0.01% takedown latency.
  • Outbox over 2PC for write→Kafka — at-least-once into the abuse pipeline; abuse worker is idempotent. 2PC was rejected on liveness (coordinator failure stalls writes).
  • Semi-sync replication on Postgres (1 sync follower + 1 async) → 30s RTO write outage on primary loss with RPO=0. Pure-async would give RTO 5s but RPO=30s, unacceptable for paste durability. Pure-sync to 2 followers would have ~1.5× write latency for negligible durability gain over 1 sync.
  • S3 lifecycle handles cold-tier deletion; metadata-DB is truth for expiry. We don't try to make S3 lifecycle real-time — the design accommodates lifecycle's 24h SLA.
  • Regional outage = read+write outage of cloud-provider-class duration. We deliberately ship single-region for the metadata DB (Citus + semi-sync replication for in-region 30s RTO, 0 RPO). Cross-region S3 CRR holds blobs (~15min RPO, hence blobstore.failover.rpoSeconds=900) but if us-east-1 is gone, the metadata DB is gone too and reads can't resolve paste_id → blob_id. Recovery is restore-from-backup, not active-active. The math: 99.9% regional uptime × N regions to upgrade to active-active is ~$5M/year vs. an expected once-per-3-year multi-hour outage. We picked the cheaper bucket and document the RTO target as 4 hours for a full regional disaster (warm DR with hourly snapshots in us-west-2).
  • Fail-open on rate-limiter for reads, fail-closed for writes. Brief abuse during a 15s ratelimit-failover gap is preferable to dropping legitimate paste reads. Counter-attacker mitigation is consistent-hash routing on client_ip at the front-door LB so an attacker pins to one gateway replica during the gap.
  • Orphan blobs sweep daily on a 24h grace, accepting up to ~120k orphan blobs/day visible to S3 cost analytics (caught and hard-deleted within 24h-7d window).

What breaks at 10× scale (~14k creates/s, 280k reads/s):

  • Single Citus shard hot-spot — re-shard from 8 to 32 shards.
  • Single Redis cluster CPU — split into hot/cold tiers, with the hot tier on Redis Enterprise / KeyDB for higher throughput.
  • Origin egress on cache miss — move public proxy-streaming off the read service onto a dedicated bytes-fleet so the metadata tier stays CPU-bound, not egress-bound.
  • Abuse worker fan-out — partition by paste_id range AND by content classification (text-only fast-path vs binary-detected slow-path).

Per-component failure stories: see "Failure scenarios we model" below.

Trace catalogue

The simulator models 8 specific paste journeys, each with budgets derived from the SLO table above:

Trace IDJourneyBudget
pastebin:read-public-cdn-hitPublic read absorbed at CDN edge (the 90%+ path)60ms
pastebin:read-public-cdn-miss-cache-hitCDN miss, origin serves from Redis cache200ms
pastebin:read-public-cache-miss-db-blobCold path: cache miss, replica + S3 GET800ms
pastebin:read-private-paste-authPrivate/password read; bypass CDN; auth-gated; S3 GET600ms
pastebin:create-pasteTwo-step: POST reserves metadata + outbox, then presigned byte upload direct to object store1500ms
pastebin:burn-after-readView-once: CAS on metadata + invalidate everywhere1000ms
pastebin:abuse-scan-fanoutOutbox → abuse worker → external scan → DB+cache+CDN30000ms
pastebin:ttl-sweep-expireSweeper picks expired rows → soft-delete + outbox30000ms

Each is a deterministic walk through the canonical hldNodes/hldEdges; the simulator verifies them against the user's diagram and surfaces gaps as concrete prompts ("you have no edge from Write Service to a stream — your abuse path can't run").

Failure scenarios we model

The simulator ships 8 chaos scenarios specific to Pastebin, each grounded in a cited real-world failure mode:

Scenario IDReal-world precedentCategorySeverity
pastebin:viral-paste-stampedePastebin.com Anonymous-dump waves (2010s); Cloudflare Argo / tiered cache rationaletraffichigh
pastebin:s3-prefix-throttleAWS S3 docs: 3,500 PUT/sec/prefix limit + 503 SlowDown class; sequential-id hot-prefix as documented in S3 perf design patternsdatamedium
pastebin:metadata-leader-failoverGitLab.com 2017-01-31 (rm -rf primary); GitHub 2018-10-21 Orchestrator failoverdatahigh
pastebin:cache-cluster-coldCache-stampede class (Wikipedia article cites Facebook 2010 4-hour outage); Discord ScyllaDB migration thundering-herd lessonsdatamedium
pastebin:cdn-edge-control-planeCloudflare 2025-06-12 KV outage (third-party storage failed → 91% KV error rate, 2h28m)infrahigh
pastebin:writesvc-bad-deployGitHub 2020-06 deploy-class availability incidents; Knight Capital 2012 (deploy reached only 7 of 8 servers)processhigh
pastebin:abuse-scan-third-party-outCloudflare 2019-07-02 regex CPU (third-party-class outage); Heroku 2022-04 cert + DNSdepsmedium
pastebin:ttl-sweeper-shard-wipeGitLab 2017-01-31 (destructive op without guard); class of "sweeper bug deletes too much"processhigh

Each scenario probes the trace lanes that matter, mutates a specific edge or node according to the canonical's authored internals (e.g. failover.rtoSeconds=30 → metadata-leader-failover applies a 30-second writes-blocked window), and returns a diff that explains why a probe degraded or broke.

Primary sources

  • Cloudflare KV outage (2025-06-12)
  • Discord — How Discord Stores Trillions of Messages
  • GitHub — sharded replicated rate limiter in Redis
  • AWS S3 — performance design patterns + lifecycle
  • Cloudflare — Tiered Cache & Argo
  • PrivateBin / 0bin — zero-knowledge paste

Now defend it

Reading a design is not the same as being able to hold one under questioning. The workspace asks the same questions an interviewer would, and the simulator disagrees with you when the diagram does not support the claim.

Work Pastebin yourself