Distributed Cron — Mass Scheduled Email

What would you ask before drawing a single box?

Ambiguity you would resolve with the interviewer: scope, scale, who uses it, what counts as done.

About Distributed Cron — Mass Scheduled Email

Single trigger, 50M recipient idempotency, catch-up.

Difficulty
advanced
Time
about 75 minutes
Stages
10
Topic
Queues, Pub/Sub & Event Streaming

How this problem is worked

Ten stages, from the questions you would ask an interviewer to the trade-offs you would defend. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.

  1. 01ClarificationsWhat would you ask before drawing a single box?
  2. 02Functional reqsWhat must this system actually do?
  3. 03Non-functionalWhat must it promise about speed, uptime and correctness?
  4. 04Capacity estimationHow much load and data does this have to hold?
  5. 05API designWhat does the outside world call, and what comes back?
  6. 06Data modelWhat gets stored, and what is it looked up by?
  7. 07Use-case breakdownHow does each requirement actually get served?
  8. 08High-level designWhich components handle a request, and in what order?
  9. 09Deep divesWhich part breaks first, and what do you do about it?
  10. 10Trade-offsWhat did this design cost, and what breaks at 10×?

Primary sources for this problem

  • Sundheim — Reliable Cron across the Planet (ACM Queue 2015, Borgcron paper)
  • LinkedIn — Air Traffic Controller: Member-First Notifications (2016)
  • LinkedIn — Hermes mass email retros (eng blog)
  • Pinterest — NEP Notification System and Relevance (Medium 2017)
  • Uber — Cherami: Uber's durable distributed task queue
  • Uber — Announcing Cadence (Temporal predecessor)
  • DoorDash — Cadence as a Fallback for Event-Driven Processing
  • Discord — How we store trillions of messages (Cassandra → ScyllaDB)
  • Stripe — Designing robust idempotency keys (Brandur)
  • Brandur — Implementing Stripe-like Idempotency Keys in Postgres
  • AWS Builders' Library — Idempotency at Scale (re:Invent ARC403 2021)
  • AWS — Handling SES throttling (Maximum sending rate exceeded)
  • AWS — SES sending quotas & dedicated IP warmup
  • Mailgun — Bulk email sending with queue management
  • Slack — Tracing notifications (Go → Kafka → Elasticsearch)
  • Shopify — High availability background jobs
  • Vallery Lancey — Kubernetes CronJob Failed For 24 Days (case study against naive cron)
  • Quartz Scheduler — JDBC JobStore clustering docs (the misfire-threshold story)
  • Cloudflare — Cron Triggers internals (Nomad-distributed schedulers)
  • Apache Kafka — KIP-429 cooperative incremental rebalancing
  • Beyer et al. — SRE Workbook ch. 21–23 (overload, cascading failures, critical state)
  • Campbell & Majors — Database Reliability Engineering ch. 9 (fencing tokens)

Browse the full problem catalog, or see what the simulator does and does not model.