Build a workflow engine (Temporal / Airflow / Cadence style)

About Build a workflow engine (Temporal / Airflow / Cadence style)

A function that survives crashes, restarts, and re-deploys — and still finishes. Build a durable execution engine where workflow code is replayed deterministically from an event history, activities retry with exponential backoff, sagas compensate on failure, and the same workflow definition runs identically a year later. Internalize why 'just retry the cron job' breaks at the second step.

Difficulty
advanced
Time
about 90 minutes
Stages
9
Topic
Consensus, Coordination & Durable Execution

How this problem is worked

Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.

  1. 01Purpose & invariantsWhat is this for, and what must always be true of it?
  2. 02Workload characterizationWho writes, who reads, and in what shapes?
  3. 03Data model & on-disk formatWhat does the data look like at rest?
  4. 04Core algorithmsHow do the write path and the read path actually work?
  5. 05Distribution & replicationHow does this scale out and survive losing a machine?
  6. 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
  7. 07Failure modes & recoveryWhat actually happens when each part fails?
  8. 08Operational characteristicsCan a human run this at three in the morning?
  9. 09Trade-offs & comparisonWhere does this sit against the alternatives?

Primary sources for this problem

  • Temporal — Durable Execution / Temporal Service concept page (the plain-English definition + the engine↔worker split): docs.temporal.io/temporal
  • Temporal — Workflow Definition, the determinism constraints (what you may NOT do in workflow code and why replay demands it): docs.temporal.io/workflow-definition
  • Temporal — Activities (the activity boundary, at-least-once semantics, why idempotency is the developer's job): docs.temporal.io/activities
  • Temporal — Retry Policies (the concrete backoff defaults: 1s, ×2, cap, jitter, non-retryable): docs.temporal.io/encyclopedia/retry-policies
  • Temporal — Versioning / patching (the longest-running-deploy problem, before you ever ship a v2): docs.temporal.io/develop/go/versioning
  • Garcia-Molina & Salem — 'Sagas', ACM SIGMOD 1987 (the origin of compensating transactions; compensation restores an acceptable approximation, not the exact prior state): dl.acm.org/doi/10.1145/38713.38742
  • Greg Young / Martin Fowler — Event Sourcing (the pattern that underpins workflow history: the log is truth): martinfowler.com/eaaDev/EventSourcing.html
  • SE Radio 596 — Maxim Fateev on Durable Execution with Temporal (the creator, ex-AWS SWF / Uber Cadence, explaining the model conversationally): se-radio.net/2023/12/se-radio-596
  • Uber Engineering — Cadence overview (Fateev & Abbas; the design rationale for durable, long-running business logic): uber.com/en/blog/open-source-orchestration-tool-cadence-overview
  • AWS Step Functions — developer guide (contrast the JSON state-machine / ASL model with code-as-workflow): docs.aws.amazon.com/step-functions/latest/dg/welcome.html
  • Apache Airflow — core concepts (DAG scheduling + task-level retry, and why it is NOT replay-deterministic): airflow.apache.org/docs/apache-airflow/stable/core-concepts
  • Azure Durable Functions — concepts (the replay model expressed via generators/yield; a useful third data point): learn.microsoft.com/azure/azure-functions/durable/durable-functions-overview
  • Stripe — Idempotency keys (the canonical real-world idempotency-key pattern; pair it with the double-charge scene): docs.stripe.com/api/idempotent_requests

Browse the full problem catalog, or see what the simulator does and does not model.