Build a stream processor (Flink / Kafka Streams style)

No scenes authored for this problem yet.

About Build a stream processor (Flink / Kafka Streams style)

Event time isn't processing time. Build a stream processor that tracks watermarks, windows by event time, holds keyed state with checkpoints, and recovers exactly-once after a node crash — and internalize why 'streaming SQL' is mostly the dataflow model in a different syntax.

Difficulty
advanced
Time
about 95 minutes
Stages
9
Topic
Queues, Pub/Sub & Event Streaming

How this problem is worked

Nine stages, from what the thing is for to how it compares with the real implementations. Each asks one question, and the simulator runs the architecture you draw against the requirements you wrote.

  1. 01Purpose & invariantsWhat is this for, and what must always be true of it?
  2. 02Workload characterizationWho writes, who reads, and in what shapes?
  3. 03Data model & on-disk formatWhat does the data look like at rest?
  4. 04Core algorithmsHow do the write path and the read path actually work?
  5. 05Distribution & replicationHow does this scale out and survive losing a machine?
  6. 06Consistency & correctnessUnder concurrency and failure, what is guaranteed?
  7. 07Failure modes & recoveryWhat actually happens when each part fails?
  8. 08Operational characteristicsCan a human run this at three in the morning?
  9. 09Trade-offs & comparisonWhere does this sit against the alternatives?

Primary sources for this problem

  • Akidau et al. — The Dataflow Model (VLDB 2015)
  • Akidau, Chernyak, Lax — Streaming Systems (O'Reilly)
  • Apache Flink docs — Stateful Stream Processing & checkpointing
  • Kafka Streams architecture (Confluent docs)
  • Chandy & Lamport — Distributed snapshots paper (1985)
  • Kreps — The Log: What every software engineer should know

Browse the full problem catalog, or see what the simulator does and does not model.