Skip to content

Latest commit

 

History

History
58 lines (51 loc) · 3.21 KB

File metadata and controls

58 lines (51 loc) · 3.21 KB

Deliberately out of scope

Each of these is a reasonable idea. Each is also the kind of "just one more thing" that left four previous repositories unfinished.

Adding to this file is encouraged. Acting on it before Stage 6 ships is not.

  • A web dashboard. The API and a CLI are enough. A frontend is a stack that would have to be relearned first, and it would not add anything an interviewer asks about.
  • Authentication beyond a static API key. Real auth is a solved problem and demonstrates nothing this project is trying to demonstrate.
  • Multi-tenancy, quotas, per-tenant fairness. Genuinely interesting, and a whole project on its own.
  • Cron expressions. A scheduled_for timestamp covers the interesting problem. Parsing cron syntax does not.
  • Workflow DAGs and job dependencies. This is the natural sequel and it is a good one — after Stage 6.
  • Kubernetes. Docker Compose demonstrates the same understanding and costs two days less.
  • Multiple language SDKs. One HTTP API is the interface.
  • Exactly-once semantics. Not achievable end to end. The project implements at-least-once delivery plus idempotent execution and says so plainly.

Added 2026-09-05, after building stages 1-6

These come from things the build actually ran into, rather than from a wish list. Same rule applies: adding is encouraged, acting is not.

  • Kafka as the claim path, not just a notification. Today Postgres decides who gets a job and Kafka only carries a wake-up. Making the partition assignment be the work assignment would remove a round trip per job, and would mean rethinking what happens when a partition moves mid-execution. Genuinely interesting; also a rewrite of the claim path.
  • A test that observes a partition moving between two live worker processes mid-batch. Stage 4's gate is only partially evidenced without it — see the evidence section in PLAN.md. This is a gap, not a feature, and it is the one thing on this list that arguably belongs above the line.
  • SKIP LOCKED batch claiming. Claiming N jobs per statement instead of one would cut round trips substantially at high worker counts. Needs care: a worker holding a batch that dies makes the recovery unit N jobs instead of 1.
  • TimescaleDB or partitioning for job_attempts. The attempt history grows without bound and nothing prunes it.
  • Backpressure on submission. POST /jobs accepts work regardless of how far behind the workers are. A queue-depth check returning 429 is the obvious fix and was never needed to prove anything.
  • Postgres LISTEN/NOTIFY instead of the idle poll. Workers currently poll when idle. NOTIFY would cut the latency floor and the idle load, and adds a second notification path to reason about alongside Kafka.
  • Prometheus metrics. /stats covers what the chaos harness needs. Real metrics are a different project's worth of yak-shaving.
  • Handler-side transactional outbox. Decision 012 shows dedup only works transactionally with the effect. The general version of that is an outbox table the handler writes in the same transaction, drained separately. It is the correct answer for effects against external systems and it is a whole subsystem.