Skip to content

Latest commit

 

History

History
427 lines (333 loc) · 20.8 KB

File metadata and controls

427 lines (333 loc) · 20.8 KB

Decision log

Append-only. Never edit or delete a past entry — supersede it with a new one that references the old by number and says what changed.

Every entry records the alternatives that lost and why. That reasoning is the part interviewers ask about, and it is much harder to reconstruct months later than the code is.

Format:

## NNN — Title
**Date:** YYYY-MM-DD · **Status:** accepted | superseded by NNN

**Context.** What forced a choice.
**Decision.** What we chose.
**Alternatives.** What lost, and why.
**Cost.** What this makes harder. Every real decision has one.

001 — Build a distributed job scheduler

Date: 2026-09-02 · Status: accepted

Context. Needed one portfolio project that demonstrates API design, low-level design, concurrency, and system-design vocabulary for SDE-1 interviews at tier-1 Indian product companies — in C++, without requiring a frontend stack.

Decision. A distributed job scheduler with at-least-once delivery and idempotent execution.

Alternatives.

  • Seat reservation system — the strongest concurrency-correctness story, but the scheduler exercises worker pools and graceful shutdown more directly, which suits the concurrency material being read alongside it.
  • Live location tracking — Kafka-heavy but thinner on low-level design.
  • Finishing matchbook — excellent depth credential, but it is a single-threaded in-memory library with no service surface, and the trading domain reads as quant to product-company interviewers.
  • Redis clone — highly legible, but building infrastructure rather than using it demonstrates a narrower slice of what SDE-1 loops actually probe.

Cost. Requires standing up three external dependencies (Postgres, Redis, Kafka), which is real setup time before any interesting code runs.


002 — Postgres holds state, Kafka is only transport

Date: 2026-09-02 · Status: accepted

Context. Both systems can technically hold job state.

Decision. Postgres is the source of truth for job records, state, and attempt history. Kafka carries work to workers and nothing more.

Alternatives. Kafka-as-database using compacted topics — fashionable, but it makes "show me every failed job from Tuesday" hard, and that query is the kind of thing a real operator needs.

Cost. Two systems to keep consistent, and a window where a message exists in Kafka but its row is not yet committed in Postgres. Stage 4 has to address this explicitly.


003 — Commit Kafka offsets after execution, not on receipt

Date: 2026-09-02 · Status: accepted

Context. The offset commit point decides the delivery guarantee.

Decision. Commit after the handler completes. This gives at-least-once delivery: a crash between execution and commit means the job is redelivered.

Alternatives. Committing on receipt gives at-most-once, which loses jobs on a crash — unacceptable for a scheduler whose entire promise is that work is not lost.

Cost. Duplicates are now possible by design, which is why deduplication in Stage 5 is mandatory rather than optional. The project therefore claims effectively-once, never exactly-once.


004 — MSVC on Windows is the primary build target

Date: 2026-09-02 · Status: accepted

Context. Stages 3–5 are process-lifecycle exercises: kill -9 a worker and have another reclaim its lease, run three worker processes through a consumer group rebalance, script a harness that kills workers and drops Redis under load. Every dependency in the stack — librdkafka, Drogon, redis-plus-plus, libpqxx — is Linux-first with Windows as a port. That argued for developing in WSL2 with clang or gcc.

Decision. Stay on MSVC/Windows as the primary target, as originally written. Dependencies come from vcpkg. The chaos harness is Windows-native — taskkill /F for kills, PowerShell for orchestration — with Docker Compose supplying Postgres, Redis and Kafka as containers.

Alternatives.

  • WSL2 with clang or gcc as primary — better dependency support and a chaos harness that matches every tutorial and Stack Overflow answer in this space. Lost because it splits the development environment from the machine the work actually happens on, and a second toolchain is its own ongoing tax.
  • Linux only, no Windows support — simplest of all, but abandons the primary development machine entirely.

Cost. vcpkg has to successfully produce librdkafka and Drogon on MSVC; if either fights back, that is time spent on packaging rather than on the scheduler. The chaos scripts will not be copy-pasteable from prior art, and process-kill semantics differ from POSIX signals — taskkill /F is closer to SIGKILL than to anything catchable, so graceful-shutdown testing needs its own separate path. Revisit this decision if Stage 3 or Stage 4 stalls on toolchain problems rather than on design problems.


005 — JobStore is tested by one contract suite over two implementations

Date: 2026-09-02 · Status: accepted

Context. Testing only against an in-memory JobStore tests a hash map behind an interface. The failure modes that matter in a scheduler are Postgres-specific: claim semantics under concurrent workers, the isolation level a state transition runs at, and whether the idempotency key is enforced by a real unique constraint or by a read-then-write that only looks atomic.

Decision. Write the JobStore tests once as a parameterized GoogleTest fixture and run the same suite against both InMemoryJobStore and PostgresJobStore. The in-memory run happens on every local build; the Postgres run happens in CI against a service container, and locally when Docker is up. JobState transitions, RetryPolicy and WorkerPool shutdown stay pure unit tests with no infrastructure.

Alternatives.

  • In-memory only — fast and infra-free, but leaves the concurrency-correctness claims of the persistence layer entirely unverified.
  • Real Postgres for everything — highest fidelity, but every test then needs Docker, the suite slows to the point where it stops being run, and it removes most of the reason for JobStore to be an interface at all.

Cost. One CI workflow to maintain, and the contract suite constrains the interface: anything InMemoryJobStore cannot honour cannot go in JobStore. That constraint is mostly a feature, but it will occasionally be inconvenient when a Postgres-only capability looks tempting.


006 — Repository is named Deadman

Date: 2026-09-02 · Status: accepted

Context. Needed a GitHub repository name before a remote could be added.

Decision. Deadman. The lease is a deadman switch — it lapses when the worker holding it stops renewing, and that lapse is the entire product. The name encodes the core mechanism. The repository description carries the explanation: "Distributed job scheduler in C++20. Kill a worker mid-job — the job still runs exactly once."

Alternatives.

  • deadman-scheduler — self-describing without the README, but the second half carries nothing the description field cannot.
  • deadman-switch — names the mechanism outright, at the cost of obscuring that the product is a job scheduler.

Cost. The bare word is not searchable and collides with unrelated projects. This does not matter for a portfolio repository reached from a resume link.


007 — .clang-format checked in before the first source file

Date: 2026-09-02 · Status: accepted

Context. Formatting arguments are worth zero interview points and cost real time. Settling layout before any code exists avoids a reformatting commit later that buries the first real diffs.

Decision. A .clang-format at the repository root: BasedOnStyle: LLVM, IndentWidth: 4, ColumnLimit: 100, PointerAlignment: Left, AllowShortFunctionsOnASingleLine: Empty, SeparateDefinitionBlocks: Always. LLVM is the base because it is the default clang-format is best tested against; 4-space indent and left-aligned pointers match MSVC-world convention; 100 columns because 80 becomes cramped once namespaced types and std::chrono durations appear in signatures. The binary ships with Visual Studio 2022 at VC\Tools\Llvm\x64\bin\clang-format.exe (19.1.1), so no extra install.

Alternatives.

  • No config at all — every contributor's editor imposes its own, and the first real refactor produces a diff that is 90% whitespace.
  • Google style base — 2-space indent, which reads as cramped in a codebase with deeply namespaced types.

Cost. .clang-format governs layout only. The naming conventions in CLAUDE.md — snake_case functions, PascalCase types, member_ trailing underscore — are clang-tidy's domain and remain unenforced by tooling, so they stay a review-discipline matter unless a .clang-tidy is added later.


008 — MIT license

Date: 2026-09-02 · Status: accepted

Context. The GitHub repository was initialised with an MIT LICENSE before the local history was pushed. Accepting or replacing it was a real choice, not a formality.

Decision. Keep MIT.

Alternatives. Apache-2.0 offers an explicit patent grant, which matters for projects expecting outside contributors or corporate adoption. Neither applies to a portfolio repository, and MIT is shorter and more universally recognised by the people who will actually read it.

Cost. No patent grant. Irrelevant at this scale, and changing it later costs nothing while the project has a single author.


009 — vcpkg for dependencies, not CMake FetchContent or Conan

Date: 2026-09-02 · Status: accepted

Context. Six third-party libraries are needed — libpqxx, Drogon, librdkafka, redis-plus-plus, GoogleTest, Google Benchmark — and Windows has no system package manager to supply them. Decision 004 makes MSVC the primary target, so whatever supplies these has to produce MSVC-linkable artifacts.

Decision. vcpkg, integrated into CMake through its toolchain file. CMake remains the build system; vcpkg only supplies the ingredients. The two are not alternatives and the seam between them is the toolchain file.

Alternatives.

  • CMake FetchContent — genuinely sufficient for GoogleTest and Google Benchmark, which are CMake-native and self-contained. It fails on the other four, because their transitive dependencies are not CMake projects. OpenSSL is required by all four and builds via Perl and its own Configure script; during the install run for this decision, vcpkg pulled down Strawberry Perl, NASM and jom to satisfy exactly that. Reproducing this by hand through ExternalProject is the week of work the stack table already identified.
  • Conan — comparable capability and arguably better version resolution, but more configuration overhead for no gain on a single-platform single-author project.
  • Hand-built dependencies — full control, and a permanent tax every time a library updates or a clean checkout happens.

Cost. A vcpkg-shaped dependency on a third party's port maintenance: if a port breaks or lags upstream, that is now our problem. Ports also carry Windows-specific patches not merged upstream, which means the build depends on patches we do not control — though that same patching is what makes decision 004 viable at all. First build is slow; the binary cache absorbs it thereafter.


010 — Postgres claims with FOR UPDATE SKIP LOCKED, not a queue table poll or an advisory lock

Date: 2026-09-05 · Status: accepted

Context. N workers must each pull a different job from the same table, and two workers receiving the same job is the failure the project exists to prevent.

Decision. One statement: UPDATE jobs SET state='running', ... WHERE id = ( SELECT id FROM jobs WHERE state IN ('pending','scheduled') AND scheduled_for <= now ORDER BY scheduled_for, id FOR UPDATE SKIP LOCKED LIMIT 1) RETURNING .... The subquery picks and locks exactly one row and the UPDATE flips it, inside the same statement, so no other transaction can observe the row as claimable after this one has chosen it.

Alternatives.

  • Plain FOR UPDATE — correct, and it serialises every worker behind the same hottest row. N workers become one. SKIP LOCKED lets each step over rows another worker is mid-claim on.
  • SELECT then UPDATE ... WHERE state='pending' — check-then-act. Two workers read the same row, one's UPDATE affects zero rows and it retries. It works, but it burns a round trip per collision and the collision rate rises with worker count, which is exactly backwards.
  • Postgres advisory locks — a second locking system to reason about, with no advantage over row locks that the query already has.

Cost. SKIP LOCKED is Postgres-specific (MySQL 8 and Oracle have it; SQLite does not), so the JobStore interface is portable but this implementation is not. That is acceptable because the interface is the seam — InMemoryJobStore gets the same guarantee from holding one mutex across read and write, and the decision-005 contract suite proves both satisfy the same contract.


011 — count_outstanding() on the JobStore interface, rather than summing per-state counts

Date: 2026-09-05 · Status: accepted

Context. "Is the queue empty" was answered by count(Pending) + count(Scheduled) + count(Running). Under load this intermittently reported zero while work was still in flight, and a timing test failed roughly one run in ten.

Decision. A single count_outstanding() on the interface, implemented as one locked scan in the in-memory store and one SQL statement in the Postgres one.

Why. The three counts are three separate observations. A job that moves from Running to Scheduled after the Running count is taken and before the Scheduled one is missed by both, and the sum reads zero. The window is small, which is worse than obvious — it made the failure look like flakiness rather than a bug. Any predicate over multiple states has to be evaluated in one observation.

Cost. One more method on an interface every implementation must supply. Worth it: the alternative is every caller reconstructing the same race.


012 — Deduplication for a side effect lives in the same database as the effect, not in Redis

Date: 2026-09-05 · Status: accepted · Supersedes the Redis-guard approach sketched in Stage 5

Context. Stage 5's design was SET dedup:{job}:{attempt} NX in Redis before performing a side effect. A 2,000-job chaos run at a 40% kill rate produced 2,000 succeeded jobs and 1,999 side effects. One was lost.

The sequence. Worker A wins the Redis key for (job, attempt 1). Worker A is killed before its INSERT commits. The lease lapses and the reaper returns the job to Pending — without incrementing the attempt, because nobody knows whether the handler ran. Worker B claims it, computes the same key (job, attempt 1), finds it held, concludes the work is done, and skips it. The job is marked succeeded and the effect never happened.

Decision. The dedup marker moves into the same Postgres transaction as the side effect: INSERT INTO side_effects ... WHERE NOT EXISTS (SELECT 1 FROM side_effects WHERE job_id = $1). One transaction covers marker and effect, so a worker killed at any point rolls back both and the retry genuinely redoes the work.

Why no ordering of the Redis version works. Mark-then-do loses the effect on a crash in between; do-then-mark duplicates it. There is no third order, and no TTL closes the gap, because the two stores cannot commit together. The general rule: a side effect can be made exactly-once only against a store you can transact with. Against one you cannot — an email gateway, a payment API — the best available is at-least-once delivery plus a provider-side idempotency key. This is precisely why the project claims effectively-once and not exactly-once.

What Redis keeps. RedisCoordinator::claim_execution still exists and is still tested. It is a valid optimisation where work is expensive and a rare duplicate is tolerable. It is not a correctness guarantee and is deliberately off the side-effect path.

Cost. The idempotency guard is now the handler's job, not the framework's, and it only works for handlers whose effect is in a database they can transact with. That is an honest limitation and the interesting half of the answer.


013 — Static libpqxx from x64-windows-static-md, not the x64-windows DLL

Date: 2026-09-05 · Status: accepted

Context. Any test executable linking both GoogleTest and the vcpkg x64-windows libpqxx fails at link with a wall of LNK2005: "public: __cdecl std::basic_string_view<char,...>::basic_string_view(void)" already defined.

Cause. vcpkg builds that port with CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS, which exports every symbol in the binary — including the MSVC STL inline functions the library instantiated. Its import library therefore provides strong definitions of std::basic_string_view<char>'s members, which collide with the COMDATs any consumer emits for the same functions. GoogleTest emits plenty of them. It is not a debug/release CRT mismatch; RelWithDebInfo fails identically.

Decision. find_package(libpqxx CONFIG REQUIRED PATHS "${DEADMAN_PQXX_ROOT}" NO_DEFAULT_PATH) against C:/vcpkg/installed/x64-windows-static-md. A static library has no import library and exports nothing, so the duplicate COMDATs merge normally. static-md rather than static keeps the dynamic CRT (/MD), matching everything else. NO_DEFAULT_PATH is required rather than tidy: this is a manifest-mode build, so vcpkg installs its own copy into build/vcpkg_installed and the toolchain gives that prefix top priority — prepending to CMAKE_PREFIX_PATH does not outrank it.

Alternatives.

  • /FORCE:MULTIPLE — makes the error go away and suppresses genuine duplicate symbols everywhere else in the binary along with it.
  • Switching the whole project to a static triplet — rebuilds all six dependencies, and would put a static OpenSSL in the same binary as Drogon's dynamic one.

Cost. libpqxx now comes from a different triplet than everything else, which is a thing a reader has to be told. PostgreSQL itself is deliberately left to resolve normally, so libpq stays shared and only one OpenSSL ends up in any binary that also links Drogon.


014 — README masthead is a public-domain Piranesi etching, not original artwork

Date: 2026-09-07 · Status: accepted

Context. The README opened on a bare # Deadman. The sibling project themis had by then established a masthead treatment — a public-domain old-master engraving floated on pure black, monospace letterspaced type, and a single accent colour used exactly once. Matching it makes the two repositories read as one body of work, which is the only audience that matters here. The brief for this one was industrial and apocalyptic rather than classical.

Decision. Plate XVI of Piranesi's Carceri d'Invenzione — the chained pier — as a full-height left panel feathered into black, with the wordmark, rule, ember rule and taglines composited on the right at 2400 × 1260. One asset, docs/img/deadman-hero.png, 949 kB. Sourcing and processing in docs/img/CREDITS.md.

The Carceri were chosen over any generic industrial scene because the series argues the project's thesis: the apparatus is the subject, the figures in it are incidental and interchangeable, and the machinery plainly predates them and will outlast them. That is the lease heartbeat stated as an image, etched in 1749.

Alternatives.

  • Themis' cutout treatment — isolate a figure with a segmentation model and float it. Not possible here and not desirable: a Carceri plate has no figure to isolate, and the point of the image is that nothing in it separates from the apparatus. Feathered into black instead.
  • Inverting the plate to white lines on black. Reads as a blueprint or a photographic negative and throws away the fact that it is ink on paper. themis rejected the same move for the same reason.
  • Three other Carceri plates — IX (the great wheel), VI (smoking fire), XII (the sawhorse). All downloaded at full resolution and rejected on composition; the reasons are recorded in docs/img/CREDITS.md.
  • Reusing Themis' amber #FBBF24. Would have made the two repositories indistinguishable at a glance. Ember #EA580C keeps the system and separates the projects.
  • Commissioned or generated artwork — no provenance that can be stated in a sentence, and the wrong medium beside Themis.

Cost. A 949 kB decorative PNG now lives in a repository whose value is the evidence in BENCHMARKS.md, and the README leads with a picture rather than that number. More to the point: this was done while the one genuinely outstanding item on the project — Stage 4's gate having no test that observes a partition moving between two running workers mid-batch, per docs/PLAN.md — remained outstanding. Recorded here so it is not mistaken for progress.