Append-only. Never edit or delete a past entry — supersede it with a new one that references the old by number and says what changed.
Every entry records the alternatives that lost and why. That reasoning is the part interviewers ask about, and it is much harder to reconstruct months later than the code is.
Format:
## NNN — Title
**Date:** YYYY-MM-DD · **Status:** accepted | superseded by NNN
**Context.** What forced a choice.
**Decision.** What we chose.
**Alternatives.** What lost, and why.
**Cost.** What this makes harder. Every real decision has one.
Date: 2026-09-02 · Status: accepted
Context. Needed one portfolio project that demonstrates API design, low-level design, concurrency, and system-design vocabulary for SDE-1 interviews at tier-1 Indian product companies — in C++, without requiring a frontend stack.
Decision. A distributed job scheduler with at-least-once delivery and idempotent execution.
Alternatives.
- Seat reservation system — the strongest concurrency-correctness story, but the scheduler exercises worker pools and graceful shutdown more directly, which suits the concurrency material being read alongside it.
- Live location tracking — Kafka-heavy but thinner on low-level design.
- Finishing
matchbook— excellent depth credential, but it is a single-threaded in-memory library with no service surface, and the trading domain reads as quant to product-company interviewers. - Redis clone — highly legible, but building infrastructure rather than using it demonstrates a narrower slice of what SDE-1 loops actually probe.
Cost. Requires standing up three external dependencies (Postgres, Redis, Kafka), which is real setup time before any interesting code runs.
Date: 2026-09-02 · Status: accepted
Context. Both systems can technically hold job state.
Decision. Postgres is the source of truth for job records, state, and attempt history. Kafka carries work to workers and nothing more.
Alternatives. Kafka-as-database using compacted topics — fashionable, but it makes "show me every failed job from Tuesday" hard, and that query is the kind of thing a real operator needs.
Cost. Two systems to keep consistent, and a window where a message exists in Kafka but its row is not yet committed in Postgres. Stage 4 has to address this explicitly.
Date: 2026-09-02 · Status: accepted
Context. The offset commit point decides the delivery guarantee.
Decision. Commit after the handler completes. This gives at-least-once delivery: a crash between execution and commit means the job is redelivered.
Alternatives. Committing on receipt gives at-most-once, which loses jobs on a crash — unacceptable for a scheduler whose entire promise is that work is not lost.
Cost. Duplicates are now possible by design, which is why deduplication in Stage 5 is mandatory rather than optional. The project therefore claims effectively-once, never exactly-once.
Date: 2026-09-02 · Status: accepted
Context. Stages 3–5 are process-lifecycle exercises: kill -9 a worker and
have another reclaim its lease, run three worker processes through a consumer
group rebalance, script a harness that kills workers and drops Redis under load.
Every dependency in the stack — librdkafka, Drogon, redis-plus-plus, libpqxx —
is Linux-first with Windows as a port. That argued for developing in WSL2 with
clang or gcc.
Decision. Stay on MSVC/Windows as the primary target, as originally written.
Dependencies come from vcpkg. The chaos harness is Windows-native — taskkill /F
for kills, PowerShell for orchestration — with Docker Compose supplying Postgres,
Redis and Kafka as containers.
Alternatives.
- WSL2 with clang or gcc as primary — better dependency support and a chaos harness that matches every tutorial and Stack Overflow answer in this space. Lost because it splits the development environment from the machine the work actually happens on, and a second toolchain is its own ongoing tax.
- Linux only, no Windows support — simplest of all, but abandons the primary development machine entirely.
Cost. vcpkg has to successfully produce librdkafka and Drogon on MSVC; if
either fights back, that is time spent on packaging rather than on the scheduler.
The chaos scripts will not be copy-pasteable from prior art, and process-kill
semantics differ from POSIX signals — taskkill /F is closer to SIGKILL than
to anything catchable, so graceful-shutdown testing needs its own separate path.
Revisit this decision if Stage 3 or Stage 4 stalls on toolchain problems rather
than on design problems.
Date: 2026-09-02 · Status: accepted
Context. Testing only against an in-memory JobStore tests a hash map behind
an interface. The failure modes that matter in a scheduler are Postgres-specific:
claim semantics under concurrent workers, the isolation level a state transition
runs at, and whether the idempotency key is enforced by a real unique constraint
or by a read-then-write that only looks atomic.
Decision. Write the JobStore tests once as a parameterized GoogleTest
fixture and run the same suite against both InMemoryJobStore and
PostgresJobStore. The in-memory run happens on every local build; the Postgres
run happens in CI against a service container, and locally when Docker is up.
JobState transitions, RetryPolicy and WorkerPool shutdown stay pure unit
tests with no infrastructure.
Alternatives.
- In-memory only — fast and infra-free, but leaves the concurrency-correctness claims of the persistence layer entirely unverified.
- Real Postgres for everything — highest fidelity, but every test then needs
Docker, the suite slows to the point where it stops being run, and it removes
most of the reason for
JobStoreto be an interface at all.
Cost. One CI workflow to maintain, and the contract suite constrains the
interface: anything InMemoryJobStore cannot honour cannot go in JobStore.
That constraint is mostly a feature, but it will occasionally be inconvenient
when a Postgres-only capability looks tempting.
Date: 2026-09-02 · Status: accepted
Context. Needed a GitHub repository name before a remote could be added.
Decision. Deadman. The lease is a deadman switch — it lapses when the
worker holding it stops renewing, and that lapse is the entire product. The name
encodes the core mechanism. The repository description carries the explanation:
"Distributed job scheduler in C++20. Kill a worker mid-job — the job still runs
exactly once."
Alternatives.
deadman-scheduler— self-describing without the README, but the second half carries nothing the description field cannot.deadman-switch— names the mechanism outright, at the cost of obscuring that the product is a job scheduler.
Cost. The bare word is not searchable and collides with unrelated projects. This does not matter for a portfolio repository reached from a resume link.
Date: 2026-09-02 · Status: accepted
Context. Formatting arguments are worth zero interview points and cost real time. Settling layout before any code exists avoids a reformatting commit later that buries the first real diffs.
Decision. A .clang-format at the repository root: BasedOnStyle: LLVM,
IndentWidth: 4, ColumnLimit: 100, PointerAlignment: Left,
AllowShortFunctionsOnASingleLine: Empty, SeparateDefinitionBlocks: Always.
LLVM is the base because it is the default clang-format is best tested against;
4-space indent and left-aligned pointers match MSVC-world convention; 100 columns
because 80 becomes cramped once namespaced types and std::chrono durations
appear in signatures. The binary ships with Visual Studio 2022 at
VC\Tools\Llvm\x64\bin\clang-format.exe (19.1.1), so no extra install.
Alternatives.
- No config at all — every contributor's editor imposes its own, and the first real refactor produces a diff that is 90% whitespace.
- Google style base — 2-space indent, which reads as cramped in a codebase with deeply namespaced types.
Cost. .clang-format governs layout only. The naming conventions in
CLAUDE.md — snake_case functions, PascalCase types, member_ trailing
underscore — are clang-tidy's domain and remain unenforced by tooling, so they
stay a review-discipline matter unless a .clang-tidy is added later.
Date: 2026-09-02 · Status: accepted
Context. The GitHub repository was initialised with an MIT LICENSE before the local history was pushed. Accepting or replacing it was a real choice, not a formality.
Decision. Keep MIT.
Alternatives. Apache-2.0 offers an explicit patent grant, which matters for projects expecting outside contributors or corporate adoption. Neither applies to a portfolio repository, and MIT is shorter and more universally recognised by the people who will actually read it.
Cost. No patent grant. Irrelevant at this scale, and changing it later costs nothing while the project has a single author.
Date: 2026-09-02 · Status: accepted
Context. Six third-party libraries are needed — libpqxx, Drogon, librdkafka, redis-plus-plus, GoogleTest, Google Benchmark — and Windows has no system package manager to supply them. Decision 004 makes MSVC the primary target, so whatever supplies these has to produce MSVC-linkable artifacts.
Decision. vcpkg, integrated into CMake through its toolchain file. CMake remains the build system; vcpkg only supplies the ingredients. The two are not alternatives and the seam between them is the toolchain file.
Alternatives.
- CMake
FetchContent— genuinely sufficient for GoogleTest and Google Benchmark, which are CMake-native and self-contained. It fails on the other four, because their transitive dependencies are not CMake projects. OpenSSL is required by all four and builds via Perl and its ownConfigurescript; during the install run for this decision, vcpkg pulled down Strawberry Perl, NASM and jom to satisfy exactly that. Reproducing this by hand throughExternalProjectis the week of work the stack table already identified. - Conan — comparable capability and arguably better version resolution, but more configuration overhead for no gain on a single-platform single-author project.
- Hand-built dependencies — full control, and a permanent tax every time a library updates or a clean checkout happens.
Cost. A vcpkg-shaped dependency on a third party's port maintenance: if a port breaks or lags upstream, that is now our problem. Ports also carry Windows-specific patches not merged upstream, which means the build depends on patches we do not control — though that same patching is what makes decision 004 viable at all. First build is slow; the binary cache absorbs it thereafter.
Date: 2026-09-05 · Status: accepted
Context. N workers must each pull a different job from the same table, and two workers receiving the same job is the failure the project exists to prevent.
Decision. One statement: UPDATE jobs SET state='running', ... WHERE id = ( SELECT id FROM jobs WHERE state IN ('pending','scheduled') AND scheduled_for <= now ORDER BY scheduled_for, id FOR UPDATE SKIP LOCKED LIMIT 1) RETURNING ....
The subquery picks and locks exactly one row and the UPDATE flips it, inside the
same statement, so no other transaction can observe the row as claimable after
this one has chosen it.
Alternatives.
- Plain
FOR UPDATE— correct, and it serialises every worker behind the same hottest row. N workers become one.SKIP LOCKEDlets each step over rows another worker is mid-claim on. SELECTthenUPDATE ... WHERE state='pending'— check-then-act. Two workers read the same row, one's UPDATE affects zero rows and it retries. It works, but it burns a round trip per collision and the collision rate rises with worker count, which is exactly backwards.- Postgres advisory locks — a second locking system to reason about, with no advantage over row locks that the query already has.
Cost. SKIP LOCKED is Postgres-specific (MySQL 8 and Oracle have it; SQLite
does not), so the JobStore interface is portable but this implementation is not.
That is acceptable because the interface is the seam — InMemoryJobStore gets
the same guarantee from holding one mutex across read and write, and the
decision-005 contract suite proves both satisfy the same contract.
Date: 2026-09-05 · Status: accepted
Context. "Is the queue empty" was answered by count(Pending) + count(Scheduled) + count(Running). Under load this intermittently reported zero
while work was still in flight, and a timing test failed roughly one run in ten.
Decision. A single count_outstanding() on the interface, implemented as one
locked scan in the in-memory store and one SQL statement in the Postgres one.
Why. The three counts are three separate observations. A job that moves from Running to Scheduled after the Running count is taken and before the Scheduled one is missed by both, and the sum reads zero. The window is small, which is worse than obvious — it made the failure look like flakiness rather than a bug. Any predicate over multiple states has to be evaluated in one observation.
Cost. One more method on an interface every implementation must supply. Worth it: the alternative is every caller reconstructing the same race.
Date: 2026-09-05 · Status: accepted · Supersedes the Redis-guard approach sketched in Stage 5
Context. Stage 5's design was SET dedup:{job}:{attempt} NX in Redis before
performing a side effect. A 2,000-job chaos run at a 40% kill rate produced 2,000
succeeded jobs and 1,999 side effects. One was lost.
The sequence. Worker A wins the Redis key for (job, attempt 1). Worker A is killed before its INSERT commits. The lease lapses and the reaper returns the job to Pending — without incrementing the attempt, because nobody knows whether the handler ran. Worker B claims it, computes the same key (job, attempt 1), finds it held, concludes the work is done, and skips it. The job is marked succeeded and the effect never happened.
Decision. The dedup marker moves into the same Postgres transaction as the
side effect: INSERT INTO side_effects ... WHERE NOT EXISTS (SELECT 1 FROM side_effects WHERE job_id = $1). One transaction covers marker and effect, so a
worker killed at any point rolls back both and the retry genuinely redoes the
work.
Why no ordering of the Redis version works. Mark-then-do loses the effect on a crash in between; do-then-mark duplicates it. There is no third order, and no TTL closes the gap, because the two stores cannot commit together. The general rule: a side effect can be made exactly-once only against a store you can transact with. Against one you cannot — an email gateway, a payment API — the best available is at-least-once delivery plus a provider-side idempotency key. This is precisely why the project claims effectively-once and not exactly-once.
What Redis keeps. RedisCoordinator::claim_execution still exists and is
still tested. It is a valid optimisation where work is expensive and a rare
duplicate is tolerable. It is not a correctness guarantee and is deliberately off
the side-effect path.
Cost. The idempotency guard is now the handler's job, not the framework's, and it only works for handlers whose effect is in a database they can transact with. That is an honest limitation and the interesting half of the answer.
Date: 2026-09-05 · Status: accepted
Context. Any test executable linking both GoogleTest and the vcpkg
x64-windows libpqxx fails at link with a wall of LNK2005: "public: __cdecl std::basic_string_view<char,...>::basic_string_view(void)" already defined.
Cause. vcpkg builds that port with CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS, which
exports every symbol in the binary — including the MSVC STL inline functions
the library instantiated. Its import library therefore provides strong
definitions of std::basic_string_view<char>'s members, which collide with the
COMDATs any consumer emits for the same functions. GoogleTest emits plenty of
them. It is not a debug/release CRT mismatch; RelWithDebInfo fails identically.
Decision. find_package(libpqxx CONFIG REQUIRED PATHS "${DEADMAN_PQXX_ROOT}" NO_DEFAULT_PATH) against
C:/vcpkg/installed/x64-windows-static-md. A static library has no import
library and exports nothing, so the duplicate COMDATs merge normally.
static-md rather than static keeps the dynamic CRT (/MD), matching
everything else. NO_DEFAULT_PATH is required rather than tidy: this is a
manifest-mode build, so vcpkg installs its own copy into
build/vcpkg_installed and the toolchain gives that prefix top priority —
prepending to CMAKE_PREFIX_PATH does not outrank it.
Alternatives.
- /FORCE:MULTIPLE — makes the error go away and suppresses genuine duplicate symbols everywhere else in the binary along with it.
- Switching the whole project to a static triplet — rebuilds all six dependencies, and would put a static OpenSSL in the same binary as Drogon's dynamic one.
Cost. libpqxx now comes from a different triplet than everything else, which is a thing a reader has to be told. PostgreSQL itself is deliberately left to resolve normally, so libpq stays shared and only one OpenSSL ends up in any binary that also links Drogon.
Date: 2026-09-07 · Status: accepted
Context. The README opened on a bare # Deadman. The sibling project
themis had by then established a masthead treatment — a public-domain
old-master engraving floated on pure black, monospace letterspaced type, and a
single accent colour used exactly once. Matching it makes the two repositories
read as one body of work, which is the only audience that matters here. The
brief for this one was industrial and apocalyptic rather than classical.
Decision. Plate XVI of Piranesi's Carceri d'Invenzione — the chained pier
— as a full-height left panel feathered into black, with the wordmark, rule,
ember rule and taglines composited on the right at 2400 × 1260. One asset,
docs/img/deadman-hero.png, 949 kB. Sourcing and processing in
docs/img/CREDITS.md.
The Carceri were chosen over any generic industrial scene because the series argues the project's thesis: the apparatus is the subject, the figures in it are incidental and interchangeable, and the machinery plainly predates them and will outlast them. That is the lease heartbeat stated as an image, etched in 1749.
Alternatives.
- Themis' cutout treatment — isolate a figure with a segmentation model and float it. Not possible here and not desirable: a Carceri plate has no figure to isolate, and the point of the image is that nothing in it separates from the apparatus. Feathered into black instead.
- Inverting the plate to white lines on black. Reads as a blueprint or a
photographic negative and throws away the fact that it is ink on paper.
themisrejected the same move for the same reason. - Three other Carceri plates — IX (the great wheel), VI (smoking fire), XII
(the sawhorse). All downloaded at full resolution and rejected on composition;
the reasons are recorded in
docs/img/CREDITS.md. - Reusing Themis' amber
#FBBF24. Would have made the two repositories indistinguishable at a glance. Ember#EA580Ckeeps the system and separates the projects. - Commissioned or generated artwork — no provenance that can be stated in a sentence, and the wrong medium beside Themis.
Cost. A 949 kB decorative PNG now lives in a repository whose value is the
evidence in BENCHMARKS.md, and the README leads with a picture rather than
that number. More to the point: this was done while the one genuinely
outstanding item on the project — Stage 4's gate having no test that observes a
partition moving between two running workers mid-batch, per docs/PLAN.md —
remained outstanding. Recorded here so it is not mistaken for progress.