You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A test-adequacy audit before v4.0.0 ships, after #710 removes the rneat alias. It covers item 2 of #803 ("Test adequacy"), split out so the release does not wait on the rest of that review.
The standard is CLAUDE.md's vetting section: tests assert content, survive a mutation check, and cover the path that can break.
Scope
Measure the harness coverage. The CLAUDE.md example under rule 6 ("scripts/delta/tests/ covers 2 of sv_pipeline.sbatch's 14 functions; score_caller and check_denominator … are untested") is out of date. The script now has 29 functions, scripts/delta/tests/ has 18 test files, and test_scorer_calibration.sh references score_caller. Record which functions are tested and which decisions those tests check, and update CLAUDE.md with the result.
Mutation spot checks on the tests guarding numbers that get quoted: denominator and recall checks, model fitting (gen-seq-error-model, gen-mut-model), and truth/reads agreement. Confirm each mutant was applied.
Existence-only assertions. Find tests whose only check is that a file exists, a count is non-zero, is_ok(), or exit 0, and give them a content assertion.
Log assertions read the right channel. simplelog's TerminalMode::Mixed sends INFO and WARN to stdout and only ERROR to stderr, so asserting a warning's presence or absence on stderr alone proves nothing (found in Stop warning that a breakend has no END/SVLEN (#497) #818). A grep of eidolon/tests/ on 2026-10-05 found no such test; check new ones.
Python without CI. Settled 2026-10-06 in Describe sbs96_compare.py as it is: stdlib-only, kept untested #820: scripts/delta/sbs96_compare.py stays as is, untested, with no current caller. Tests get added if it is needed again; otherwise it goes when the validation scripts are archived. Not part of this audit.
Output
Small fixes land directly. Larger findings go in their own tickets. The measured state replaces the stale example in CLAUDE.md, and new case histories go in docs/claude_engineering_audit.md.
A test-adequacy audit before v4.0.0 ships, after #710 removes the
rneatalias. It covers item 2 of #803 ("Test adequacy"), split out so the release does not wait on the rest of that review.The standard is CLAUDE.md's vetting section: tests assert content, survive a mutation check, and cover the path that can break.
Scope
scripts/delta/tests/covers 2 ofsv_pipeline.sbatch's 14 functions;score_callerandcheck_denominator… are untested") is out of date. The script now has 29 functions,scripts/delta/tests/has 18 test files, andtest_scorer_calibration.shreferencesscore_caller. Record which functions are tested and which decisions those tests check, and update CLAUDE.md with the result.#[ignore]tests pass. Runcargo test --workspace -- --ignored. Each must pass; a known-failing target does not belong there (gate2_realigned_dup is red on develop and main — the gate's denominator, not the duplication #582).gen-seq-error-model,gen-mut-model), and truth/reads agreement. Confirm each mutant was applied.is_ok(), or exit 0, and give them a content assertion.TerminalMode::Mixedsends INFO and WARN to stdout and only ERROR to stderr, so asserting a warning's presence or absence on stderr alone proves nothing (found in Stop warning that a breakend has no END/SVLEN (#497) #818). A grep ofeidolon/tests/on 2026-10-05 found no such test; check new ones.Python without CI.Settled 2026-10-06 in Describe sbs96_compare.py as it is: stdlib-only, kept untested #820:scripts/delta/sbs96_compare.pystays as is, untested, with no current caller. Tests get added if it is needed again; otherwise it goes when the validation scripts are archived. Not part of this audit.Output
Small fixes land directly. Larger findings go in their own tickets. The measured state replaces the stale example in CLAUDE.md, and new case histories go in
docs/claude_engineering_audit.md.