Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions docs/development/review-checklist.md
Original file line number Diff line number Diff line change
Expand Up @@ -167,3 +167,7 @@ And then note where the fix landed. Someone hit the audience gap, felt it, and b
## Reading a mutation's result

46. **An arm's verdict is its test TOTAL compared against BASE, not the presence of the word `failed` — a suite that fails to COMPILE reports zero failures and simply stops contributing tests.** Jest counts a collection failure at the *suite* level only: `Test Suites:` names it, and `Tests:` is a sum over the suites that actually ran, so a corrupted tree prints a line with no failures on it and a total that is quietly smaller. The number stays plausible, which is what makes this worse than the empty case our existing discipline already covers — `Tests: 0 total` or a missing `Tests:` line announces itself, and `875 passed, 875 total` does not. Worse still, **the suite you are watching is usually not one of the casualties.** An invariant suite that reads the module as a *string* (`read('../landing/V2LandingPage.tsx')`) never imports it, so it collects, runs and passes over the corruption; the suites that die are the ones that `import` the module, and they are elsewhere in the tree. So the arm you are pointed at goes green while a tenth of the run vanishes. The defence costs one line: **have every arm print its own total and flag any total ≠ BASE's**, and treat a mismatch as a broken instrument rather than a result. Grepping `^Tests:` alone is not enough — that is the line engineered to look fine. The generalisation is worth holding onto beyond jest: **any per-test aggregate is blind to a suite that produced no tests**, which is also how a coverage threshold, a snapshot count or a "N tests added" check can be satisfied by a tree that does not build. *(Earned: 2026-09-29, TASK-209, during the #2022 gate, and it had already been live in two of that day's gate harnesses before anyone noticed. Reproduced at main `abe19fe0` with `npx jest` in `frontend/`: BASE is `Test Suites: 120 passed, 120 total` and `Tests: 1108 passed, 1108 total`; corrupting **one** line — line 48 of `V2LandingPage.tsx`, a broken JSX tag, which is what a mis-quoted mutation regex produces — gives `Test Suites: 5 failed, 115 passed, 120 total` beside `Tests: 1005 passed, 1005 total`. Zero failed on the line a harness reads, and **103 tests absent**. Name the invocation whenever you quote a pair like this, because the totals are a property of the filter: the same corruption under `npx jest src/v2` reads `5 failed, 83 passed, 88 total` and `875 passed, 875 total` — different totals, the **same 103**, because the absent count belongs to the corrupted module’s importers and not to the filter that happened to be running. The five are exactly the suites that import the module — `landingBarInset`, `landingInstallCopy`, `landingHeroContent`, `V2LandingPage.stats`, `V2Login` — each a `SyntaxError` at collection; `v2-layout-invariants`, the suite nearly every CSS and TSX arm is aimed at, is **not** among them, because it reads the file rather than importing it. The reviewer's own harness printed only the `Tests:` line, so the first run of this corruption was recorded as `870 passed, 870 total` and read as a clean green; the arm that produced it was reported as a real result for several minutes before the suite-level line was looked at. Its own entry rather than a rider on 45, by rule 41's test: 45 sends you to confirm the edit's SITE, and every one of those checks passes here — the site was right, printed, and inside the intended braces — while the run it produced was not a run.)*

## Reaching a mutation's site

47. **A mutation has to be sited in a statement the arm's own trace reaches, or its green is unreadable — rule 45's file:line is necessary and not sufficient.** The mutation applies, the occurrence guard confirms it applied, the site the ledger prints is the site the arm names, and the suite still stays green, because the statement that was mutated is never executed on the path the named arm takes. The green is then read as *this arm is unpinned* — a finding filed against an arm that was correct, which is rule 45's error direction entered from a different door, and worse here because a ledger that reports "20 mutants, 0 survivors, 0 equivalent" cannot distinguish this from a well-pinned arm. Three shapes were measured in one afternoon: code appended after a `return` (unreachable by construction), an append that changes nothing (`void 0;`), and a mutation aimed at the right function whose edited span belongs to a branch the fixture's own values never select. The requirement this adds to rule 45: a mutant must change a branch or a value **on the path the named arm takes**, so the arm's trace reaches the edit — print the site, and confirm the arm's path includes it. The consequence for equivalence claims is the load-bearing half: **"nothing died because nothing ran" is not an equivalence**, and an equivalence is a measured statement about the code, so a zero-kill mutant has to be investigated before it is recorded as anything — **a survivor is a question, not a classification**. **A survivor stays unclassified until it is investigated, because a reached mutant can survive for a reason that is not the instrument's fault.** Three outcomes, not two. If the edit cannot run on the arm's path — a statement after a `return`, a value nothing reads — the ledger has an `INSTRUMENT DEFECT`, and the reachability that was missing goes in the record. If it runs, changes the result the arm's path computes, and the suite stays green anyway, the instrument is not broken at all: that is a **test gap**, the arm's assertion being weaker than the behaviour it names, and filing it as an instrument defect disposes of a finding against the code. `result = n + 1` mutated to `n + 2` under an `assert(result > 1)` executes, changes the value, and passes both ways — nothing about that mutant is unreachable or inert. Only the third outcome is a proven equivalence, and it needs the most evidence: the mutant shown *reached* **and** shown semantically unchanged, so an equivalent is a measured statement about the code rather than a claim about reachability. Note also that a ledger's own anchor-count assertion does not catch any of this: the anchors match, which is why this is a sibling of 45 rather than a rider on the mutation-discipline rules — every check those send you to run passes here. *(Earned: 2026-09-29, TASK-172 §10 step 6 (#2035 at `a48fdfbe`). The first cut of the arm asserting that a failed provider revoke is reported as a failed removal appended its mutation after the `return` that had already answered the question, and the first cut of the arm asserting that unreadable material refuses rather than deletes appended `void 0;`; both then sat in a 20-mutant ledger as survivors, and both passed rule 45's site check, which printed `connectionRemovalService.ts:226 in export const removeConnection` for work that could not run. Rewritten to change the value the branch returns and to hoist the read above its guard, both died on the arm that names them. Nothing distinguished the two greens from a genuinely unpinned arm, including the honest reading that a survivor is a question rather than an equivalence — that reading saved this case only because the survivor was asked about. Its own rule rather than a rider on 45: 45 asks WHERE the edit landed, this asks whether the arm can reach it, and each passes while the other fails.)*
Loading