Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/development/review-checklist.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,7 +166,7 @@ And then note where the fix landed. Someone hit the audience gap, felt it, and b

## Reading a mutation's result

46. **An arm's verdict is its test TOTAL compared against BASE, not the presence of the word `failed` — a suite that fails to COMPILE reports zero failures and simply stops contributing tests.** Jest counts a collection failure at the *suite* level only: `Test Suites:` names it, and `Tests:` is a sum over the suites that actually ran, so a corrupted tree prints a line with no failures on it and a total that is quietly smaller. The number stays plausible, which is what makes this worse than the empty case our existing discipline already covers — `Tests: 0 total` or a missing `Tests:` line announces itself, and `875 passed, 875 total` does not. Worse still, **the suite you are watching is usually not one of the casualties.** An invariant suite that reads the module as a *string* (`read('../landing/V2LandingPage.tsx')`) never imports it, so it collects, runs and passes over the corruption; the suites that die are the ones that `import` the module, and they are elsewhere in the tree. So the arm you are pointed at goes green while a tenth of the run vanishes. The defence costs one line: **have every arm print its own total and flag any total ≠ BASE's**, and treat a mismatch as a broken instrument rather than a result. Grepping `^Tests:` alone is not enough — that is the line engineered to look fine. The generalisation is worth holding onto beyond jest: **any per-test aggregate is blind to a suite that produced no tests**, which is also how a coverage threshold, a snapshot count or a "N tests added" check can be satisfied by a tree that does not build. *(Earned: 2026-09-29, TASK-209, during the #2022 gate, and it had already been live in two of that day's gate harnesses before anyone noticed. Reproduced at main `abe19fe0` with `npx jest` in `frontend/`: BASE is `Test Suites: 120 passed, 120 total` and `Tests: 1108 passed, 1108 total`; corrupting **one** line — line 48 of `V2LandingPage.tsx`, a broken JSX tag, which is what a mis-quoted mutation regex produces — gives `Test Suites: 5 failed, 115 passed, 120 total` beside `Tests: 1005 passed, 1005 total`. Zero failed on the line a harness reads, and **103 tests absent**. Name the invocation whenever you quote a pair like this, because the totals are a property of the filter: the same corruption under `npx jest src/v2` reads `5 failed, 83 passed, 88 total` and `875 passed, 875 total` — different totals, the **same 103**, because the absent count belongs to the corrupted module’s importers and not to the filter that happened to be running. The five are exactly the suites that import the module — `landingBarInset`, `landingInstallCopy`, `landingHeroContent`, `V2LandingPage.stats`, `V2Login` — each a `SyntaxError` at collection; `v2-layout-invariants`, the suite nearly every CSS and TSX arm is aimed at, is **not** among them, because it reads the file rather than importing it. The reviewer's own harness printed only the `Tests:` line, so the first run of this corruption was recorded as `870 passed, 870 total` and read as a clean green; the arm that produced it was reported as a real result for several minutes before the suite-level line was looked at. Its own entry rather than a rider on 45, by rule 41's test: 45 sends you to confirm the edit's SITE, and every one of those checks passes here — the site was right, printed, and inside the intended braces — while the run it produced was not a run.)*
46. **An arm's verdict is its test TOTAL compared against BASE, not the presence of the word `failed` — a suite that fails to COMPILE reports zero failures and simply stops contributing tests.** Jest counts a collection failure at the *suite* level only: `Test Suites:` names it, and `Tests:` is a sum over the suites that actually ran, so a corrupted tree prints a line with no failures on it and a total that is quietly smaller. The number stays plausible, which is what makes this worse than the empty case our existing discipline already covers — `Tests: 0 total` or a missing `Tests:` line announces itself, and `875 passed, 875 total` does not. Worse still, **the suite you are watching is usually not one of the casualties.** An invariant suite that reads the module as a *string* (`read('../landing/V2LandingPage.tsx')`) never imports it, so it collects, runs and passes over the corruption; the suites that die are the ones that `import` the module, and they are elsewhere in the tree. So the arm you are pointed at goes green while a tenth of the run vanishes. The defence costs one line: **have every arm print its own total and flag any total ≠ BASE's** — then, before reporting anything, **find out which suites' counts moved** (`jest --json`, compare each result's `assertionResults.length` against the base run). A mismatch is a signal to investigate, **not by itself evidence of a corrupted tree**: a suite that throws at MODULE scope — a top-level `const` calling a helper that raises when what it reads is absent — fails to collect for the same reason a syntax error does, so a *genuine* red surfaces as absent tests too, and the `Tests:` line cannot tell the two apart. Both shapes can occur in **one run**, so the question is never which of the two it was. Two things the summary will not show you. The **suite** total does not move — a suite that fails to collect is still counted, so `88 total` reads the same before and after and only the `Tests:` total drops; do not look for a shrinking suite count. And the comparator is the **tree** under test, not the branch name: the base figure has to come from the unmutated head you are mutating, because a head that adds a test reads one higher than `main` and every arm compared against `main` would then look like a missing test. Reporting either one as the other is the failure this rule exists to stop: call the arm broken and you dismiss a real finding; call it a result and you publish a green. Grepping `^Tests:` alone is not enough — that is the line engineered to look fine. The generalisation is worth holding onto beyond jest: **any per-test aggregate is blind to a suite that produced no tests**, which is also how a coverage threshold, a snapshot count or a "N tests added" check can be satisfied by a tree that does not build. *(Earned: 2026-09-29, TASK-209, during the #2022 gate, and it had already been live in two of that day's gate harnesses before anyone noticed. Reproduced at main `abe19fe0` with `npx jest` in `frontend/`: BASE is `Test Suites: 120 passed, 120 total` and `Tests: 1108 passed, 1108 total`; corrupting **one** line — line 48 of `V2LandingPage.tsx`, a broken JSX tag, which is what a mis-quoted mutation regex produces — gives `Test Suites: 5 failed, 115 passed, 120 total` beside `Tests: 1005 passed, 1005 total`. Zero failed on the line a harness reads, and **103 tests absent**. Name the invocation whenever you quote a pair like this, because the totals are a property of the filter: the same corruption under `npx jest src/v2` reads `5 failed, 83 passed, 88 total` and `875 passed, 875 total` — different totals, the **same 103**, because the absent count belongs to the corrupted module’s importers and not to the filter that happened to be running. The five are exactly the suites that import the module — `landingBarInset`, `landingInstallCopy`, `landingHeroContent`, `V2LandingPage.stats`, `V2Login` — each a `SyntaxError` at collection; `v2-layout-invariants`, the suite nearly every CSS and TSX arm is aimed at, is **not** among them, because it reads the file rather than importing it. The reviewer's own harness printed only the `Tests:` line, so the first run of this corruption was recorded as `870 passed, 870 total` and read as a clean green; the arm that produced it was reported as a real result for several minutes before the suite-level line was looked at. Its own entry rather than a rider on 45, by rule 41's test: 45 sends you to confirm the edit's SITE, and every one of those checks passes here — the site was right, printed, and inside the intended braces — while the run it produced was not a run. The investigate-before-reporting clause was earned separately, on the very next gate — while gating **#2038 at `7271de4c`, which is TASK-213’s PR**, and filed as TASK-214: an arm renaming `@media (max-width: 680px)` to `681px` moved the total 981 → 978, and the cause was not corruption — `landingAnchorInsets.test.ts` builds its fixture at module scope (`const PHONE = mediaBlock(BARE, '@media (max-width: 680px)')`) with a private helper that THROWS when the at-rule is absent, so that suite's three tests left the total while a second suite's ten assertions reddened normally. Both reds were real. The first draft of this rule said to treat the mismatch as a broken instrument, which would have dismissed one of them.)*

## Reaching a mutation's site

Expand Down
Loading