From 9c191741da548a5cb8e2fb861f7aca2631b68e6a Mon Sep 17 00:00:00 2001 From: Lily Shen <115414357+lilyshen0722@users.noreply.github.com> Date: Sat, 22 Aug 2026 13:47:33 -0700 Subject: [PATCH 1/2] =?UTF-8?q?docs(ax):=20entry=2043=20=E2=80=94=20a=20kn?= =?UTF-8?q?own-red=20suite=20absorbs=20the=20next=20failure?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit @sprint-review (56978) named the cost of "ignore those, they're the Node 26 thing": a permanently-red suite does not just lose its own signal, it absorbs any new failure landing in it. Measured on main rather than argued, and it is worse than the framing: baseline Tests: 9 failed, 9 total add a trivially-passing test Tests: 9 failed, 1 passed, 10 total simulate a real boot regression Tests: 9 failed, 9 total The regression's output is byte-identical to the healthy baseline — the suite already dies at import, so a second reason to die is invisible. An unrelated passing test, meanwhile, does move the number. The count is therefore not weakly informative, it is anti-informative: it moves for changes that do not matter and holds still for the one that does. "Still 9, same as always" is watching the one statistic guaranteed not to respond. Scope stated precisely, because overclaiming this would be its own version of the error: the blindness is LOCAL. CI is Node 22, the suite passes there, and the simulated regression would have gone red in CI. No real boot bug reaches main through this. The exposure is a person or agent deciding something from a local run between pushes, which is what local runs are for. Distinct from entry 42: there the count sits in the verdict slot and means nothing; here the count is real and the suite genuinely ran, but the baseline is non-zero so a delta has nowhere to appear. Co-Authored-By: Claude Opus 5 --- docs/development/agent-experience-audit.md | 63 ++++++++++++++++++++++ 1 file changed, 63 insertions(+) diff --git a/docs/development/agent-experience-audit.md b/docs/development/agent-experience-audit.md index c12c72a19..5ad18a698 100644 --- a/docs/development/agent-experience-audit.md +++ b/docs/development/agent-experience-audit.md @@ -2573,3 +2573,66 @@ agent path since it was written. - Companion rule, on the method that missed it: reviewer-checklist rule 17 — a mutation proves a term matters to the suite, not that the suite's shape is real. +## 44. A known-red suite absorbs the next failure (2026-08-22, pod-architect + sprint-review) + +> Renumbered 43 → 44 on rebase. #1164's dual-auth entry took 42 on main while +> this sat open, pushing #1142 to 43. **This branch is cut from main, not +> stacked on #1142**, so 44 is correct only if #1142 lands first; merge it the +> other way round and 43 is a hole. #1122 (39) and #1132 (40) are still open +> and reserve those numbers — if they merge in another order, renumber this +> one rather than them. + +`backend/__tests__/unit/server.test.js` fails 9/9 on `main` locally, and has +for a while. The cause is environmental and known: local Node 26 versus CI's +Node 22, where `jsonwebtoken` → `jws` → `jwa` → `buffer-equal-constant-time` +throws at import, taking `server.ts:9` with it. CI is green; nothing is +actually broken. The standing advice is "ignore those, they're the Node 26 +thing." + +@sprint-review (56978) named the cost of that advice: a known-red suite does +not merely lose its own signal, it **absorbs any new failure that lands in +it**. Measured on `main`, and the result is worse than the framing: + +``` +baseline Tests: 9 failed, 9 total +add one trivially-passing test Tests: 9 failed, 1 passed, 10 total +simulate a real boot regression Tests: 9 failed, 9 total + (throw at the top of config/db-pg.ts, + which server.ts requires) +``` + +**The regression's output is byte-identical to the healthy baseline.** The +suite already dies at import, so a second reason to die changes nothing you +can see. Meanwhile adding an unrelated passing test *does* move the number. + +So the count is not weakly informative here — it is anti-informative. It moves +for changes that do not matter and holds still for the one that does. A +developer watching "still 9, same as always" is watching the one statistic +guaranteed not to respond. + +**Scope, stated precisely.** This is a *local* blindness. CI runs Node 22, the +suite passes there, and the simulated regression would have gone red in CI. +Nothing about this lets a real boot bug reach `main` on its own. The exposure +is a person or agent using a local run to decide something — "I didn't break +anything, the failures are the usual ones" — which is exactly how local runs +get used between pushes. + +Related to entry 42 but distinct. There the count sits in the verdict slot and +means nothing (`0 total`). Here the count is real, the suite genuinely ran, and +the failures are genuine — the defect is that the **baseline is non-zero**, so +a delta has nowhere to show up. + +**Rules earned.** + +- **A permanently-red suite is a hole in coverage the size of that suite, not + a nuisance.** "Known failures" is a status, and statuses need an expiry or a + quarantine. Either fix the environment, pin the runtime locally, or mark the + suite skipped-with-reason so a *new* failure in it cannot hide behind an old + one. +- **Compare counts against a recorded baseline, never against memory.** "Still + 9" is only meaningful if 9 was written down and the suite's test count has + not changed. Both of those move. +- **Do not accept an unchanged failure count as evidence you broke nothing.** + On any suite that fails at import, it is evidence of nothing at all. Run the + specific suites your change touches, or push and read CI, which is the tier + that is actually measuring. From febc2f707962ee4169d1c053d8e09fc0f9f444fd Mon Sep 17 00:00:00 2001 From: Lily Shen <115414357+lilyshen0722@users.noreply.github.com> Date: Sat, 22 Aug 2026 15:34:41 -0700 Subject: [PATCH 2/2] docs(ax): entry 43's reassurance did not hold for its own incident MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit @sprint-review (57015) supplied the clause I dropped when filing this: combined with those branches having no CI at all, the boot bug had precisely zero instruments pointed at it. I had scoped the entry as a local-only blindness on the grounds that CI runs Node 22 and would go red. True in general, and false for the case that motivated the entry. #1123 — which dropped the branches:[main] filter and so gave stacked PRs any CI at all — merged at 15:20:27Z. The boot bug lived on a stacked branch before that, where pull_request workflows did not dispatch. So in that window the local suite could not show a new boot failure (this entry) and the remote suite was not running (entry 41, mechanism 1). Each entry's stated mitigation is the other entry's failure mode, and the two coincided on one defect. The correction is the sentence "nothing about this lets a real boot bug reach main on its own", which I have removed. "CI would catch it" is not a property of the repo. It is a property of whether CI ran on that PR — which is exactly what entry 41 exists to say you cannot assume, and I assumed it one entry later. Co-Authored-By: Claude Opus 5 --- docs/development/agent-experience-audit.md | 23 ++++++++++++++++++---- 1 file changed, 19 insertions(+), 4 deletions(-) diff --git a/docs/development/agent-experience-audit.md b/docs/development/agent-experience-audit.md index 5ad18a698..6d7fda0db 100644 --- a/docs/development/agent-experience-audit.md +++ b/docs/development/agent-experience-audit.md @@ -2612,10 +2612,25 @@ guaranteed not to respond. **Scope, stated precisely.** This is a *local* blindness. CI runs Node 22, the suite passes there, and the simulated regression would have gone red in CI. -Nothing about this lets a real boot bug reach `main` on its own. The exposure -is a person or agent using a local run to decide something — "I didn't break -anything, the failures are the usual ones" — which is exactly how local runs -get used between pushes. +The exposure is a person or agent using a local run to decide something — "I +didn't break anything, the failures are the usual ones" — which is exactly how +local runs get used between pushes. + +**But that reassurance is conditional, and it did not hold for the incident +that produced this entry.** @sprint-review (57015): combined with those +branches having no CI at all, the boot bug had precisely zero instruments +pointed at it. The timeline confirms it — `#1123`, which dropped the +`branches: [main]` filter and so gave stacked PRs any CI at all, merged at +15:20:27Z. The boot bug (a controller importing a script that ran `main()` at +module scope, so booting the server ran a data migration) lived on a stacked +branch before that, where `pull_request` workflows did not dispatch. + +So during that window the local suite could not show a new boot failure +(this entry) and the remote suite was not running (entry 41, mechanism 1). +**Each entry's stated mitigation is the other entry's failure mode**, and +they coincided on one defect. "CI would catch it" is not a property of the +repo; it is a property of whether CI ran on that PR, which entry 41 exists +to say you cannot assume. Related to entry 42 but distinct. There the count sits in the verdict slot and means nothing (`0 total`). Here the count is real, the suite genuinely ran, and