fix(pg): retry the boot connect, mount in place, and make readiness assert the mount (TASK-168) - #1901
Conversation
…eds (TASK-168)
WHAT HAPPENED. On 2026-09-25 a production pod's startup PG connect timed out
("Connection terminated due to connection timeout"). `server.ts` called
`connectPG()` once and treated a null result as "PostgreSQL is not available on
this pod, forever": the pod reported `/api/pg/status` → available:false for its
whole life, 404'd `/api/pg/messages` because the routes were never mounted (so
chat history would not load), sent socket chat writes to Mongo, and never
started pg-retention or installation-cleanup — while PG itself answered a probe
from that same pod in 331ms. A rollout restart on the same image fixed it, i.e.
the only recovery was a human noticing.
A transient connect failure at boot is not a property of the deployment, so it
must not be decided once.
WHAT THIS COMMIT DOES (acceptance 1 and 3).
- `services/pgBootService.ts`: a retrying boot. Synchronous attempts with
exponential backoff (defaults: 5 attempts, 500ms base, 30s cap), then a
background retry every 30s until it succeeds, and the routes mount on the
first success. All four are env-tunable (PG_BOOT_RETRY_ATTEMPTS,
_BASE_DELAY_MS, _MAX_DELAY_MS, _INTERVAL_MS). It never throws, and it mounts
only after a connect AND a schema initialization both succeed — a pod must
not serve 500s on a route it only pretends to have. `pgBootState` is the
observable result (mounted / attempts / fellBack / lastError / mountedAt).
- `server.ts`: the 98-line boot block (one connect, five copies of a
`{available:false}` placeholder status handler, one per failure branch)
becomes a `createPgBoot(...).start()`. `/api/pg/status` is now mounted
unconditionally and NOT as a placeholder: `checkStatus` reads the pool and
the schema itself, so it is correct in every state — including the retry
window, where the placeholders could never tell the truth after a late mount.
Its POST /sync-user is only ever called after a GET said available:true
(SocketContext.tsx:45-49), i.e. only when the pool is up.
- `__tests__/unit/services/pgBootService.test.js` (8 cases): the acceptance-3
case makes the FIRST connect time out and proves the SECOND attempt mounts
the routes; plus the fallback-then-background-success case, the
schema-not-applied case, exponential backoff and its cap, mount idempotence,
a throwing connector, post-mount work only on success, and the timer being
cancelled once it succeeds. Clock, scheduler, connector and initializer are
all injected, so the retry SHAPE is asserted, not just the end state.
- `__tests__/unit/server.test.js`: the `server pg status route` block asserted
the placeholder contract (available:false on every failure branch) and is
replaced by a `server pg boot routes` block that asserts what server.ts owns
— the status router is mounted for every configuration, the message routes
stay unmounted while the boot connect keeps failing and while schema
initialization fails, and a first connect that times out followed by a
success mounts them. The mocked pg-messages router gained a route so a
mounted router is distinguishable from an absent one.
ACCEPTANCE 2 (readiness) IS NOT IN HERE, deliberately. It is ruled on
separately: readiness is `/api/health/ready` (values.yaml:86) and
`values-dev.yaml:23` sets replicaCount 1, so making readiness fail whenever PG
is unreachable would take the only pod out of the Service and turn a PG blip
into a total API outage — including Mongo-backed routes — which is the opposite
of what a written decision in values.yaml:80-83 and the comment in
`routes/health.ts` chose. The fork is with lily-shen on the row.
EVIDENCE. `lint:ts` 0 errors. Full backend suite 445 suites / 4148 tests, with
one unrelated pre-existing flake (`__tests__/service/uploads.signedurl.integration.test.js`,
a per-token rate-limit case that passes on its own and alongside this suite;
that file builds its own express app and never imports server.ts).
Nine mutations, each red on the assertion it should red, each restored
byte-identical with a green baseline after: one boot attempt only; no
background retry scheduled; routes mounted before the schema is initialized;
no cap on the backoff; mount() no longer idempotent; a throwing connector
propagating; the background canceller never called on success; an eager
`app.use('/api/pg/messages', …)` at boot (reds the two negative server-level
cases, i.e. the incident's own defect shape); and — the harness case that
matters — the mount-timeout bump, which is what surfaced that this suite
re-requires server.ts four times and needed an explicit budget under full-run
load rather than jest's 30s default.
NOT VERIFIED. The retry has not been exercised against a real PG that is down
at boot; the production path it fixes is a 2-minute connect timeout, and the
delays here are the configured ones rather than that wall-clock shape.
…ASK-168)
Acceptance 2 of the production incident: a pod whose boot connect timed out had
no PG routes at all — `/api/pg/messages` 404'd, so chat history could not load —
and it answered `/api/health/ready` 200 anyway, so a rollout sent it 100% of the
traffic while the pod it replaced was healthy. Recovery required a human.
`98f20db9` made that pod retry and mount in place; this makes it say so.
WHAT WAS WRONG WITH THE OLD GATE. The handler asked whether a live PG query
worked and returned 200 when it failed (`degraded: ['postgresql']`). It could not
see the defect: `config/db-pg.ts` creates the pool lazily, so `SELECT 1` answered
the moment the transient passed while the routes stayed unmounted. Asking PG how
it feels is not the same question as whether this pod can serve chat, and lily's
20:57Z scope note asks for the second one: "readiness must not reuse that block's
outcome as its own evidence; assert the mount."
WHAT IT ASSERTS NOW. `/ready` returns 503 `{status:'not_ready', reason:'PostgreSQL
routes are not mounted on this pod'}` when `PG_HOST` is set and the PG router is
absent from the app — and 200 the moment the retry mounts it, with no restart.
PG's liveness after a successful mount stays ungated on purpose: failing every
pod's readiness on a shared-dependency blip would turn partial degradation into a
full API outage (values.yaml, readinessProbe; `health.ready.test.js` still pins
the degraded-200).
THE PROBE IS THE ROUTE TABLE, NOT A FLAG. `routerIsMounted(app, pgMessageRoutes)`
walks the Express stack for the router object by identity, so it cannot drift
from what is actually mounted, and it stays correct if the path changes. It is
deliberately not `pgBootState.mounted`/`pgAvailable`: a pod whose boot block
believes it is fine while the route table disagrees is exactly the 2026-09-25
shape. `server.ts` installs it before `createPgBoot`, so the answer is live.
A STRUCTURAL ASSERTION, and why. The states reachable in `server.test.js` cannot
tell the flag from the table — no PG, both false; PG working, both true. A
rewiring to `pgAvailable` would keep every behavioural case green while restoring
the exact blindness that caused the outage, so the test also reads `server.ts`
and refuses a probe wired to a flag (`routerIsMounted(app, pgMessageRoutes)`
required; `setPgMountProbe(() => pgAvailable|pgBootState)` refused).
EVIDENCE. Full backend suite 445 suites / 4158 tests green (`npx jest
--forceExit`, Node 22); `npm run lint:ts` 0 errors; `tsc --noEmit` clean (only
the pre-existing root-level `test-discord-interactions.ts` errors). Six
mutations, each anchored, each restored byte-identical with a green baseline
after: mount gate deleted → 2 red; gate back on the pool proxy → 2 red; probe
wired to `pgAvailable` → the structural assertion; probe wired to `true` → the
unmounted-route assertion; `routerIsMounted` answering true for anything → 4 red;
`mountRoutes` not registering the router → 1 red.
NOT VERIFIED. Not exercised against a real PG that is down at boot: every case
above injects the failure, so what is proven is the logic and the wiring, not the
end-to-end timing of a real connection timeout. `routeRateLimitGuard` and the
deploy-dev probe in #1899 are untouched.
DELIBERATELY NOT IN THIS COMMIT: deleting the Mongo chat fallback (`pgAvailable`
in the socket branch at `server.ts:607`, `pgAvailable()` in
messageController.ts). lily's note asks for it and the argument is sound — the
Mongo path's newest document is 2026-08-05, so it hides failures while serving
stale rows — but it is a user-visible behaviour change on the write path, not
part of "make the pod report the capability it is missing", and it should be
reviewed as its own change rather than inside an incident fix. Flagged on the row.
WHY THIS IS IN THIS PR. The new readiness gate turned #1901's kind smoke test red on the deploy step — the backend pod stayed 0/1 for the full 300s while every other pod was Ready, and `helm --wait` timed out. The pod was not slow; it was in a state it could never leave. The cause is this file: SSL disabled — ... (old) "Using CA certificate from: /app/certs/ca.pem" "SSL configuration added with CA certificate" "PostgreSQL connection error: The server does not support SSL connections" The chart sets `PG_SSL_ENABLED` (false locally, true in dev/prod) and the local secrets file carries the comment "Empty cert — not used when PG_SSL_ENABLED=false", but `db-pg.ts` never read that variable: it turned SSL ON whenever the CA *path* existed as a file. The local chart mounts a zero-byte `ca.pem` under a real path, so the smoke cluster's backend went into SSL mode against an in-cluster Postgres with no TLS and could not connect at all. That has been true on main for as long as the local chart has existed. Nothing reported it because the old readiness probe ignored PG: the pod served Mongo-fallback chat, returned 200, and the smoke suite passed. So a green kind smoke never proved the PG routes were mounted — which is the same blindness, in the same place, that made the 2026-09-25 production incident possible. The gate did not create this bug; it is the first thing to see it. THREE FAILURES IN ONE DIRECTION, all "SSL on for a server that has none": 1. `PG_SSL_ENABLED` ignored — now a first-class input. Unset still means ON, so dev and prod (both set "true" with a real Aiven CA) are byte-identical in behaviour. 2. An EMPTY CA file counted as a CA, because `fs.existsSync` was the entire test. `{ ca: '' }` both forces TLS and gives Node nothing to verify against. Any instance whose CA secret materializes empty lands here, not just kind. 3. The decision can now be read/tested as a pure function (`resolvePgSsl`) rather than only as a side effect at import time — the chart documents the intended rule in two comments, and a rule that lives in a side effect is how it drifted from them. WHAT WAS DELIBERATELY NOT CHANGED. A missing or unreadable CA still disables SSL rather than blocking boot: an unstoppable pod is worse than an unencrypted one that fails its own readiness gate, and the reasons are now logged with the decision instead of looking like a decision that was made. `rejectUnauthorized: true` is unchanged where SSL is on. EVIDENCE. With a real Postgres and the app run from source, the smoke cluster's exact shape — `PG_SSL_ENABLED=false` plus a zero-byte `ca.pem` at a real path — went from "The server does not support SSL connections" / `connectPG() → null` to `PostgreSQL connected: PostgreSQL 17.10` / `connectPG() → pool`. 9 new cases in `db-pg.ssl.test.js` (including the empty, whitespace-only, missing, unreadable and explicitly-true shapes, and one structural assertion that the module calls the exported function rather than deciding inline a second time). Backend suite with node 22: 444 suites / 4166 tests passed, the 2 reds being `agentsRuntime.rateLimitTiers` ("socket hang up") and `oauthController` (TASK-133) under parallel load — both pass in isolation, neither imports this file. `tsc --noEmit` clean; `lint:ts` 0 errors. Also fixed while here: a comment in routes/health.ts still named the probe `pathIsMounted`, which is not what it is called.
…ext (TASK-168) sprint-review caught this on #1901 and they are right: the empty-CA guard changed DIRECTION. Before it, a zero-byte ca.pem meant TLS against a server with no TLS — a failed handshake, loud, no plaintext. The first version of this fix turned that same shape into `ssl: false`, i.e. a *successful plaintext* connection. In the file whose entire job is transport security, that is a downgrade traded for a green build. A configured `PG_SSL_CA_PATH` with an unusable file is a misconfiguration, never a request for plaintext. So the unusable-CA branches now return `{ rejectUnauthorized: true }` with no custom CA: TLS stays required, verification falls to the system store, and the handshake fails loudly (readiness then reports it) rather than a pod reading and writing chat in the clear because a secret materialized empty. The dev/prod shape this was written for — `PG_SSL_ENABLED` explicit "true" plus a real Aiven CA — is untouched. Plaintext now has exactly two doors, both deliberate, and neither is "a CA file we could not use": an explicit `PG_SSL_ENABLED=false` (the kind cluster, which is why the smoke fix still holds — the flag short-circuits before the CA is consulted), and an unset `PG_SSL_CA_PATH`, which is the pre-existing behaviour and is left alone rather than tightened in this PR. Tests: 13 cases, and they assert the reason STRING as well as the ssl decision, so each branch pins why it ran — a decision about TLS that is only visible as an object equality is one refactor away from being unverifiable. A new case pins the whitespace-only CA path (a chart that renders an unset value into a space), and one names the direction explicitly: the pairing that yields plaintext. Mutations (each restored byte-identical, file sha 15a6c7e7608a before and after): attaching the empty CA again — the pre-fix shape — reds all 5 unusable-CA cases; ignoring PG_SSL_ENABLED reds 3; deleting the empty guard reds; removing the path trim reds the new whitespace case. Recorded for the next reader: my first attempt at mutation 1 returned "Tests: 0 total" — a suite that fails to RUN, not an assertion that fails, because `ssl: false` in that branch is a type error. That is the same trap as a mutation that never applied; the mutation has to compile.
Two findings from the gate on `288591e7`, both in code this PR just wrote.
1. THE LEVEL. Main warned for a missing CA file (`console.warn`, and
`console.error` when the read threw). The rewrite flattened every branch into
one `SSL disabled — …` line at `console.log`, so the three ways an explicit
TLS intent can fail reported as routine config — the silent-config-failure
shape this row exists to kill, reintroduced by its own fix. `level` is now
part of the returned decision and the module logs with it, so the caller
cannot lose the distinction: deliberate config is info, a failed intent is
warn.
2. PRESENT-BUT-BLANK IS NOT ABSENT. `unset` was decided on the TRIMMED value, so
a whitespace-only `PG_SSL_CA_PATH` took the no-path branch and downgraded to
plaintext — one line below the empty-FILE case that correctly fails closed.
The discriminator is the key: absent means not configured, anything present
but unusable means configured wrong and keeps TLS required.
WHAT FAIL-CLOSED COSTS, and the two configs it broke. Any environment that sets
`PG_SSL_CA_PATH` without a CA behind it now fails its connect instead of quietly
going unencrypted. The chart and the CI workflows already said which they wanted
(`values-local.yaml: pgSslEnabled: "false"`, `PG_SSL_ENABLED=false` in
tests.yml/playwright.yml). docker-compose and backend/.env.example declared
`/app/ca.pem`, a path the backend image does not contain, and relied on the
silent downgrade — so both now state `PG_SSL_ENABLED=false`.
Measured against a real Postgres 17 with no TLS, the exact compose shape
(`PG_SSL_CA_PATH` set, file absent):
flag unset → `SSL enabled — CA file not found … keeping TLS on …` (warn)
`The server does not support SSL connections`, connectPG() null, rc 1
flag false → `SSL disabled — PG_SSL_ENABLED=false` (info)
`PostgreSQL connected: PostgreSQL 17.10`, connectPG() → pool, rc 0
Evidence: ssl suite 15/15; config + health.ready + pgBootService + server 64/64;
`tsc --noEmit` clean for `config/`; eslint 0 errors. Seven mutations, baseline
green on both sides, each restored byte-identical: blank-path treated as unset
again → 4 red; warn flattened to info → 8 red; empty CA counted as a CA → 3;
`PG_SSL_ENABLED` ignored → 3; unusable CA downgraded to plaintext → 9; level
flattened at the call site → 1; the CA file never read at all → 6 (the
non-vacuity control).
Conflict: backend/__tests__/unit/server.test.js, on the one describe this PR replaces. #1914 (TASK-129) rewrote the four 'server pg status route' tests in place to await the mount instead of a fixed number of turns; this PR deletes those four and replaces them with the pg-boot-route suite, because the real status router is now mounted unconditionally rather than only after a successful connect. Resolution takes the branch's describe block and keeps #1914's non-conflicting hunks: 'delete process.env.PG_HOST' in the afterEach of 'server route precedence' and 'server websocket authorization helpers'.
|
Arrival check against a moved base — main advanced into BOTH files this PR edits. Recorded because a clean auto-merge of two edits to one file is exactly where a semantic break hides.
Conflict: none. Semantic check, by grep on the reconstructed tree:
Reconstructed merge is green: Instrument note. The first run of that command looked red (8 failures, Stale-base distance is 28 against the ceiling 40; head unchanged at |
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75962) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets: four branches that merged main into themselves carry one extra commit and one fewer bullet (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). Census over 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates for the subject can be told apart: one non-merge commit landed the commit's subject 32 of 32; two or more landed the PR title 50 of 50. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Corrections this revision carries, both of them mine. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the subject box and generalised a rule from it; with the setting read, #2031 is the one-commit default. The second draft then reported "50 of 51" with #1964 as an unexplained residual: that was a classifier artifact, `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The doc's classifier, its recipe and its census are all stated as non-merge counts now, and it says plainly that an editable box which has never been seen being edited is not evidence that a given landing was one. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749): #1645 matches NEITHER the PR title nor the commit's subject (a hand-typed hybrid, 0 title renames); #1644 landed the FIRST COMMIT's subject on a branch of 2 non-merge commits, where the count says the title; #1623 landed the PR TITLE where the count of 1 non-merge commit says the commit's subject (its body carries 0 bullets). Zero overrides in the 297 PRs measured below (#1696-#2053). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 32 of 32; two or more landed the PR title 50 of 50. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Three corrections this revision carries, all mine. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third is this revision: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The doc now carries the override as a measured finding with its signatures, and every claim it makes names the window it was measured in. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749): #1645 matches NEITHER the PR title nor the commit's subject (a hand-typed hybrid, 0 title renames); #1644 landed the FIRST COMMIT's subject on a branch of 2 non-merge commits, where the count says the title; #1623 landed the PR TITLE where the count of 1 non-merge commit says the commit's subject (its body carries 0 bullets). Zero overrides in the 297 PRs measured below (#1696-#2053). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 194 landings took the commit's subject, 32 of them on a one-commit branch, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 32 of 32; two or more landed the PR title 50 of 50. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Three corrections this revision carries, all mine. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third is this revision: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The doc now carries the override as a measured finding with its signatures, and every claim it makes names the window it was measured in. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749). #1645 matches NEITHER the PR title nor the commit's subject (0 rename events) and is the only one of the three an edit is needed to explain; #1623 landed the PR TITLE and #1644 landed commit 1's subject, each BYTE-EXACT, where the count predicts the other candidate - so a box seeded when the dialog rendered explains those two without an edit (#1644's timings: commit 1 at 20:24:15Z, commit 2 at 20:29:26Z, the press at 20:52:58Z). They establish that the prediction can fail; only #1645 establishes why. Zero overrides in the 297 PRs measured below (#1696-#2053). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 194 landings took the commit's subject, 32 of them on a one-commit branch, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 32 of 32; two or more landed the PR title 50 of 50. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Three corrections this revision carries, all mine. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third is this revision: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The doc now carries the override as a measured finding with its signatures, and every claim it makes names the window it was measured in. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749). #1645 matches NEITHER the PR title nor the commit's subject (0 rename events); #1623 landed the PR TITLE and #1644 landed commit 1's subject, each BYTE-EXACT, where the count predicts the other candidate. A stale prefill - the dialog populates its subject field when it RENDERS, not when it is pressed - accounts for #1644 (commit 1 at 20:24:15Z, commit 2 at 20:29:26Z, the press at 20:52:58Z, so a box opened in that window carried commit 1's subject) and for NEITHER of the other two: #1623's branch held one non-merge commit for its whole life with 0 force-push events, so no render of it could have offered the title. Five branches in the same window carrying main-merges landed the commit's subject where counting every commit would predict the title (#1678 1+1, #1647 1+2, #1574 1+1, #1558 1+2, #1539 1+1) - which is what makes #1623 an override and not a conforming multi-commit landing. Zero overrides in the 297 merged PRs measured below (#1750-#2053, contiguity-checked). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 96 discriminating landings took the commit's subject, 32 of them on a branch holding one non-merge commit, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 33 of 33; two or more landed the PR title 49 of 49. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Four corrections this revision carries, all mine. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The fourth, also ux-lead's: this revision claimed only #1645 needed an edit, treating a stale prefill as an explanation for #1623 - whose one non-merge commit never grew, so no prefill could have offered its title - and mislabelled the census window #1696-#2053 when the fetch, capped at 300 rows sorted by UPDATED, spanned #1750-#2053 and missed three merged rows inside it (#1752, #1753, #1765, since read and conforming). Both fixed; every claim the doc makes now names the window it was measured in. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749). #1645 matches NEITHER the PR title nor the commit's subject (0 rename events); #1623 landed the PR TITLE and #1644 landed commit 1's subject, each BYTE-EXACT, where the count predicts the other candidate. A stale prefill - the dialog populates its subject field when it RENDERS, not when it is pressed - accounts for #1644 (commit 1 at 20:24:15Z, commit 2 at 20:29:26Z, the press at 20:52:58Z, so a box opened in that window carried commit 1's subject) and for NEITHER of the other two: #1623's branch held one non-merge commit for its whole life with 0 force-push events, so no render of it could have offered the title. Five branches in the same window carrying main-merges landed the commit's subject where counting every commit would predict the title (#1678 1+1, #1647 1+2, #1574 1+1, #1558 1+2, #1539 1+1) - which is what makes #1623 an override and not a conforming multi-commit landing. Zero overrides in the 297 merged PRs measured below (#1750-#2053, contiguity-checked). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 96 discriminating landings took the commit's subject, 32 of them on a branch holding one non-merge commit, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 33 of 33; two or more landed the PR title 49 of 49. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Four corrections this revision carries, all mine. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The fourth, also ux-lead's: this revision claimed only #1645 needed an edit, treating a stale prefill as an explanation for #1623 - whose one non-merge commit never grew, so no prefill could have offered its title - and mislabelled the census window #1696-#2053 when the fetch, capped at 300 rows sorted by UPDATED, spanned #1750-#2053 and missed three merged rows inside it (#1752, #1753, #1765, since read and conforming). Both fixed; every claim the doc makes now names the window it was measured in. (5) An addition rather than a correction: `git log --no-merges -1 <branch>` means "newest non-merge commit REACHABLE", so on a merge-tipped branch it hands back another PR's commit on main (#1623 -> "...(#1626)", #1574 -> "...(#1576)"); the recipe now ranges from the merge base and carries #1964 (tip not a merge) as the control. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749). #1645 matches NEITHER the PR title nor the commit's subject (0 rename events); #1623 landed the PR TITLE and #1644 landed commit 1's subject, each BYTE-EXACT, where the count predicts the other candidate. A stale prefill - the dialog populates its subject field when it RENDERS, not when it is pressed - accounts for #1644 (commit 1 at 20:24:15Z, commit 2 at 20:29:26Z, the press at 20:52:58Z, so a box opened in that window carried commit 1's subject) and for NEITHER of the other two: #1623's branch held one non-merge commit for its whole life with 0 force-push events, so no render of it could have offered the title. Five branches in the same window carrying main-merges landed the commit's subject where counting every commit would predict the title (#1678 1+1, #1647 1+2, #1574 1+1, #1558 1+2, #1539 1+1) - which is what makes #1623 an override and not a conforming multi-commit landing. Zero overrides in the 297 merged PRs measured below (#1750-#2053, contiguity-checked). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 96 discriminating landings took the commit's subject, 32 of them on a branch holding one non-merge commit, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 33 of 33; two or more landed the PR title 49 of 49. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Four corrections this revision carries, all mine - plus an addition at (5) and a correction to it at (6), and that one is sprint-review's catch. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The fourth, also ux-lead's: this revision claimed only #1645 needed an edit, treating a stale prefill as an explanation for #1623 - whose one non-merge commit never grew, so no prefill could have offered its title - and mislabelled the census window #1696-#2053 when the fetch, capped at 300 rows sorted by UPDATED, spanned #1750-#2053 and missed three merged rows inside it (#1752, #1753, #1765, since read and conforming). Both fixed; every claim the doc makes now names the window it was measured in. (5) An addition rather than a correction: `git log --no-merges -1 <branch>` means "newest non-merge commit REACHABLE", so on a merge-tipped branch it hands back another PR's commit on main (#1623 -> "...(#1626)", #1574 -> "...(#1576)"); the recipe now ranges from the merge base. (6) A correction to (5) itself, and it is sprint-review's: the control half of (5) was wrong. #1964's tip IS a merge (3e7ce46, parents 3a65832 own + ff8de0f merged-in), so merge-tip is NECESSARY AND NOT SUFFICIENT for the trap. What spares #1964 is DATE ORDER, by 40 seconds - its own commit at 11:17:28Z against the merged-in 11:16:48Z - while #1623 loses the same race by 30 minutes (own a9c6cb2 at 06:55:46Z against main's 9f75906, #1626, at 07:25:57Z), which is why #1623's unranged read returns the SEO guide. #1964 is therefore a NEAR MISS, not a clean control, and the clean control is a tip that is not a merge at all, where main is unreachable: #2051, #2053, #1752, #1753, #1765, #1645, #2031, each with one parent, each reading its own commit unranged. The doc carries the corrected condition in both the recipe and the habits bullet. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749). #1645 matches NEITHER the PR title nor the commit's subject (0 rename events); #1623 landed the PR TITLE and #1644 landed commit 1's subject, each BYTE-EXACT, where the count predicts the other candidate. A stale prefill - the dialog populates its subject field when it RENDERS, not when it is pressed - accounts for #1644 (commit 1 at 20:24:15Z, commit 2 at 20:29:26Z, the press at 20:52:58Z, so a box opened in that window carried commit 1's subject) and for NEITHER of the other two: #1623's branch held one non-merge commit for its whole life with 0 force-push events, so no render of it could have offered the title. Five branches in the same window carrying main-merges landed the commit's subject where counting every commit would predict the title (#1678 1+1, #1647 1+2, #1574 1+1, #1558 1+2, #1539 1+1) - which is what makes #1623 an override and not a conforming multi-commit landing. Zero overrides in the 297 merged PRs measured below (#1750-#2053, contiguity-checked). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 96 discriminating landings took the commit's subject, 32 of them on a branch holding one non-merge commit, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 33 of 33; two or more landed the PR title 49 of 49. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Six corrections this revision carries, four of them mine and two sprint-review's, plus an addition at (5) that both of theirs correct. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The fourth, also ux-lead's: this revision claimed only #1645 needed an edit, treating a stale prefill as an explanation for #1623 - whose one non-merge commit never grew, so no prefill could have offered its title - and mislabelled the census window #1696-#2053 when the fetch, capped at 300 rows sorted by UPDATED, spanned #1750-#2053 and missed three merged rows inside it (#1752, #1753, #1765, since read and conforming). Both fixed; every claim the doc makes now names the window it was measured in. (5) An addition rather than a correction: `git log --no-merges -1 <branch>` means "newest non-merge commit REACHABLE", so on a merge-tipped branch it hands back another PR's commit on main (#1623 -> "...(#1626)", #1574 -> "...(#1576)"); the recipe now ranges from the merge base. (6) A correction to (5) itself, and it is sprint-review's: the control half of (5) was wrong. #1964's tip IS a merge (3e7ce46, parents 3a65832 own + ff8de0f merged-in), so merge-tip is NECESSARY AND NOT SUFFICIENT for the trap. What spares #1964 is DATE ORDER, by 40 seconds - its own commit at 11:17:28Z against the merged-in 11:16:48Z - while #1623 loses the same race by 30 minutes (own a9c6cb2 at 06:55:46Z against main's 9f75906, #1626, at 07:25:57Z), which is why #1623's unranged read returns the SEO guide. #1964 is therefore a NEAR MISS, not a clean control, and the clean control is a tip that is not a merge at all, where main is unreachable: #2051, #2053, #1752, #1753, #1765, #1645, #2031, each with one parent, each reading its own commit unranged. The doc carries the corrected condition in both the recipe and the habits bullet. (7) Also sprint-review's, and again about (5): the doc measured the trap "on the seven merge-tipped branches in this population, six returned a different PR's commit". That population - every merge-carrying row of #1534-#1749 holding one non-merge commit - contains 22 such branches, all 22 merge-tipped, and 21 of them return a different commit, so 6-of-7 understated the hazard 3.5x; "this population" also had no stated boundary in the doc. The one quiet row is #1651, and it is quiet for the same reason #1964 survives: its own 1270011 (05:26:06-07:00) is newer than the tip it merged in, f3ae379 (00:49:19-07:00), so the date-ordered walk never leaves the branch. Coincident candidates do not make a row quiet - 15 of the 16 rows in that population whose title equals their first commit subject are among the 21 that fire. The doc now states the window, the ratio, and #1651's measured reason. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
…PR takes its subject from the commit, two or more take the PR title (TASK-223) Written by sprint-impl, a Commonly agent. Pod thread: https://commonly.me/v2/pods/6a692a1be833c668acdb84cf (message 75968) Two surfaces compose a squash message and a repository setting decides which one supplies the subject. `squash_merge_commit_title: COMMIT_OR_PR_TITLE` means one non-merge commit takes the commit's subject and two or more take the PR title; `squash_merge_commit_message: COMMIT_MESSAGES` means the body is always the branch's commit messages. Both read from the repo, so the default composition is decidable before a press. The PR body is not an input. The count is of NON-MERGE commits, and merge commits are excluded from both the count and the body bullets (#1965 7+1 -> 7, #1901 6+1 -> 6, #1905 5+1 -> 5, #1906 4+1 -> 4). THE SUBJECT CAN BE OVERRIDDEN, and that is measured rather than assumed. Three landings took a subject the setting does not predict, all merged 2026-09-08, in an older window of 194 merged PRs (#1534-#1749). #1645 matches NEITHER the PR title nor the commit's subject (0 rename events); #1623 landed the PR TITLE and #1644 landed commit 1's subject, each BYTE-EXACT, where the count predicts the other candidate. A stale prefill - the dialog seeding its subject field when it RENDERS, so a box opened earlier is pressed carrying an older subject; INFERRED, never read, since no API field records the client or the box - accounts for #1644 (commit 1 at 20:24:15Z, commit 2 at 20:29:26Z, the press at 20:52:58Z, so a box opened in that window carried commit 1's subject) and for NEITHER of the other two unless that inference holds: #1623's branch held one non-merge commit for its whole life with 0 force-push events, so no render of it could have offered the title if the dialog prefills by the count the landings follow. Five branches in the same window carrying main-merges landed the commit's subject where counting every commit would predict the title (#1678 1+1, #1647 1+2, #1574 1+1, #1558 1+2, #1539 1+1) - which is what makes #1623 an override and not a conforming multi-commit landing. Zero overrides in the 297 merged PRs measured below (#1750-#2053, contiguity-checked). So an override is detectable — a landed subject the setting and the count do not predict — and a census is a claim about its own window, not about the practice. The "the setting was different then" reading is pre-empted by the same window: 33 of its 96 discriminating landings took the commit's subject, 32 of them on a branch holding one non-merge commit, which a PR-title setting cannot produce without an override each. Census, 297 merged PRs, restricted to the discriminating set — the 82 PRs whose title and first commit subject differ, the only set where the two candidates can be told apart: one non-merge commit landed the commit's subject 33 of 33; two or more landed the PR title 49 of 49. No exceptions. Six of the nine merges of 2026-09-30 had title == first commit subject and could not discriminate at all; the three that could (#2024, #2031, #2049) follow the setting. Eight corrections this revision carries: four of them mine, two sprint-review's and two ux-lead's, plus an addition at (5) that the later ones correct. The first draft read #2031's landed subject (the commit's, not the title's) as a presser overwriting the box and generalised from it; with the setting read, #2031 is the one-commit default. The second reported "50 of 51" with #1964 as an unexplained residual — a classifier artifact, since `gh pr view --json commits` counts merge commits, and #1964's branch is 1 non-merge + 1 merge and landed in the one-commit shape with no bullets. The third: the commit message said an editable box "has never been seen being edited", which is a claim about the whole history made from a 297-PR window, and ux-lead refuted it with #1644 and #1645. The fourth, also ux-lead's: this revision claimed only #1645 needed an edit, treating a stale prefill as an explanation for #1623 - whose one non-merge commit never grew, so no prefill following the landings' count could have offered its title - and mislabelled the census window #1696-#2053 when the fetch, capped at 300 rows sorted by UPDATED, spanned #1750-#2053 and missed three merged rows inside it (#1752, #1753, #1765, since read and conforming). Both fixed; every claim the doc makes now names the window it was measured in. (5) An addition rather than a correction: `git log --no-merges -1 <branch>` means "newest non-merge commit REACHABLE", so on a merge-tipped branch it hands back another PR's commit on main (#1623 -> "...(#1626)", #1574 -> "...(#1576)"); the recipe now ranges from the merge base. (6) A correction to (5) itself, and it is sprint-review's: the control half of (5) was wrong. #1964's tip IS a merge (3e7ce46, parents 3a65832 own + ff8de0f merged-in), so merge-tip is NECESSARY AND NOT SUFFICIENT for the trap. What spares #1964 is DATE ORDER, by 40 seconds - its own commit at 11:17:28Z against the merged-in 11:16:48Z - while #1623 loses the same race by 30 minutes (own a9c6cb2 at 06:55:46Z against main's 9f75906, #1626, at 07:25:57Z), which is why #1623's unranged read returns the SEO guide. #1964 is therefore a NEAR MISS, not a clean control, and the clean control is a tip that is not a merge, where the branch's own commit is the newest thing reachable: #2051, #2053, #1752, #1753, #1765, #1645, #2031, one parent each and origin/main an ancestor of none of those tips, each reading its own commit unranged. The doc carries the corrected condition in both the recipe and the habits bullet. (7) Also sprint-review's, and again about (5): the doc measured the trap "on the seven merge-tipped branches in this population, six returned a different PR's commit". That population - every merge-carrying row of #1534-#1749 holding one non-merge commit - contains 22 such branches, all 22 merge-tipped, and 21 of them return a different commit, so 6-of-7 understated the hazard 3.5x; "this population" also had no stated boundary in the doc. The one quiet row is #1651, and it is quiet for the same reason #1964 survives: its own 1270011 (05:26:06-07:00) is newer than the tip it merged in, f3ae379 (00:49:19-07:00), so the date-ordered walk never leaves the branch. Coincident candidates do not make a row quiet - 15 of the 16 rows in that population whose title equals their first commit subject are among the 21 that fire. The doc now states the window, the ratio, and #1651's measured reason. (8) Both of ux-lead's. (a) The contiguity split was wrong: two of the range's 304 numbers are ISSUES, not PRs (#1821 closed, #1959 open), so the split is 302 pull requests - 297 merged, 3 closed unmerged (#1784, #1903, #1967), 2 open (#1751, #1768) - and 2 issues, checked one at a time with the `pull_request` key on `gh api .../issues/N`, not "297 merged, 3 unmerged, 4 open" of 304 PR numbers. (b) The doc and this message stated the prefill as measured while the PR ask called it an inference, and the ask was the honest one: no prefill was ever read, because the measured rule is about what LANDS. #1623's exclusion is therefore conditional on the dialog prefilling by the same count, and if it counts merges instead then #1623 needed no edit and the merge-carrying landings that took a one-commit-predicted subject are the ones to explain. Every surface now says inferred; ux-lead notes their own 75d82be wording was where the overclaim entered. Also recorded because the doc quotes them: `git log --grep="(#N)$"` resolves the WRONG commit when a later commit on main also names that PR (#1677 returned de5fc4d, actual f8ad3bd; #1877 8235be3 vs e3d9550), so the recipe uses the PR's mergeCommit; `gh pr view --json commits` truncates headlines (17 of 22 in the nine-PR set, each to 69 characters plus an ellipsis, longest intact 72); and amending the tip does not amend earlier commits — #2049's merge message carries commit 1's retracted framing at line 13 and its retraction at line 58. Scope: this is the message, not the tree. Fidelity of what landed to what was reviewed stays rule 32's patch-id comparison; the tree/origin check is rule 49.
Incident
lily-shen, 20:45Z: a transient Postgres timeout at boot breaks production chat until someone restarts the pod.
backend/server.tsranconnectPG()once at boot. If that attempt returned null — the observed failure wasConnection terminated due to connection timeout— the PG routes were never mounted and nothing retried, ever./api/pg/messages404'd for the life of the pod, so chat history could not load and socket writes fell through to Mongo. Nothing surfaced it: five placeholder handlers under/api/pg/statuseach answered{available:false}for their own failure branch,/api/health/readyreturned 200 (degraded: ['postgresql']), and the readiness probe gated only Mongo. A rollout therefore replaced the healthy pod with this one and sent it 100% of the traffic (replicaCount: 1). The only remedy was a human restart.Acceptance 1 + 3 — retry and mount in place (
98f20db9)backend/services/pgBootService.ts, new: acreatePgBoot(deps)factory with injectedmountRoutes / connect / initialize / onMounted / sleep / schedule / log / state, and module-levelpgBootState(mounted,attempts,fellBack,lastError,mountedAt).500ms → 30scap;PG_BOOT_RETRY_ATTEMPTS,PG_BOOT_RETRY_BASE_DELAY_MS,PG_BOOT_RETRY_MAX_DELAY_MS), then a 30s background interval that keeps trying.pgBootState.lastErrorand the process stays up.server.ts:pgStatusRoutesis required at the top and mounted unconditionally (the five placeholders are gone), and the retention + installation-cleanup crons moved intoonMounted, because they need the same guarantee the routes do.Acceptance 2 — the pod reports the capability it is missing (
aa3c1bb3)The old gate asked whether a live PG query worked. It structurally could not see this defect:
config/db-pg.tscreates the pool lazily, soSELECT 1started succeeding the moment the transient passed while/api/pg/messagesstayed unmounted. Per lily's 20:57Z scope note — "readiness must not reuse that block's outcome as its own evidence; assert the mount":/api/health/readyreturns 503{status:'not_ready', reason:'PostgreSQL routes are not mounted on this pod'}whenPG_HOSTis set and the PG router is absent from the app, and 200 the moment the retry mounts it — no restart, which is the half of the incident that needed a human.health.ready.test.jskeeps pinning the degraded-200, andvalues.yamlnow states both halves of the policy.routerIsMounted(app, pgMessageRoutes)— a walk of the Express stack for the router object, notpgBootState.mounted/pgAvailableand not a path string. Identity rather than prefix-parsing: a pod whose boot block believes it is fine while the route table disagrees is the 2026-09-25 shape.A structural assertion, and why it exists. No reachable state in
server.test.jsseparates the flag from the table — no PG, both false; PG working, both true. A rewiring topgAvailablewould keep every behavioural case green while restoring exactly the blindness that caused the outage, so that test also readsserver.ts:setPgMountProbe(() => routerIsMounted(app, pgMessageRoutes))is required, andsetPgMountProbe(() => pgAvailable|pgBootState)is refused.Evidence
npx jest --forceExit, Node 22)npm run lint:tstsc --noEmittest-discord-interactions.tserrorspgBootService13,server.test.js9,health.ready6Nine mutations across both commits, each anchored, each restored byte-identical with a green baseline after:
AC1/AC3 — connect-once without retry → red; mount despite a failed connect → red; mount before
initializesucceeds → red; mount not idempotent → red; backoff cap removed → red; throwing connector escapingstart()→ red; eager mount on the failed path → red.AC2 — mount gate deleted → 2 red; gate back on the pool proxy → 2 red; probe wired to
pgAvailable→ the structural assertion; probe wired totrue→ the unmounted-route assertion;routerIsMountedanswering true for anything → 4 red;mountRoutesnot registering the router → 1 red.One instrument note, since it cost a cycle: the first form of the AC2 probe parsed the mount prefix back out of Express's
Layer.regexp.source, and its own regex literal was malformed — the suite reported "Test suite failed to run", which reads like a held guard and is red for the wrong reason. The identity walk has no parsing step to get wrong.NOT VERIFIED
routeRateLimitGuardbaseline and the ci(deploy-dev): fail the run when the new backend cannot serve chat #1899 deploy-dev probe are untouched by this PR.Deliberately not in this PR
Deleting the Mongo chat fallback (
pgAvailablein the socket branch atserver.ts:607,pgAvailable()inmessageController.ts). lily's note asks for it — the Mongo path's newest document is 2026-08-05, so it hides failures while serving stale rows — and the argument is sound. It is a user-visible behaviour change on the chat write path rather than part of "make the pod report the capability it is missing", so it belongs in a review of its own, not inside an incident fix. Flagged on the row for a scope call.Deploy note
No migration.
values.yamlprobe config is unchanged — the comment now states both halves of the policy.The kind smoke red this PR caused, and what it turned out to be (
853413ae,35c048f9)The first push failed Smoke Tests (kind cluster) on the deploy step:
backendpod 0/1 Running for the full 300s, every other pod Ready,helm --waittimed out. That is the new gate doing its job and the gate being unable to explain itself, so both halves are fixed here.The pod was not slow — it was in a state it could never leave. Root cause, in
backend/config/db-pg.ts, which this PR now also fixes:The chart sets
PG_SSL_ENABLED(falselocally,truein dev/prod) andlocal-secrets.yamlcarries the comment "Empty cert — not used when PG_SSL_ENABLED=false" — but the code never read that variable. The SSL decision was made by the presence of the CA path, and the local chart mounts a zero-byteca.pemunder a real path, so the smoke cluster's backend went into SSL mode against an in-cluster Postgres with no TLS.connectPG()returned null, the PG routes were therefore never mounted, and readiness — correctly — never passed.This has been true on
mainfor as long as the local chart has existed, and a green smoke never showed it. The old readiness probe ignored PG entirely: the pod served Mongo-fallback chat, returned 200, and the suite passed. So the smoke suite has never proven the PG routes were mounted — the same blindness, in the same place, that made the 2026-09-25 production incident possible. This gate did not create the bug; it is the first instrument to see it.The fix reads
PG_SSL_ENABLED(unset still means ON, so dev and prod are behaviourally identical), and refuses an empty CA file instead of configuring{ ca: '' }— which forces TLS and gives Node nothing to verify against, a shape any instance gets if its CA secret materializes empty, not just kind. The decision is now a pure exported function (resolvePgSsl) with 9 cases, rather than a side effect at import time that had drifted from the two comments documenting the intended rule.Reproduced and fixed against a real Postgres, app run from source, using the smoke cluster's exact shape (
PG_SSL_ENABLED=false+ a zero-byte CA at a real path):The server does not support SSL connectionsPostgreSQL connected: PostgreSQL 17.10connectPG()nullpoolPostgreSQL routes registered for chat functionality/api/pg/status{"available":false}{"available":true}/api/health/readynot_readyreadyAnd the smoke job can now explain itself: on failure it prints the backend deployment's last 120 log lines, the backend pod's events, and an in-cluster
curl /api/health/ready— instead ofError: context deadline exceededwith akubectl describethat, on a five-pod install, is entirely postgres.if: failure()only.Two full-suite reds under node 22 are load flakes, not this change:
agentsRuntime.rateLimitTiers("socket hang up" in the 3,000-request case) andoauthController(TASK-133) both pass in isolation and neither importsdb-pg. Measured: 444 suites / 4166 tests passed with those two red; 23/23 in isolation.