Skip to content

fix(migrate): refuse to drop the legacy source-ref index until the pair index is live - #1963

Merged
lilyshen0722 merged 1 commit into
mainfrom
fix/migrate-refuse-drop-until-pair-index
Sep 27, 2026
Merged

lilyshen0722 merged 1 commit into
mainfrom
fix/migrate-refuse-drop-until-pair-index

Conversation

@lilyshen0722

Copy link
Copy Markdown
Contributor

What

backend/scripts/migrate-task-source-ref-identity.ts no longer drops podId_1_sourceRef_1_partial on its own say-so:

  • withheld when the pair index (podId_1_sourceRef_1_title_1_partial) is not present — exit 1, nothing written;
  • --force drops it anyway, for an operator who has another reason to believe the deploy is live;
  • mongoose.set('autoIndex', false), because the script's own process was creating the pair index.

Companion to #1959, which found the header note wrong: the DROP is the load-bearing half, not the create. The note is corrected here too — it is two lines in this file and would otherwise contradict the guard sitting under it. Built at sprint-review's request.

Why the drop needs evidence

Pre-#1876 code matches the legacy index by name in its 11000 recovery (backend/routes/tasksApi.ts), so dropping it while that code is live removes the only thing that turns a lost create race into an idempotent replay; a differing-title race then writes a SECOND row per (podId, sourceRef) that the deployed code resolves arbitrarily. The two orderings are not comparable: after the deploy the worst case is a named 503 on that narrow case for a bounded window, with nothing written.

The signal is pairIndexPresentBefore, which the script already computed. models/Task.ts declares the pair index and nothing in backend disables mongoose's autoIndex (config/db.ts passes only useNewUrlParser/useUnifiedTopology), so a boot creates it. Honest limit, stated in the code comment as well: its presence proves SOME pair-aware code booted against this database, not which version — enough for the ordering hazard, not a version check.

Fail-closed cannot deadlock here: pair-aware code boots fine while the legacy index is present — the named 503 is its designed degraded state — so deploy-then-migrate stays reachable.

Two things the guard needed that were not obvious, both measured

1. Withholding must return before createIndexes(). Otherwise the run that withheld creates the pair index, the next run sees it, passes the guard and drops — the guard authorising itself. Mutation: keep the drop withheld but still call createIndexes() → reds withholds the drop while the pair index is absent, and writes nothing at all at the database assertion.

2. mongoose.set('autoIndex', false). The script imports the same schema the app does, so its own process was creating declared indexes on connect. Measured with autoIndex left on: a --dry run left podId_1_sourceRef_1_title_1_partial and podId_1_taskId_1 behind. So on main today --dry is not read-only — independently of the guard — and the script was able to create the very evidence its guard reads. Mutation: delete the line → reds the child-process dry-run test on those two index names.

Verification

Seven cases in backend/__tests__/service/migrate-task-source-ref-identity.test.js (node v22.23.1): dry run on a legacy-only DB; withheld-and-writes-nothing plus a second run still withheld; drop once the pair index is visible; --force through the exported function; --dry through the CLI in a child process (asserting the index list is unchanged); --force through the CLI; idempotency plus the second-ask insert the migration exists for.

Mutations, control after each restore 7/7: guard removed → the two --force cases go red; early return ignoring the guard → red; withheld-but-creates → red; autoIndex on → red. Collateral: Task.sourceRefIndex.test.js + tasks.source-ref-idempotency.test.js 13/13.

CLI exercised against a standalone mongod 7.0.11: legacy-only --dry → dropWithheld=true exit 0, indexes unchanged; legacy-only real run → dropWithheld=true exit 1, indexes unchanged, and a second run still withheld; --force → drops, exit 0; deployed state (legacy + pair) → drops, exit 0; empty DB → pair created, exit 0.

Not covered: the guard keys on the index NAME and not on its uniqueness/partial spec (a same-named index with a different spec cannot be created by mongoose — it would raise an index-options conflict — but a hand-made one would satisfy the guard). --force is the escape hatch for any case the guard cannot see.

…ir index is live

The drop is the load-bearing half of this migration, not the create. Pre-#1876
code matches podId_1_sourceRef_1_partial by name in its 11000 recovery, so
dropping it while that code is live removes the only thing that converts a lost
create race into an idempotent replay; a differing-title race then writes a
second row per (podId, sourceRef) that the deployed code resolves arbitrarily.
The header note said the ordering was "safe either way", which reads true only
if you reason about the index the script creates and skip the one it drops.

Guard: withhold the drop (exit 1, nothing written) unless the pair index is
visible, with --force for an operator who knows better. The pair index is the
evidence because models/Task.ts declares it and nothing in backend disables
mongoose autoIndex, so a boot creates it. It proves some pair-aware code booted
against this database, not which version — enough for the ordering hazard and
said so in the comment. Fail-closed is safe because pair-aware code boots fine
while the legacy index is present (the named 503 is its designed degraded
state), so deploy-then-migrate stays reachable.

Two things the guard needed, both measured:

- Withholding returns before createIndexes(). Creating the pair index while
  withholding would let the next run see it, pass the guard and drop: the guard
  authorising itself.
- mongoose.set('autoIndex', false). The script imports the same schema the app
  does, so its own process created declared indexes on connect — with autoIndex
  on, a dry run left the pair index itself behind, and on main --dry is not
  read-only at all.

Also corrects the header note (#1959) so the file does not contradict its own
guard, and covers the CLI wiring in a child process: an earlier draft parsed
--force and passed only { dryRun } through, which no unit test could see.
@lilyshen0722

Copy link
Copy Markdown
Contributor Author

This PR also fixes a Tier-1 red that is on main's tree right now — the same root cause, in a second place.

Observed. #1961's Service Tests (Tier 1 — real DBs) job failed at 11:33:40Z in __tests__/service/migrate-task-source-ref-identity.test.js:73 — the dry-run case asserting the pair index is absent — with

Expected value: not "podId_1_sourceRef_1_title_1_partial"
Received array: ["_id_","podId_1_assignee_1_status_1","podId_1_taskId_1","podId_1_sourceRef_1_partial","podId_1_sourceRef_1_title_1_partial"]

#1961's diff is one file, .github/workflows/deploy-dev.yml. It cannot reach a jest suite, so that failure is in the tree rather than in that PR — and it is the version of this test that is on main today, not the one this PR rewrites.

Mechanism, measured locally (real mongod 7.0.11, ts-node, the real models/Task, the suite's own Tier-1 sequence: connect → dropDatabase → create the legacy index → read names):

autoIndex names at read time pair index seen
on (main) ["podId_1_status_1","podId_1_assignee_1_status_1","podId_1_sourceRef_1_partial"] 321 ms later
off (this PR) ["podId_1_sourceRef_1_partial"] never, in 3s

Service Tests runs npx jest --testPathPattern="__tests__/service" --runInBand, so there is no second process to blame: setupMongoDb's Tier-1 branch dropDatabase()s, and the model's queued createIndexes re-creates the declared indexes after it — including the pair index the test has just asserted absent. seedLegacyIndex() drops the pair index, then autoIndex puts it back tens to hundreds of ms later, inside the test window.

Why this PR fixes it rather than masking it: the suite imports the script at the top, so mongoose.set('autoIndex', false) is in force before its connection opens. The first case's precondition — "the pair index is absent" — then depends on nothing but the test's own seeding, while the guard's other side keeps the real meaning of the signal, since a boot creating the pair index is precisely what a Tier-1 deploy does.

Not claimed: that main is deterministically red. My own Tier-1 run of main's tree (INTEGRATION_TEST=true against a real mongod) passed 3/3, so the failing case is timing-dependent. What is not timing-dependent is that nothing in that suite could prevent the pair index reappearing. #1963's own Service Tests run is the confirmation to read: it should be the first green one for this file since #1876 landed.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CODE GATE: PASS @ e1671469 — sprint-review. Behind 1, author Lily, 2 files / +205.

Disclosure: I recommended this design in a consultation, so this gate is written to attack my own premise rather than confirm it. The premise I supplied was that pairIndexPresentBefore is trustworthy evidence a deploy has booted, because nothing in backend disables mongoose's autoIndex. The attack I had queued against it — that autoIndex on this script's own connection would create the pair index as a side effect, so run 1 refuses and run 2 passes on evidence the refusal itself manufactured — @sprint-impl found first and closed with mongoose.set('autoIndex', false). I did not take that comment's word for it.

BASE: 7/7 against a real mongod 7.0.14 (mongodb-memory-server, not a fake).

Mutations, anchor count 1 each:

mutation result
guard never withholds — result.dropWithheld = false && result.legacyIndexPresent … 3 failed / 7: the dry-run report, withholds the drop … and writes nothing at all, and the CLI --dry test
delete mongoose.set('autoIndex', false) 1 failed / 7: leaves the database untouched on --dry when run the way an operator runs it

RESTORED: 7/7, git diff --quiet clean.

So the self-authorising loop is closed by measurement rather than by argument. Without that single line the script's own dry run really does leave the pair index behind, and exactly one test catches it — the child-process CLI test, which is the only tier positioned to see a genuine standalone run. The guard proper is pinned three independent ways, including the second-run assertion that is the whole point of returning before createIndexes().

Two details in the tests worth crediting, because both are the kind of thing that usually rots silently:

  • await Task.collection.drop() before the --dry case. clearMongoDb removes documents and leaves indexes, so a sibling test that ran Task.createIndexes() would have left the declared indexes in before and made the final toEqual(before) insensitive to anything the child created. The comment says the first version of the test passed for exactly that reason — that is an assertion whose sides move together, caught by the author.
  • Exercising both flags through spawnSync rather than only the exported function. main() parsing --force and then passing only { dryRun } is precisely the wiring defect no unit test on the export can see.

Carry: behind 1, and the commit is #1958 (ff8de0fa8), touching backend/services/pgBootService.ts and its test — disjoint from this PR's two files. git merge-tree --write-tree against current main is clean. No rebase needed.


One residual finding. Non-blocking, and it is the inverse of the bug this PR fixes.

The guard's premise, stated in its own comment, is that "nothing in backend turns mongoose's autoIndex off … so a boot creates it." As of this PR that is no longer quite true: this file is the first thing in backend that turns autoIndex off. It is safe today — I checked the import graph and only backend/__tests__/service/migrate-task-source-ref-identity.test.js requires the module; no production path does, and package.json invokes it as a standalone script.

But the premise has changed from being protected by absence to being protected by an import graph. mongoose.set is global to the process, so if this module is ever imported from production code — a future consolidated-migrations runner, a boot-time "pending migrations" check — then autoIndex is off app-wide. The pair index is then never created at boot, so this guard blocks forever, and separately every declared index stops being created. The failure is silent and it would present as this guard being broken rather than as its cause.

Cheap to pin: a test asserting no non-test file imports this script. There is precedent in the adjacent TASK-172 work, which used jest --listTests to assert a helper is not collected as a suite. Worth a follow-up, not a change here.

@lilyshen0722

Copy link
Copy Markdown
Contributor Author

Addendum to my gate above — a measurement of the bug this PR fixes, taken against main, plus a production instance of the sequence it protects.

My gate proved the mongoose.set('autoIndex', false) line is load-bearing by mutation. This is the same bug observed from the other side: what main's script does today, run the way an operator runs it.

Seeded a database with only the legacy index, then ran main's CLI with --dry via ts-node in a child process:

stdout: [migrate-task-source-ref-identity] DRY-RUN (no changes written) legacyIndexPresent=true pairIndexPresentBefore=false pairIndexPresentAfter=false droppedLegacy=false
before: ["_id_", "podId_1_sourceRef_1_partial", "podId_1_status_1"]
after:  ["_id_", "podId_1_assignee_1_status_1", "podId_1_sourceRef_1_partial", "podId_1_status_1"]
createdByADryRun: ["podId_1_assignee_1_status_1"]

So on main, --dry writes while printing that it does not.

The detail worth recording is that the reach is a race, not a fixed set. @sprint-impl measured the same run leaving podId_1_sourceRef_1_title_1_partial and podId_1_taskId_1 behind; I got only podId_1_assignee_1_status_1. Mongoose creates the declared indexes one after another in the background and the process exits when the migration returns, so how far the sequence gets varies per run. That means you cannot predict from the outside whether a given past --dry reached the pair index specifically — which is the one this PR's guard reads as evidence that pair-aware code has booted.

This is the stronger argument for the guard than the one I originally gave in consultation. I justified pairIndexPresentBefore on the grounds that only a boot creates the pair index. That was incomplete in a way that mattered: before this PR, the migration's own diagnostic step could create it. So the sequence

  1. operator runs --dry to inspect index state before deploying,
  2. the dry run creates the pair index as a side effect,
  3. the real run sees pairIndexPresentBefore: true and drops the legacy index,

would have passed the guard on evidence the guard's own diagnostic manufactured — the pre-deploy ordering this guard exists to refuse, waved through by the check meant to prevent it. The withheld-path early return and the autoIndex line close that together: one stops the script writing at all, the other stops a withheld run from leaving evidence for its successor.

Production sequence, for the record. The TASK-063 migration was run on prod today using main's script, with a --dry first. It was harmless, and only because the ordering was right: it ran after #1958's deploy had landed pair-aware code, so every index that dry run could have created is one the new boot creates anyway. Had the same two commands run in the other order, step 2 above was live. That is the case this PR removes.

Verdict unchanged: PASS. Recording this because it is a measurement of the defect from the consumer side, and because the reasoning I contributed to the design was weaker than the implementation that resulted from it.

@lilyshen0722
lilyshen0722 added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit b14ba3c Sep 27, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant