Skip to content

[1433] Prisma $transaction isolation level defaults to ReadCommitted allowing phantom reads in allocation rebalance - #1517

Merged
Junirezz merged 8 commits into
Junirezz:mainfrom
ToryMic:fix/1433-prisma-transaction-isolation-level-defaults-to-readcommitted-allowing-phantom-reads-in-allocation-rebalance
Oct 3, 2026
Merged

Junirezz merged 8 commits into
Junirezz:mainfrom
ToryMic:fix/1433-prisma-transaction-isolation-level-defaults-to-readcommitted-allowing-phantom-reads-in-allocation-rebalance

Conversation

@ToryMic

@ToryMic ToryMic commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

📋 Description

Stacked PR. Branched from fix/1431-auth-middleware-accepts-expired-jwt-for-60s-due-to-clocktolerance-misconfiguration (PR #1516), so it carries the #1430 and #1431 work plus the main breakage repair. Review 03a87e18..HEAD for this issue alone. After #1515 → #1516 merge, this collapses to a single commit.

Goal

The allocation rebalance is a read-compute-write over one vault's Allocation rows: read the current allocations, compute the target weights, upsert them. Prisma defaults $transaction to ReadCommitted, so a concurrent rebalance over the same vault can change or insert a row between our read and our write. The loser then overwrites the winner's weights, and the per-vault invariant sum(weight) = 100 is violated — silently, because nothing downstream checks it.

Closes #1433

Worth flagging up front: the invariant was not even representable. Allocation had no weight column and no uniqueness constraint on (vaultId, strategyId), so two rows for the same strategy pair could coexist. Both are fixed here.

Changes

backend/src/services/allocation.ts (new)

  • The whole read-compute-write runs inside one prisma.$transaction(..., { isolationLevel: 'Serializable' }). Every input to the write is read inside that transaction — vault AUM, deletedAt, the current allocation set — so a concurrent commit cannot land mid-rebalance. Serializable is also the strongest level Postgres accepts, and it is the only level Prisma exposes for this repo's SQLite datasource, so one setting is correct on both.
  • A lost race is "someone else went first", not a bug, so it is retried once with exponential backoff and counted in rebalance_serialization_retry_total. Any other error is surfaced untouched. Exhausting the attempts raises a typed RebalanceConflictError carrying the underlying error.
  • Target weights are validated against the invariant before the transaction opens, so a malformed request never contends for a write lock, and re-validated inside it.
  • Amounts are derived with Decimal; the rounding residual goes to the largest-weight target so amounts sum to exactly the AUM.
  • Allocations for strategies dropped from the target list are retired — otherwise their weight persists and breaks the invariant on the next read.
  • getAllocationSummary() is a read and is deliberately left outside the transaction, as are the existing exposureGuardrails.ts query paths.

backend/src/metrics.ts

  • rebalance_serialization_retry_total (partitioned by attempt) — a sustained rise means rebalances collide often enough to need jitter or coarser locking; a value that never moves means the isolation level is not doing its job.
  • rebalance_total (by outcome: ok, ok_after_retry, error, conflict).

backend/prisma/schema.prisma + migration

  • Allocation.weight, @@unique([vaultId, strategyId]), @@index([vaultId, weight]).
  • The migration is hand-written as additive steps — add column with default → fold duplicate amounts into the surviving row → remove duplicates → backfill weight from current exposure → add constraints. Prisma's generated version rebuilds the table (DROP TABLE), which the repo's migration-safety linter rejects. This version drops nothing, renames nothing, and every ADD COLUMN carries a DEFAULT.
  • Backfilling weight from current exposure means the invariant is true for every pre-existing row and the first rebalance after the migration is a no-op relative to reality. A vault whose allocations are all zero gets weight 0 and must be rebalanced before it means anything.

Deliberately not in this PR

POST /api/v1/vault/strategy is the natural caller, but it is a live money-adjacent endpoint and wiring a new write path into it (with its own authorisation, idempotency and rollback story) is a separate decision from fixing the transaction. The acceptance criteria scope this to the transaction, so this PR is the correctness fix only.


🔗 Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • 🔒 Security improvement

🛡️ Risk Assessment

Risk Level

  • 🟠 High: Core contract logic change, access control modification, financial accounting / share math

Allocation weights are economic state, so the tier is High. The blast radius is bounded by the fact that nothing calls this service yet — see the "Deliberately not in this PR" note.

Blast Radius & Impact Analysis

  • Contract storage layout / data key migration involved
  • Value transfer, deposit/withdraw flow, or vault share calculation affected
  • External integration (Oracle, Soroban RPC, Bridge, Token contract) affected
  • Database schema migration or data backfill required
  • Breaking API or interface change affecting downstream clients
  • Zero blast radius (isolated tooling / documentation only)

Detailed Risk & Blast Radius Notes:

1. Schema (the only change that touches existing rows)
   - Allocation.weight is added NOT NULL DEFAULT 0: old code inserting
     without it keeps working, which is what the canary checker requires.
   - @@unique([vaultId, strategyId]) is the one constraint that can *fail* to
     apply: if a deployment somehow has duplicate (vaultId, strategyId) rows,
     CREATE UNIQUE INDEX errors. The migration therefore dedupes first —
     folding the non-surviving rows' amounts into the oldest row of each group
     so no exposure is lost — and the committed test database has been
     checked. Worst case if it were missed: the migration fails loudly and
     nothing is applied, because Prisma runs each migration in a transaction.
   - weight backfill is 100 * amount / sum(amount) within the vault, so the
     invariant holds immediately. A vault whose amounts sum to 0 gets 0.

2. Rebalance semantics
   - Allocations for strategies no longer in the target list are deleted.
     Before this PR nothing wrote Allocation at all, so there is no historical
     behaviour to regress; the backfill sets weight from the amounts that
     exist, so a rebalance that keeps the same target is a no-op.
   - Rebalance is not yet reachable from any route, so no request path
     changes behaviour in this PR.

3. Concurrency
   - Two concurrent rebalances of the SAME vault: exactly one commits, the
     other retries and then commits. Invariant held in every case tested.
   - Two concurrent rebalances of DIFFERENT vaults: no interaction. Tested.
   - REBALANCE_MAX_ATTEMPTS (default 2) and REBALANCE_BASE_BACKOFF_MS
     (default 25) are env-tunable so an operator can widen the retry window
     without a deploy.

4. Latency
   - Every rebalance now takes a Serializable transaction. Under Postgres
     SSI that can add a serialization retry, which is why the retry and the
     metric exist. Rebalances are rare operator actions, not request-path
     traffic, so the added cost is acceptable.

🔄 Rollback Plan

Rollback Strategy & Feasibility

  • Clean Git Revert: Revertable with zero persistent state drift
  • Database Migration Revert: Reversible migration down-script tested and verified
  • Contract Upgrade Rollback: Tested rollback to previous contract WASM hash / implementation
  • Feature Flag / Circuit Breaker: Feature can be toggled off instantly without redeployment
  • Emergency Pause: Contract pause / freeze mechanism available to halt affected functions
  • Forward-Only / Irreversible: State migration cannot be cleanly reversed; emergency recovery runbook linked below

Rollback Trigger Criteria

- rebalance_total{outcome="conflict"} above ~0 for any sustained period
- rebalance_serialization_retry_total climbing monotonically (a rebalance
  loop, not a collision)
- sum(weight) != 100 on any vault after a rebalance (would be a defect in
  this PR, not a tuning problem)
- Rebalance p99 above 1s under normal contention

Step-by-Step Rollback Procedure

  1. git revert the single commit. Nothing calls the service, so reverting removes the behaviour entirely.
  2. To also revert the schema, run a down-migration:
    DROP INDEX "Allocation_vaultId_weight_idx";
    DROP INDEX "Allocation_vaultId_strategyId_key";
    Leaving Allocation.weight in place is the safer half-revert: it is nullable-with-default and harmless, and dropping it would be the one irreversible step.
  3. npx prisma migrate deploy && npm run prisma:generate to resync the client.

⚡ Performance Impact

Performance & Resource Assessment

  • Backend API latency (p95/p99) and database query execution plans verified
  • Database indexing verified for newly queried columns (no table scans)
  • Memory allocation and leak checks verified (no memory leaks in long-running services)
  • No measurable performance impact (documentation, tests, or trivial changes)
  • Smart contract gas / compute units benchmarked — N/A, no Solidity touched
  • Frontend bundle size and TTI verified — N/A, no frontend change

Performance & Gas Profiling Summary:

Per rebalance: 1 findUnique(vault) + 1 findMany(allocations)
              + N upserts + at most 1 deleteMany, inside one transaction.
Previously (hypothetically) the same statements ran at ReadCommitted, so
the statement count is unchanged — only the isolation level and the
retry-on-conflict differ.

New indices, both serving the reads above:
  Allocation_vaultId_strategyId_key  (unique — also the upsert target)
  Allocation_vaultId_weight_idx      (covers "list this vault's
                                      allocations by weight")

Migration cost: one ADD COLUMN (metadata-only on SQLite/Postgres 11+),
two full-table UPDATEs during migration, one DELETE of duplicate rows,
two CREATE INDEX. All on a table that is written only by rebalance, i.e.
low-volume. The migration-safety checker's only findings are warnings
("index creation without CONCURRENTLY", "data backfill or mass update"),
which is the same class of warning as 24 pre-existing migrations.

🔒 SECURITY REVIEW

Smart-contract sections do not apply — no Solidity touched.

  • Input validation: targets are validated before the transaction (non-empty, no duplicate strategyId, finite, non-negative, sum within tolerance of 100). NaN and Infinity weights are rejected explicitly, and Number.isFinite guards the coercion.
  • Authorization: out of scope and deliberately not changed — the service is the data-integrity boundary, and callers keep their existing authorisation. Documented in the module docblock.
  • TOCTOU: every value the write depends on is read inside the same transaction that performs the write, which is the specific race the issue describes.
  • Unbounded loop / DoS: the upsert loop is bounded by the caller's target list, which is validated to be non-empty and free of duplicates; a duplicate list would otherwise have caused unbounded upserts against one key.
  • Invariant enforcement: the sum(weight) = 100 check is a typed error surfaced to the caller, not a log line. A caller that wants a different invariant must change this code deliberately.

Option A: Fixed in This PR ✅

  • Vulnerability identified and resolved
  • Test case added to verify fix
  • Explain fix below:
`prisma.$transaction(fn)` with no options runs at ReadCommitted. The
rebalance read the vault and its allocations, then wrote, so a concurrent
rebalance could commit in between. There was also nothing stopping two
Allocation rows for the same (vaultId, strategyId), which is the phantom
the issue describes.

Fixed by: running the read-compute-write at Serializable, upserting on a
new (vaultId, strategyId) unique key, and retrying exactly once on P2034
so the loser of a race converges instead of corrupting the weights.

📝 Testing

Functional Testing

  • Unit tests added/updated for changes
  • Integration tests passing
  • End-to-end (E2E) tests passing — no frontend or route surface changed
  • Manual testing completed and documented below
cd backend
npx prisma migrate deploy && npx prisma generate
npm run build                                   # tsc → 0 errors
npm run lint                                    # 0 errors
npx jest --runInBand
#   Test Suites: 96 passed, 96 total
#   Tests:       1432 passed, 1432 total        (0 failures)

npm run snapshots:check                                         # ✅
node scripts/check-migrations.js                                # ✅ warnings only
npm run check:migrations:canary                                 # ✅
npx prisma migrate status                                       # Database schema is up to date!

Security Testing

  • Access control test: unauthorized request rejected — N/A by design (no route added); the service boundary is the invariant check, covered by 6 rejection cases
  • Boundary / edge-case test: weight sum 99.99 (rejected), 100 ± tolerance/2 (accepted), duplicate strategyId, empty list, -10, NaN, +Infinity, zero-AUM vault, unknown vault, soft-deleted vault, REBALANCE_MAX_ATTEMPTS override
  • Concurrency test: two rebalances on the same vault, three-way collision, cross-vault isolation, amounts summing to AUM

Test Coverage

  • All new code paths have test coverage
  • Security-critical paths have comprehensive test cases

allocationRebalance.test.ts — 28 unit cases (Prisma double)

Against the issue's acceptance criteria:

  • Serializable is actually requested — asserts $transaction receives a second argument equal to { isolationLevel: 'Serializable' }, that the argument exists at all (so it cannot silently default to ReadCommitted), and that every attempt re-opens at Serializable.
  • Retry once with backoff + metric — first attempt throws P2034, second succeeds; asserts attempts === 2, retried === true, and rebalance_serialization_retry_total === 1. Also: exhausting attempts throws RebalanceConflictError after exactly 2 attempts, REBALANCE_MAX_ATTEMPTS=3 gives 3 attempts and 2 retries, the underlying error is carried, and a non-serialization error (P2002) is not retried and not counted.
  • isSerializationFailure recognises P2034, P2028, SQLITE_BUSY/database is locked, and rejects P2002, connection resets, null and non-objects.
  • Invariant — sum(weight) = 100 reported and persisted; stale allocations retired; nothing deleted when nothing is stale; upsert targets the unique key; over/under-allocation rejected without opening a transaction; negative/NaN/Infinity rejected; missing strategyId rejected.
  • Read path unchanged — getAllocationSummary() runs no transaction and no extra findMany.

allocationRebalanceConcurrency.test.ts — 7 integration cases (real SQLite database)

  • Two concurrent rebalances on the same vault → persisted weights sum to 100 within 1e-6, and rows === distinct (no phantom row), whichever caller won.
  • Three concurrent rebalances with different target sets → invariant still holds.
  • Concurrent rebalances of three different vaults → no cross-interference.
  • Invalid weights during a concurrent window → rejected, and the existing allocation set is byte-for-byte unchanged.
  • Dropping a strategy retires its weight instead of accumulating.
  • Amounts sum to exactly the vault AUM.
  • "one retries": two real concurrent rebalances with a P2034 fault-injected into the first attempt only → exactly one returns retried: true, $transaction was called 3 times for 2 operations, rebalance_serialization_retry_total === 1, and the invariant still holds.

Documented limitation: SQLite takes a database-wide write lock and Prisma waits on SQLITE_BUSY rather than aborting, so a naturally occurring P2034 is not reproducible in CI (verified empirically). The retry assertion therefore injects the failure at the transaction boundary while keeping the reads and writes real. This is stated in the docblock of both test files so nobody later mistakes it for a real race.


🚀 Deployment Notes

Mainnet Readiness

  • This code is ready for production deployment
  • All critical tests pass — see the known-issues note below
  • Security review approved
  • No temporary debug code
  • No TODO comments

Environment variables added (all optional, with defaults)

Variable Default Purpose
REBALANCE_MAX_ATTEMPTS 2 Total transaction attempts (1 = no retry)
REBALANCE_BASE_BACKOFF_MS 25 Backoff base; attempt N waits base * 2^(N-1), capped at REBALANCE_MAX_BACKOFF_MS (500)

Known issues inherited from main (out of scope, flagged for maintainers)

These fail identically on main and are unrelated to this issue — full triage table is on PR #1515:

  1. npm test coverage gate — jest.config.js requires 80% global coverage; main sits at 37.11%, this branch at ~68%. Closing that gap is a repo-wide programme.
  2. npm run prisma:schema-check — ScopedAdminToken index mismatch; the zero-diff fix needs a DROP INDEX, which check-migrations.js classifies as an error, so the two gates contradict each other. (My migration is clean here: prisma migrate diff reports no Allocation difference.)
  3. npm audit / missing pnpm-lock.yaml / CodeQL (Rust, TypeScript) — dependency and infrastructure issues on untouched paths.

✅ Reviewer Checklist

  • PR author completed Risk Assessment and Rollback Plan ✓
  • Performance and gas impact evaluated and verified ✓
  • PR author completed security checklist ✓ (non-contract sections)
  • All findings documented and categorized (fixed/false positive/excluded)
  • Inline security comments are clear and justified
  • Tests cover security-critical code paths
  • No external calls bypass return value checks
  • Access control is properly enforced
  • State updates follow CEI pattern — every value the write depends on is read inside the transaction that performs the write
  • Input validation is comprehensive
  • Follow-up actions (if any) tracked — wiring POST /api/v1/vault/strategy is the next step, deliberately left out of this PR

📞 Questions or Issues?

ToryMic added 7 commits September 29, 2026 23:10
main is currently unbuildable and its test suite is red, so every
back-end CI job fails before it reaches the code under review. The
breakage all traces back to two bad merges (9a15975, 7300a48) that
truncated a Prisma model block, dropped a closing brace, renamed an
imported limiter and replaced the idempotency replay cache.

Prisma schema
- close `WalletTenantAssociation` and separate it from `IdempotencyKey`
  (a missing `}` made `prisma generate` fail, so `@prisma/client`
  stayed an uninitialised stub and every suite crashed on import)
- add `SessionAuditLog`, `Transaction.deletedAt` and
  `WebhookEndpoint.tenantId`, all of which sessionAudit.ts and the
  tenant boundary guard already query
- add the matching migration and refresh the committed SQLite dev.db

Type errors
- import `readsLimiter` in vaultEndpoints.ts (used by GET /receipts)
- restore the idempotency replay cache as idempotencyStore.ts and
  re-export it from idempotency.ts, so transferOrchestrator.ts,
  vaultEndpoints.ts, index.ts and idempotencyRetention.ts resolve
- fix the unterminated `if` in diffSchemaShapes and a possibly-undefined
  `baseline.properties` read in apiContractSnapshots.ts
- build the OTel resource through whichever factory the installed
  @opentelemetry/resources major exposes (v1 `new Resource()`,
  v2 `resourceFromAttributes()`)
- `ZodObject._shape` → `.shape` for Zod 4
- return the strategy-switch 200 response, order receipts by `timestamp`,
  serialise session metadata, and type the request mocks in src/tests

Error contract
- restore `summary`, `errors` and 404 `path` on the error envelope

Tests
- issues711 "newly added required fields" mutated the live schema
  instead of the baseline snapshot, so it asserted against a scenario
  that can never produce the expected diff
An unauthenticated caller could pass `?limit=100000` and have it
forwarded straight into Prisma's `take`, forcing the API to materialise
100k vault rows and OOM-kill the 512MB container.

Chosen behaviour (documented in the route, the spec and the tests):
reject rather than silently clamp. A caller asking for 100000 rows has a
bug or is probing, and quietly returning 50 hides that behind a response
that looks like a complete page. So `limit > 50` fails fast with
`400` and `code: 'LIMIT_EXCEEDED'`, and the body repeats the ceiling.
`page` is the opposite case — a benign mistake — so it is clamped into
1..1000 and the effective value is echoed in the pagination envelope.

- new `middleware/paginationGuard.ts`: `DEFAULT_PAGE_SIZE` (20),
  `MAX_PAGE_SIZE` (50), `MAX_PAGE` (1000), `resolvePagination()` and the
  `enforcePaginationLimits()` middleware
- new `routes/vaults.ts`: `GET /api/v1/vaults`, rate limited, reads
  `limit + 1` rows for the lookahead, and never exposes `tenantId`
- `parsePaginationQuery` gains `maxPage` and clamps out-of-range pages;
  `PaginationQuerySchema.page` accepts a signed integer so those requests
  reach the clamp instead of being rejected
- OpenAPI: reusable `pageSize`/`pageNumber` parameters carrying the
  ceilings, plus a `GET /api/v1/vaults` path documenting the
  LIMIT_EXCEEDED response; `openapi.json` regenerated
- `GET /api/v1/vaults` joins the committed contract snapshots, so the
  response shape is now covered by the backward-compatibility check
- tests spy on the Prisma client to prove no read is ever issued with a
  `take` above the ceiling, that an oversized limit is rejected before
  any query runs, and that the committed `openapi.json` still matches
…request

The verifier applied one shared tolerance to every time claim, so a
token stayed acceptable for 60s past `exp`. A stolen bearer token
therefore kept authorising `POST /vault/:id/withdraw` for a full minute
after the user logged out, and nothing consulted a revocation list at all
on the request path — `/auth/logout` only returned 200 without revoking
anything.

Time claims (backend/src/auth.ts)
- add `assertTimeClaims()`, `TokenExpiredError` and
  `TokenNotYetValidError`, and an optional `nbf` claim on JwtPayload
- `exp` is checked with **zero** tolerance and is inclusive of the
  expiry second (RFC 7519: a token is invalid at and after `exp`)
- only `nbf`/`iat` get `CLOCK_SKEW_TOLERANCE_SECONDS` (5s) of slack,
  for clock skew between the pod and whatever minted the token

Revocation (backend/src/tokenRevocation.ts)
- `RevocationStore` gains `revokeWalletBefore()` and
  `isWalletRevokedBefore()`: a per-wallet high-water mark rather than an
  id list, so `logout-all` stays O(1) per wallet. The marker never moves
  backwards, so a later, broader revocation always wins
- `revokeAllForWallet()` no longer just deletes records — it recorded
  nothing, which silently un-revoked every token it touched
- Redis store implements the new methods and falls back to its in-process
  store on error, including for `revokeAllForWallet` (was `return 0`)
- the store is now actually wired to Redis when `REDIS_URL` is set, so a
  logout handled by one pod is visible to the others

Request path
- `requireAuth` checks the revocation list on every authenticated request
  and answers 401 `TOKEN_REVOKED`; it fails **closed** with 503 if the
  store itself is unreachable rather than downgrading to "token is fine"
- `/auth/logout` revokes the presented `jti` and, when a refreshToken is
  supplied, the whole refresh family
- `/auth/logout-all` writes the wallet marker and revokes every refresh
  family for the wallet
- OpenAPI: the `bearerAuth` scheme documents the zero-tolerance `exp`,
  the 5s `nbf` skew and the per-request revocation check

Tests: 29 new cases in `jwtExpiryAndRevocation.test.ts`, including a
token signed to expire *now* and read 10s later with a mocked
`Date.now` (401, not 200), the 59s case the old 60s tolerance allowed, a
5s `nbf` skew that must still pass, logout/logout-all over HTTP, and the
Redis-error fallback path.
…comment

Two review-level corrections to the revocation path added in the previous
commit:

- `isAccessTokenRevoked` looked the wallet marker up under the raw `sub`
  claim. Revocations are written under the canonical (upper-case) address
  via `normalizeWalletAddress`, so a token whose `sub` happened to be
  lower-case would not have matched a wallet-wide revocation. Normalise
  the lookup key the same way the write path does.
- requireAuth's comment claimed a `res.headersSent` guard that was never
  implemented. Replace it with the reasoning that actually matters: Express
  ignores middleware return values, so returning before `next()` is safe for
  route usage, and the single direct caller passes a synchronous next().
…s path

`transferOrchestrator.test.ts` exercises the happy path end to end, but
with no `REDIS_URL` the `redisClientManager` never reports ready, so
`IdempotencyStore.redis` returns null and every Redis helper is skipped —
the store's Redis branch, and more importantly its error fallback, had no
coverage at all.

These tests drive the store through a stub client so they reach: the
Redis get/set/del round trip, conflict detection across instances, the
read/write/delete failure fallbacks, and `pruneStaleKeys` across expired
TTL, unparsable payloads and dryRun. Plus the in-process semantics that
money-moving code relies on: single execution, replay, fingerprint
conflict, in-flight coalescing, and that a rejected operation frees the
pending slot for a retry.

idempotencyStore.ts: 70.9% -> 94.5% statements, 50.9% -> 81.8% branches.
…d-allows-limit-100000-to-oom-the-api' into fix/1431-auth-middleware-accepts-expired-jwt-for-60s-due-to-clocktolerance-misconfiguration
…034 retry

The rebalance is a read-compute-write over one vault's `Allocation`
rows: read the current allocations, compute target weights, upsert them.
Prisma defaults `$transaction` to ReadCommitted, so a concurrent
rebalance over the same vault could change a row between our read and our
write. The loser then overwrote the winner's weights and
`sum(weight) = 100` was violated — silently, because nothing downstream
checks it.

The invariant was not even representable: `Allocation` had no `weight`
column, and nothing stopped two rows existing for the same
(vaultId, strategyId).

- new `services/allocation.ts`
  - the whole read-compute-write runs inside one
    `prisma.$transaction(..., { isolationLevel: 'Serializable' })`; every
    input to the write is read inside that transaction, so a concurrent
    commit cannot land mid-rebalance
  - a lost race is a "someone else went first", not a bug, so it is
    retried once with exponential backoff and counted in
    `rebalance_serialization_retry_total`. Any other error is surfaced
    untouched; exhausting the attempts raises a typed
    `RebalanceConflictError` carrying the underlying error
  - targets are validated against the invariant *before* the transaction
    opens, so a malformed request never contends for a write lock
  - amounts are derived with Decimal and the rounding residual is given to
    the largest-weight target, so the amounts sum to exactly the AUM
  - allocations for strategies dropped from the target list are retired,
    otherwise their weight would persist and break the invariant
  - `getAllocationSummary()` is a read and is deliberately left outside
    the transaction, as are the existing `exposureGuardrails.ts` queries
- metrics: `rebalance_serialization_retry_total` (partitioned by attempt)
  and `rebalance_total` (by outcome)
- schema: `Allocation.weight`, `@@unique([vaultId, strategyId])` and a
  `(vaultId, weight)` index. The migration is additive
  (add column with default -> fold duplicate amounts -> remove duplicates ->
  backfill weight from current exposure -> add constraints), so no row is
  dropped and no existing ADD COLUMN lacks a DEFAULT, keeping it inside
  the repo's migration-safety policy

Tests: 28 unit cases (isolation level, retry/metric/typed conflict,
invariant, vault state, amount derivation, read path) plus 7 integration
cases against the real database covering two concurrent rebalances on the
same vault, repeated collisions, cross-vault isolation, and that amounts
sum to AUM. The retry assertion fault-injects P2034 on the first attempt
because SQLite waits on SQLITE_BUSY rather than aborting, so a real
P2034 is not reproducible in CI — that limitation is documented in both
test files.
@ToryMic

ToryMic commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

🧭 CI triage — same 12 pre-existing failures as #1515 / #1516, none from this diff

main is red on all of these (verified against the main push run 36430911829 at 711c4332). Every red check below fails identically on main.

Check main This branch Blocked by
Backend test suites 33 failed suites / 49 failed tests 0 failed / 96 suites, 1432 tests — fixed by the stacked commits
Backend build (tsc) fails ✅ pass — fixed by the stacked commits
backend-governance ❌ ❌ prisma:schema-check, only the pre-existing ScopedAdminToken index — the log shows no Allocation difference, so the new migration is in sync
Backend lint + test ❌ ❌ npm audit --audit-level=high (5 high, needs breaking majors)
Backend Test Coverage (>= 80%) ❌ ❌ coverage threshold
Frontend Test Coverage (>= 70%) ❌ ❌ frontend untouched
CodeQL (TypeScript) ❌ ❌ frontend/package-lock.json out of sync
CodeQL (Rust) / Cargo Security Audit ❌ ❌ cargo build failure in contracts/ (untouched)
Governance & PR Standards / Dependency Security Scan / Dependency Vulnerability Audit ❌ ❌ repo has no pnpm-lock.yaml, so pnpm install --frozen-lockfile can never succeed
NPM Audit / GitHub Dependency Review ❌ ❌ dependency advisories

New in this PR, and green

  • prisma migrate diff reports zero difference for Allocation — the new weight column, the (vaultId, strategyId) unique key and the (vaultId, weight) index are all satisfied by the committed migration.
  • node scripts/check-migrations.js → exit 0 (warnings only: index creation without CONCURRENTLY and data backfill, the same class as 24 pre-existing migrations).
  • npm run check:migrations:canary → passed. The first draft failed it, because the canary checker greps raw file text and a comment of mine said "no DROP TABLE / DROP COLUMN". Reworded; the migration now drops and renames nothing and every ADD COLUMN has a DEFAULT.
  • npm run snapshots:check, npm run build, npm run lint → all pass.

✅ Final local verification

cd backend
npx prisma migrate deploy && npx prisma generate
npm run build                                   # tsc → 0 errors
npm run lint                                    # 0 errors
npx jest --runInBand
#   Test Suites: 96 passed, 96 total
#   Tests:       1432 passed, 1432 total        (0 failures)

npx prisma migrate status                       # Database schema is up to date!
npm run snapshots:check                                         # ✅
node scripts/check-migrations.js                                # ✅
npm run check:migrations:canary                                 # ✅

Branch status: upstream/main is an ancestor of this branch (git rev-list --count HEAD..upstream/main → 0), so the PR is MERGEABLE with zero conflicts. Stacked on #1516 → #1515; once those merge, this PR's diff collapses to a single commit.

…defaults-to-readcommitted-allowing-phantom-reads-in-allocation-rebalance
@Junirezz
Junirezz merged commit f847960 into Junirezz:main Oct 3, 2026
6 of 17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Prisma $transaction isolation level defaults to ReadCommitted allowing phantom reads in allocation rebalance

2 participants