From d1d75055396cff3ea15c6ef6c45d9363440b4175 Mon Sep 17 00:00:00 2001 From: BigSimmo <87357024+BigSimmo@users.noreply.github.com> Date: Tue, 18 Aug 2026 12:30:08 +0800 Subject: [PATCH 1/2] docs(issues): reconcile 17 queued ledger requests; mark G1 merged and record the S1d canary/#2065 revert state MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One fresh-base issues:reconcile over the complete pending inbox (17 requests, 2 cancellation decisions): closes #J912J9 (governance question) and #0MSNT8 (G1 task) via G1 PR #2053, closes #6BG9X2 (R2+R3) via S1c PR #2052; adds the three S5 adversarial-divergence pins, the RAG_TELEMETRY_EXTENDED owner decision, and the non-RAG captures queued since #2045. HANDOVER: G1 merged (125e98526), S1d merged with its red canary bisected to #2065 (revert PR #2088), S1c follow-up note, S2 blocker re-keyed; COORDINATION §7 wave-1 status. Visual register refreshed. Co-Authored-By: Claude Fable 5 --- .../34c0f9bf-22fe-495b-a828-f73bbd4cfddf.json | 0 .../3565972b-02c2-4ad8-b899-3fbde0d54725.json | 0 .../44e79b9e-5530-4b14-a431-b7a87683cf30.json | 0 .../48745dbc-b7a1-40d5-aec7-818274814291.json | 0 .../51e687a9-406a-4f2e-bb70-f6ba9aebef53.json | 0 .../5cc39bc6-4cf4-4a96-9ce7-a0ee46020b90.json | 0 .../632c50f9-6e6a-4247-8b75-11572137e579.json | 0 .../63e7cebb-9a42-4c98-9520-9a1b7f5874a8.json | 0 .../71399b63-cce0-46d4-9fc3-c5fdc617289d.json | 0 .../776405e0-c2d9-4dec-b688-e26c22143f04.json | 0 .../c68dae81-559c-45ad-af70-f1b0334a49c3.json | 0 .../d6ce8a1d-518d-48fc-8a9a-7796f060a46e.json | 0 .../d92786de-de31-4118-84cc-0a9098e7f2e0.json | 0 .../e790f80b-efb3-4684-8195-77ff872ad014.json | 0 .../e7f18a92-046d-4fb7-98be-d6aca938657d.json | 0 .../e8a5480e-0b0e-4ab1-b7bd-6090699dd7c1.json | 0 .../fb7d42c0-ee67-426e-907e-f2c306afc38b.json | 0 docs/outstanding-issues.md | 16 ++++++-- docs/rag-improvement/COORDINATION.md | 16 ++++---- docs/rag-improvement/HANDOVER.md | 38 +++++++++---------- 20 files changed, 40 insertions(+), 30 deletions(-) rename docs/outstanding-issues-inbox/{ => applied}/34c0f9bf-22fe-495b-a828-f73bbd4cfddf.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/3565972b-02c2-4ad8-b899-3fbde0d54725.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/44e79b9e-5530-4b14-a431-b7a87683cf30.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/48745dbc-b7a1-40d5-aec7-818274814291.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/51e687a9-406a-4f2e-bb70-f6ba9aebef53.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/5cc39bc6-4cf4-4a96-9ce7-a0ee46020b90.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/632c50f9-6e6a-4247-8b75-11572137e579.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/63e7cebb-9a42-4c98-9520-9a1b7f5874a8.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/71399b63-cce0-46d4-9fc3-c5fdc617289d.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/776405e0-c2d9-4dec-b688-e26c22143f04.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/c68dae81-559c-45ad-af70-f1b0334a49c3.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/d6ce8a1d-518d-48fc-8a9a-7796f060a46e.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/d92786de-de31-4118-84cc-0a9098e7f2e0.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/e790f80b-efb3-4684-8195-77ff872ad014.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/e7f18a92-046d-4fb7-98be-d6aca938657d.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/e8a5480e-0b0e-4ab1-b7bd-6090699dd7c1.json (100%) rename docs/outstanding-issues-inbox/{ => applied}/fb7d42c0-ee67-426e-907e-f2c306afc38b.json (100%) diff --git a/docs/outstanding-issues-inbox/34c0f9bf-22fe-495b-a828-f73bbd4cfddf.json b/docs/outstanding-issues-inbox/applied/34c0f9bf-22fe-495b-a828-f73bbd4cfddf.json similarity index 100% rename from docs/outstanding-issues-inbox/34c0f9bf-22fe-495b-a828-f73bbd4cfddf.json rename to docs/outstanding-issues-inbox/applied/34c0f9bf-22fe-495b-a828-f73bbd4cfddf.json diff --git a/docs/outstanding-issues-inbox/3565972b-02c2-4ad8-b899-3fbde0d54725.json b/docs/outstanding-issues-inbox/applied/3565972b-02c2-4ad8-b899-3fbde0d54725.json similarity index 100% rename from docs/outstanding-issues-inbox/3565972b-02c2-4ad8-b899-3fbde0d54725.json rename to docs/outstanding-issues-inbox/applied/3565972b-02c2-4ad8-b899-3fbde0d54725.json diff --git a/docs/outstanding-issues-inbox/44e79b9e-5530-4b14-a431-b7a87683cf30.json b/docs/outstanding-issues-inbox/applied/44e79b9e-5530-4b14-a431-b7a87683cf30.json similarity index 100% rename from docs/outstanding-issues-inbox/44e79b9e-5530-4b14-a431-b7a87683cf30.json rename to docs/outstanding-issues-inbox/applied/44e79b9e-5530-4b14-a431-b7a87683cf30.json diff --git a/docs/outstanding-issues-inbox/48745dbc-b7a1-40d5-aec7-818274814291.json b/docs/outstanding-issues-inbox/applied/48745dbc-b7a1-40d5-aec7-818274814291.json similarity index 100% rename from docs/outstanding-issues-inbox/48745dbc-b7a1-40d5-aec7-818274814291.json rename to docs/outstanding-issues-inbox/applied/48745dbc-b7a1-40d5-aec7-818274814291.json diff --git a/docs/outstanding-issues-inbox/51e687a9-406a-4f2e-bb70-f6ba9aebef53.json b/docs/outstanding-issues-inbox/applied/51e687a9-406a-4f2e-bb70-f6ba9aebef53.json similarity index 100% rename from docs/outstanding-issues-inbox/51e687a9-406a-4f2e-bb70-f6ba9aebef53.json rename to docs/outstanding-issues-inbox/applied/51e687a9-406a-4f2e-bb70-f6ba9aebef53.json diff --git a/docs/outstanding-issues-inbox/5cc39bc6-4cf4-4a96-9ce7-a0ee46020b90.json b/docs/outstanding-issues-inbox/applied/5cc39bc6-4cf4-4a96-9ce7-a0ee46020b90.json similarity index 100% rename from docs/outstanding-issues-inbox/5cc39bc6-4cf4-4a96-9ce7-a0ee46020b90.json rename to docs/outstanding-issues-inbox/applied/5cc39bc6-4cf4-4a96-9ce7-a0ee46020b90.json diff --git a/docs/outstanding-issues-inbox/632c50f9-6e6a-4247-8b75-11572137e579.json b/docs/outstanding-issues-inbox/applied/632c50f9-6e6a-4247-8b75-11572137e579.json similarity index 100% rename from docs/outstanding-issues-inbox/632c50f9-6e6a-4247-8b75-11572137e579.json rename to docs/outstanding-issues-inbox/applied/632c50f9-6e6a-4247-8b75-11572137e579.json diff --git a/docs/outstanding-issues-inbox/63e7cebb-9a42-4c98-9520-9a1b7f5874a8.json b/docs/outstanding-issues-inbox/applied/63e7cebb-9a42-4c98-9520-9a1b7f5874a8.json similarity index 100% rename from docs/outstanding-issues-inbox/63e7cebb-9a42-4c98-9520-9a1b7f5874a8.json rename to docs/outstanding-issues-inbox/applied/63e7cebb-9a42-4c98-9520-9a1b7f5874a8.json diff --git a/docs/outstanding-issues-inbox/71399b63-cce0-46d4-9fc3-c5fdc617289d.json b/docs/outstanding-issues-inbox/applied/71399b63-cce0-46d4-9fc3-c5fdc617289d.json similarity index 100% rename from docs/outstanding-issues-inbox/71399b63-cce0-46d4-9fc3-c5fdc617289d.json rename to docs/outstanding-issues-inbox/applied/71399b63-cce0-46d4-9fc3-c5fdc617289d.json diff --git a/docs/outstanding-issues-inbox/776405e0-c2d9-4dec-b688-e26c22143f04.json b/docs/outstanding-issues-inbox/applied/776405e0-c2d9-4dec-b688-e26c22143f04.json similarity index 100% rename from docs/outstanding-issues-inbox/776405e0-c2d9-4dec-b688-e26c22143f04.json rename to docs/outstanding-issues-inbox/applied/776405e0-c2d9-4dec-b688-e26c22143f04.json diff --git a/docs/outstanding-issues-inbox/c68dae81-559c-45ad-af70-f1b0334a49c3.json b/docs/outstanding-issues-inbox/applied/c68dae81-559c-45ad-af70-f1b0334a49c3.json similarity index 100% rename from docs/outstanding-issues-inbox/c68dae81-559c-45ad-af70-f1b0334a49c3.json rename to docs/outstanding-issues-inbox/applied/c68dae81-559c-45ad-af70-f1b0334a49c3.json diff --git a/docs/outstanding-issues-inbox/d6ce8a1d-518d-48fc-8a9a-7796f060a46e.json b/docs/outstanding-issues-inbox/applied/d6ce8a1d-518d-48fc-8a9a-7796f060a46e.json similarity index 100% rename from docs/outstanding-issues-inbox/d6ce8a1d-518d-48fc-8a9a-7796f060a46e.json rename to docs/outstanding-issues-inbox/applied/d6ce8a1d-518d-48fc-8a9a-7796f060a46e.json diff --git a/docs/outstanding-issues-inbox/d92786de-de31-4118-84cc-0a9098e7f2e0.json b/docs/outstanding-issues-inbox/applied/d92786de-de31-4118-84cc-0a9098e7f2e0.json similarity index 100% rename from docs/outstanding-issues-inbox/d92786de-de31-4118-84cc-0a9098e7f2e0.json rename to docs/outstanding-issues-inbox/applied/d92786de-de31-4118-84cc-0a9098e7f2e0.json diff --git a/docs/outstanding-issues-inbox/e790f80b-efb3-4684-8195-77ff872ad014.json b/docs/outstanding-issues-inbox/applied/e790f80b-efb3-4684-8195-77ff872ad014.json similarity index 100% rename from docs/outstanding-issues-inbox/e790f80b-efb3-4684-8195-77ff872ad014.json rename to docs/outstanding-issues-inbox/applied/e790f80b-efb3-4684-8195-77ff872ad014.json diff --git a/docs/outstanding-issues-inbox/e7f18a92-046d-4fb7-98be-d6aca938657d.json b/docs/outstanding-issues-inbox/applied/e7f18a92-046d-4fb7-98be-d6aca938657d.json similarity index 100% rename from docs/outstanding-issues-inbox/e7f18a92-046d-4fb7-98be-d6aca938657d.json rename to docs/outstanding-issues-inbox/applied/e7f18a92-046d-4fb7-98be-d6aca938657d.json diff --git a/docs/outstanding-issues-inbox/e8a5480e-0b0e-4ab1-b7bd-6090699dd7c1.json b/docs/outstanding-issues-inbox/applied/e8a5480e-0b0e-4ab1-b7bd-6090699dd7c1.json similarity index 100% rename from docs/outstanding-issues-inbox/e8a5480e-0b0e-4ab1-b7bd-6090699dd7c1.json rename to docs/outstanding-issues-inbox/applied/e8a5480e-0b0e-4ab1-b7bd-6090699dd7c1.json diff --git a/docs/outstanding-issues-inbox/fb7d42c0-ee67-426e-907e-f2c306afc38b.json b/docs/outstanding-issues-inbox/applied/fb7d42c0-ee67-426e-907e-f2c306afc38b.json similarity index 100% rename from docs/outstanding-issues-inbox/fb7d42c0-ee67-426e-907e-f2c306afc38b.json rename to docs/outstanding-issues-inbox/applied/fb7d42c0-ee67-426e-907e-f2c306afc38b.json diff --git a/docs/outstanding-issues.md b/docs/outstanding-issues.md index f63a0b2c52..f17e5f24e5 100644 --- a/docs/outstanding-issues.md +++ b/docs/outstanding-issues.md @@ -209,13 +209,20 @@ removed after current-main verification; it is not missing recommended work. | #341 | P2 | task | Route the remaining ~22 unguarded source-slice test windows through the guarded helper | PARTIALLY DONE 2026-08-15 by PR #1985, which added tests/helpers/source-contract.ts and migrated the three worst files. The hazard this closes is a silent pass, not fragility: the idiom source.slice(source.indexOf(start), source.indexOf(end)) returns -1 for a missing end marker, and slice(n, -1) does not throw — it returns the rest of the file bar one character. A renamed end marker therefore converts a scoped assertion into a whole-file assertion and every positive toContain in it keeps passing for the wrong reason. The mirror case, a missing start, yields slice(-1, n) so every negative assertion passes vacuously. Neither shows up as a failure. The helper throws on a missing start, a missing end, and an AMBIGUOUS start (a window anchored on a string that appears twice silently covers only the first hit). MIGRATED: search-route-ownership.test.ts (three windows, including one whose end marker is an indentation depth and one anchored on a comment string), document-detail-performance.test.ts (end marker was the next literal 'useEffect' token, of which DocumentViewer has several), therapy-compass-responsive-contract.test.ts. REMAINING, roughly 22 windows across ~11 files, none migrated: audit-navigation-auth-regressions.test.ts is the densest and was deliberately skipped because PR #1983 edits the same file and the anti-churn rule prefers one late sync to a merge fight — do it once #1983 lands. Also tools-search-directions-mockups.test.ts, in-page-nav-playwright-contract.test.ts and document-section-nav-contract.test.ts (the last two slice ui-smoke.spec.ts between Playwright test titles, so renaming OR reordering an unrelated spec silently rescopes them), and rag-retrieval-parallelism.test.ts, which is left for a session that flags the RAG surface first per AGENTS.md. A REAL COVERAGE HOLE was found while surveying and is NOT yet fixed: audit-navigation-auth-regressions.test.ts around line 285 anchors on '{showUniversalAlsoMatches &&', which occurs TWICE in ClinicalDashboard.tsx (the second around line 3832 sits outside the window), so its not.toContain check does not enforce the named contract across the file. The new helper would reject that anchor outright, which is how it was found. Fix it in the same pass as that file's migration. STOP: do not loosen the demo-data boundary pins in favourites-demo-boundary.test.ts. The exact conditional-spread form '...(demoMode ? prototypeFavouriteItems : [])' with its paired negative is the live-vs-demo privacy contract, and its strictness is the point. Likewise leave the SQL windows ending on '$$;' and header-scroll-hide-contract.test.ts anchoring on the matching '' closing tag — those are true structural terminators and are the model the rest should move toward. | session 2026-08-15; PR #1985; tests/helpers/source-contract.ts | 2026-08-15 | | #342 | P2 | issue | Recurring 'Unhandled server request error' on /api/search and /api/search/universal is untriaged | Three Sentry issue groups in clinibase-xz over 24h (JAVASCRIPT-NEXTJS-Y, -Z, -10), 17 events, 0 users impacted, all titled 'Error: Unhandled server request error' with culprit chunk 1261.js:2:4801. Top frames are /api/search/route.js and /api/search/universal/route.js. First seen 2026-08-14T08:44:37Z on release c9b089c92c975297c10649b005401d5ae337cf48, roughly six hours BEFORE PR #1946 merged, so it is not caused by the retrieval row contract; the post-merge group is the same error refingerprinted by the release change. The error string does not appear anywhere in repo source, so it likely originates in a dependency or an instrumentation wrapper — origin unidentified. Nobody owns this. Next step: identify what throws it, then decide whether it is a bot/scanner artefact or a real request-handling gap. | Sentry clinibase-xz, reviewed 2026-08-15 | 2026-08-15 | | #343 | P3 | task | Make the retrieval row contract's source_metadata pin structural, not data-guaranteed | rag-row-contracts.ts pins source_metadata to a JSON object via z.record(...), but documents.metadata is bare jsonb and permits arrays and scalars. Measured against the live project (sjrfecxgysukkwxsowpy) on 2026-08-15: all 2851 documents are object-typed, so nothing breaks today and no live errors exist. The guarantee is data, not schema — a future ingest path could violate it and take retrieval down for that document's chunks. Fix is either a check (jsonb_typeof(metadata) = 'object') constraint on public.documents, or loosening the pin. Every other required field in that contract is backed by a not-null constraint. | PR #1946 review + live Supabase verification 2026-08-15 | 2026-08-15 | -| #6BG9X2 | P2 | task | R2 + R3: claim-support strictness rejects verbatim-faithful guideline restatements (directive normativity; topic-overlap dilution) — packet S1c | R2: normativeDirectiveActions in src/lib/rag/rag-claim-support.ts has no pattern for 'usual / recommended ... dose is ...' guideline phrasing, so imperative claims ('start lithium at 500 mg nocte') fail against descriptive norms; reproduced offline on the EMHS lithium chunk. R3: a claim synthesising two adjacent source bullets fails the >=50% single-segment topic-overlap requirement even when every atom matches; reproduced offline. Fix R2 with a small pattern addition plus adversarial negatives; MEASURE R3 before loosening. One PR after S1b merges and its canary is green; RAG impact behaviour change; canary pair. Packet: docs/rag-improvement/HANDOVER.md S1c. Stop: no grounding-gate weakening beyond the two named artefacts. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | -| #0MSNT8 | P3 | task | Governance Option B decided: tag document-summary rows with similarity_origin 'document_context', keep the confidence label — implement as packet G1 | Owner decision 2026-08-17 on the question queued by #212 tranche 3 (buildDocumentSummaryResults stamps similarity: 1 with no similarity_origin; deriveConfidence excludes only synthetic_text). Option B: add a new similarity_origin value document_context to the union in src/lib/types.ts and to src/lib/answer-stream-contract.ts, stamp it in buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts), keep deriveConfidence unchanged and pin it with discriminating tests, keep rag.ts synthetic_similarity_count from counting it, and update docs/clinical-hazard-analysis.md H5a. Rationale: the only caller is the document-summary route where the query is the document itself and citation support is still verified; Option A (tag as synthetic_text so summaries cap at medium) rejected as a label downgrade without a measured safety gain. RAG impact: no retrieval behaviour change; no canary. Packet: docs/rag-improvement/HANDOVER.md G1. Closes the P1 governance question row once landed. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | | #DP6M3G | P1 | task | R1: unbudgeted strong escalation makes provider_timeout the dominant lithium fallback — route the dosing class to strong before the deadline (packet S1b) | S1 (PR #2022) post-fix live probes: 'Lithium dosing?' 4/4 source-only, 3/4 as provider_timeout. fast_unsupported_retry_strong launches a strong generation into the fast route's leftover ~10-13 s; only the truncation self-heal is deadlineAllowsGenerationRetry-gated. Ladder rung 3 (README A1): route medication_dose_risk / dosing to the strong route in chooseAnswerRoute (src/lib/rag/rag-routing.ts) BEFORE the route deadline is created — not in shouldRetryWithStrongAfterFast, and NOT a budget change (#231 stop condition stands). Own PR, RAG impact behaviour change, canary pair, Clinical Governance Preflight, check:production-readiness. Owner decided 2026-08-17 this lands before S2 (A2/A3 add length; length under the unbudgeted retry pushes more dosing queries into timeout). Packet: docs/rag-improvement/HANDOVER.md S1b. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | | #BTVMVK | P2 | issue | Recurring 'Unhandled server request error' on /api/search and /api/search/universal in Sentry — unowned, pre-dates #1946 | Sentry (clinibase-xz): three issue groups in 24h, 17 events, 0 users impacted, on /api/search and /api/search/universal, all with culprit chunk 1261.js:2:4801. First seen 2026-08-14T08:44:37Z on release c9b089c9, about six hours before PR #1946 merged, so not caused by the row contracts. Recorded in the #212 tranche-1 handover; never captured durably until now. Next: triage the Sentry groups (read-only Sentry MCP or dashboard), map chunk 1261.js to source via the release's source maps, reproduce locally with the request shapes Sentry recorded. Stop: do not silence the error path; search routes are clinical output. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | -| #J912J9 | P1 | issue | Decide whether a fabricated similarity of 1 on document-summary rows may earn the high confidence label a clinician reads | buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts) stamps similarity: 1 on document-summary rows -- a fabricated score, not a measured cosine -- and does NOT set similarity_origin. deriveConfidence (src/lib/rag/rag-answer-support.ts:32-36) computes strongestNonSynthetic by EXCLUDING rows tagged similarity_origin === "synthetic_text", so an untagged fabricated 1.0 IS counted, and line 36 gates the "high" verdict on strongestNonSynthetic >= 0.82 with at least 2 accepted citations. A document summary therefore reaches the "high" confidence label on a score nobody measured. Note the mechanism precisely: buildDocumentSummaryResults does NOT set the tag -- the three call sites that DO tag synthetic scores are in rag-candidate-sources.ts (lines 626, 721, 934), and it is the ABSENCE of the tag here that admits the fabricated score. docs/clinical-hazard-analysis.md H5a records the adjacent document-lookup fast-path hazard. QUESTION: should a fabricated similarity count toward "high"? Adding the tag would demote these answers to medium or low. This is a clinical-governance decision before it is an engineering one: either way it needs its own design, discriminating offline tests that separate tagged from untagged rows, and a live eval-canary before/after pair per docs/rag-behaviour/. Deliberately excluded from PR #1981 and from the #212 tranche 3 PR rather than bundled, because it changes clinical output. Verified on main at d0276718. | ledger #212 tranche 3 (src/app/api row contracts), session 2026-08-17 | 2026-08-17 | | #TYJ0XP | P3 | rec | eval-canary.yml is post-merge only (repository_dispatch + Sunday cron, no ref input) — record this in docs/rag-behaviour so sessions stop expecting a branch canary | .github/workflows/eval-canary.yml triggers on repository_dispatch type eval-canary and schedule cron 0 18 * * 0; it always loads the default branch and has no workflow_dispatch or ref input, so a canary can only ever measure main. Consequence for the RAG programme: 'canary pair' means latest green run on main before the merge -> a dispatch after the merge (gh api repos/BigSimmo/Database/dispatches -f event_type=eval-canary), compared with npm run eval:retrieval:compare -- --fail-on-regression on the eval-canary-output artifacts. Recorded correctly for S1 (baseline run 31964560921 -> post run 32025082010, zero regressions). Next: add one paragraph to docs/rag-behaviour/safeguards.md canary protocol; no workflow change. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | | #ND10QT | P3 | rec | source_metadata pin in rag-row-contracts.ts is data-backed only — add check (jsonb_typeof(metadata) = 'object') on documents or loosen the pin | src/lib/rag/rag-row-contracts.ts requires source_metadata to be a JSON object (z.record) while documents.metadata jsonb permits arrays and scalars. Live query 2026-08-14 (project sjrfecxgysukkwxsowpy): select jsonb_typeof(metadata), count(*) from public.documents group by 1 returned a single row, object = 2851, so nothing breaks today — but the guarantee is data, not schema, and a future ingest could make retrieval throw RetrievalRowShapeError for that document. Options: a check constraint via a new migration (role postgres; run check:migration-role) or loosen the pin. The module's doc comment also claims every required field is 'not null in supabase/schema.sql', which is true for nine fields and false for source_metadata — one-clause docs fix in the same PR. | RAG programme coordinator, S1 PR #2022 residuals and #212 handover follow-ups, 2026-08-17 | 2026-08-17 | +| #4TBHS8 | P3 | issue | Advisory UI mockup spec 'phone filter sheet follows the shared local-filter behavior' fails on main | tests/ui-tools-search-mode-mockup.spec.ts:188 fails at line 198 waiting for '2 showing' inside [data-testid=tools-search-filter-sheet] after searching 'Safety' at 390px. Reproduced locally under --project=chromium-mockups on BOTH claude/card-review-optimize-h0pidc and origin/main (dc7e518), so it is pre-existing and NOT caused by the card branch — attribution was checked before any fix was attempted. It surfaced now only because the ui-advisory lane fires on advisory_ui_changed (a mockup surface changed or the flake ledger is non-empty) and had been skipped on every earlier run of that PR. It is non-blocking: ui-advisory carries continue-on-error true and is absent from pr-required's needs list in ci.yml, and verify:ui excludes @mockup via --grep-invert, which is why a 429-pass local run never touched it. Next: open the trace at test-results/ui-tools-search-mode-mocku-b6163-hared-local-filter-behavior-chromium-mockups/trace.zip and decide whether the expected count of 2 is stale against the current tools catalogue or the facet hint genuinely miscounts; the mockup renders the production ToolsSearchResultsPage, so a real miscount would affect /tools too. Stop: do not change the expected number to match observed output without establishing which is correct. | session 2026-08-18; PR #2060 Advisory UI run 32090358678; reproduced on origin/main dc7e518 | 2026-08-18 | +| #2AB2NJ | P3 | task | Owner decision: enable RAG_TELEMETRY_EXTENDED (verification_latency_ms projection) in production once a dashboard consumer exists | Packet S5 (PR #2056, merge 093f9340c) landed the B1 telemetry gap assessment: the one proven gap is verification_latency_ms, now persisted behind RAG_TELEMETRY_EXTENDED (typed, default false) via the allow-listed projection module with canary-absence tests. Enabling it in production is an owner decision gated on a dashboard consumer existing (no consumer today), and is a Railway env change (provider-backed, explicit approval; rollback = set false). Next: when a dashboard question needs verification latency, set RAG_TELEMETRY_EXTENDED=true on the Database service after confirming the canary-absence tests are still green on main. Stop: do not enable speculatively; do not add unproven fields. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | +| #C2D9JF | P2 | issue | Adversarial divergence (S5 harness pin): scope-other-owner-document — abstains in substance but the review fallback still cites in-scope evidence | Pinned in tests/rag-adversarial-harness.test.ts KNOWN_DIVERGENCES (self-expiring). Observed shape: grounded false, confidence unsupported, but cited chunk ids [syn-scope-owner-a] — the answer correctly abstains from the other-owner document, yet the review fallback attaches an in-scope citation to an unsupported answer. Fixture: scripts/fixtures/rag-adversarial-cases.v1.json case scope-other-owner-document (category scope_or_tenant). Tenancy/no-read invariant held (the other-owner content is never read). Next: decide whether an unsupported abstention may carry any citation; if not, strip citations on the abstention path (RAG-surface change; own PR; harness pin flips; canary pair). Stop: do not delete the pin without the behaviour change. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | +| #71NT23 | P2 | task | No mobile-WebKit or display-mode:standalone Playwright project exists; phone coverage is a narrow viewport on desktop engines | playwright.config.ts defines chromium, chromium-mockups, firefox and webkit, and the webkit project uses devices["Desktop Safari"]. Every phone assertion in the suite is therefore a narrow viewport on a desktop engine, and nothing exercises display-mode: standalone at all — despite globals.css:3570-3606 carrying the only phone-AND-standalone-exclusive CSS in the repo (bounded overflow:hidden shell, -webkit-overflow-scrolling:touch scrollport). PR #2046 added per-test devices["iPhone 14"] emulation in tests/ui-phone-motion.spec.ts as a cheap partial, but that is still Playwright WebKit, not the iOS engine. Next: decide between a dedicated mobile-WebKit project (CI cost) and per-test emulation as the standing pattern, and add standalone display-mode coverage for the phone shell rules. Relates to #280 (physical iPhone acceptance debt). | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | +| #NTAV3D | P2 | issue | Adversarial divergence (S5 harness pin): scope-guessed-chunk-id — review fallback returns a grounded source pointer echoing the query instead of refusing | Pinned in tests/rag-adversarial-harness.test.ts KNOWN_DIVERGENCES (self-expiring). Observed shape: grounded true, cited [syn-scope-guess-a], and the guessed (never-retrieved) chunk id syn-not-retrieved-zzz is never resolved into content — the no-read invariant holds — but the review fallback returns a grounded source pointer that echoes the query text rather than refusing the guessed-id request. Fixture: scripts/fixtures/rag-adversarial-cases.v1.json case scope-guessed-chunk-id (category scope_or_tenant). Next: decide whether a query naming an unretrieved chunk id should refuse rather than fall back to a source pointer (RAG-surface change; own PR; harness pin flips; canary pair). Stop: do not delete the pin without the behaviour change. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | +| #VXB8XA | P2 | issue | Adversarial divergence (S5 harness pin): cite-mismatched-attribution — offline document-match listing cites every retrieved document, not only the claim-bearing one | Pinned in tests/rag-adversarial-harness.test.ts KNOWN_DIVERGENCES (self-expiring: the harness asserts the normative B0 fixture contract still FAILS; when behaviour reaches the fixture expectation the pin goes red and must be deleted). Observed shape: cited chunk ids [syn-cite-attrib-a, syn-cite-attrib-b], grounded true, answerQualityTier source_only — the document-match listing attributes both retrieved documents although only one carries the claim. Fixture: scripts/fixtures/rag-adversarial-cases.v1.json case cite-mismatched-attribution (category citation_fabrication). Safety invariants (network, budget, canary absence, forbidden substrings, tenancy) hold; only citation attribution precision diverges. Next: decide whether the document-match listing should cite only claim-bearing documents (RAG-surface change; own PR; RAG impact line; offline harness proves the pin flips; canary pair). Stop: do not delete the pin without the behaviour change; do not weaken the fixture. | RAG programme coordinator, packet S5 (PR #2056) follow-ups, 2026-08-17 | 2026-08-17 | +| #S4K1GA | P3 | task | Physical iPhone acceptance owed for the answer-progress motion fix (Safari + installed PWA, Motion=Full) | PR #2046 fixed the reported defect (OS Reduce Motion froze every animation and set the ECG trace to opacity:0) and added a Motion preference whose "full" value opts back in over the OS setting. All executed browser evidence ran on Chromium 1194 in a Cloud container — the repo's own verify:ui gate could not run because check:playwright-browser-revision reports the known #255 drift (expects 1234). Playwright WebKit is not the iOS engine either. Acceptance: on the physical iPhone, in Safari and as the installed PWA, with Settings > Motion set to Full, confirm the ECG strip visibly travels and the current-step spinner rotates; with Motion left on System, confirm the trace stays visible and static rather than blank. Failure to confirm means the defect class is unclosed, which is exactly how #1974/#1989/#1995 were each declared fixed. Relates to #255, #280. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | +| #Q5JHBJ | P2 | task | Deploy the 20260818090000 schema_drift_snapshot v2 history probe (Phase 6.1) in an approved production window after Phase 4, triage the first unguarded no-statements report, and author fail-fast guard migrations for the pre-contract 2026-07-01..02 and 2026-07-12 rows | Repo side of remediation plan Phase 6 landed (probe migration, guard-migration contract in docs/database-drift-detection.md + AGENTS.md, tests/migration-history-guards.test.ts, tests/search-health-index-coverage.test.ts + supabase/search-health-unmonitored-indexes.json). The migration is NOT deployed: it needs the owner-approved production migration deploy window (plan approval map, Phase 6.1, after Phase 4). Expected first live run: the ~9 remaining section 1.1 versions plus the 2026-07-12 batch (20260712165915..20260712173000) are reported as unguarded no_statements findings because no repo-provable guard exists; each needs a validation guard migration (20260804110240 pattern) + a migration_history allowlist entry, not a bare allowlist. Also decide the 8 monitor-candidate indexes in the ratchet file by a required_indexes migration (Phase 4.4). Do not touch #316/#056 for this; those rows are owned by the Phase 1.2 / Phase 2 sessions. | Phase 6 worker chat 2026-08-18; docs/database-remediation-plan.md section 6; docs/audit/live-drift-forensics-2026-08.md Phase 6 | 2026-08-17 | +| #43SSS0 | P3 | rec | Three spring easing tokens in globals.css are dead: zero var() references and zero utility usage | --spring-tight, --spring-bouncy and --spring-gentle (src/app/globals.css:222-224) are declared in the @theme block but have no var() consumer in any stylesheet and no generated-utility consumer in src/. Tailwind v4.3.3 tree-shakes unused theme variables, so they never reach the compiled CSS — they are source noise, not shipped weight. Found while confirming (during PR #2046) that --animate-answer-ecg survives that same tree-shaking because it IS referenced via var() from the project's own CSS; --ease-spring is the working precedent for that pattern. Next: delete the three tokens, or wire them to the motion surfaces they were intended for. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | +| #75JA0P | P2 | issue | Playwright runs the whole suite with reducedMotion:"reduce", so no gate reflects the default user configuration | playwright.config.ts:61 sets contextOptions: { reducedMotion: "reduce" } suite-wide, and every motion assertion has to opt out per-test via page.emulateMedia({ reducedMotion: "no-preference" }). That inversion is why three consecutive PRs (#1974, #1989, #1995) shipped green while a physical iPhone with OS Reduce Motion on showed a frozen, blank answer-progress panel: the suite never exercised the reported configuration. PR #2046 added tests/ui-phone-motion.spec.ts to cover that one surface, but the suite-wide default remains inverted for every other motion behaviour. Next: decide whether the suite default should be no-preference with reduce opted into per-test (the safer direction), or keep the current default and add a contract test that fails when a motion assertion has no explicit emulateMedia call. | PR #2046 phone/PWA answer-progress animation defect, 2026-08-17 | 2026-08-17 | ## Resolved / archive @@ -481,3 +488,6 @@ Move resolved rows here with the resolution date and a one-line outcome. Keep th | #192 | task | X6: Raise clinical/retrieval/answer coverage floors | Resolved by PR #1964 (commit adc5182a7): X6 added four clinical/retrieval/answer coverage groups plus global/broad floors and CI enforcement, raising evidence/verification branches 81 to 84 and core RAG branches 72 to 76 without lowering any floor. Current-main static CI scope check passes. A fresh exclusive coverage run is useful after later RAG commits but is freshness evidence, not remaining X6 implementation. | 2026-08-15 | | #162 | task | Redesign Tools search results state (Compact Results Instrument) | Resolved on current main 512b8c20e by commit 88e3117ff3e078f0c88cd2e5e5c360ce4efca665: Tools is an all-results directory with query-as-H1, shared results band, dense rows, filters, and one-composer ownership. Stale local branch disposition remains separate hygiene work. | 2026-08-15 | | #238 | rec | Visual pass for Sheet portal default on settings, sidebar, and answer overlays | Completed on clean current main 512b8c20 with real Chromium journeys across all five required Sheet host contexts. Repository wrapper result: 5 passed in 23.8s, covering settings via ClinicalSidebar at phone widths with close/Escape, source-backed answer Sources/safety sheets with focus return, answer Sources/Clinical notes/Evidence sheets, the mobile launcher detail sheet, and the form section-navigation sheet. This is host-level browser evidence, not generic Sheet unit coverage. | 2026-08-15 | +| #0MSNT8 | task | Governance Option B decided: tag document-summary rows with similarity_origin 'document_context', keep the confidence label — implement as packet G1 | Implemented as packet G1 on branch claude/g1-rag-document-context-qn9ubx. Every element of the queued Option B scope landed: document_context added to the similarity_origin union in src/lib/types.ts and to src/lib/answer-stream-contract.ts (as an allow-set, so the union and the stream validator cannot drift), stamped in buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts), deriveConfidence left unchanged, rag.ts synthetic_similarity_count left counting only synthetic_text, and docs/clinical-hazard-analysis.md H5a updated to mark the decision implemented. Four discriminating pins, each mutation-checked against the change it exists to catch: summary rows carry the tag (tests/rag-retrieval-row-contract.test.ts); two document_context citations at >= 0.82 still yield high while the identical scores tagged synthetic_text still yield medium, and the telemetry counter ignores the new value (tests/rag-score.test.ts); the stream validator accepts every declared union member with a compile-time exhaustiveness guard (tests/answer-incremental-delivery.test.ts). Gates: vitest 58/58 across the five touched suites, full unit 642 files / 6879 passed, eval:rag:offline 36 golden cases / 603 tests, check:rag:fixtures, verify:pr-local all 18 stages green. No canary per the decision. HANDOVER status row updated. Closes the paired governance question row #J912J9. | 2026-08-17 | +| #6BG9X2 | task | R2 + R3: claim-support strictness rejects verbatim-faithful guideline restatements (directive normativity; topic-overlap dilution) — packet S1c | Fixed in PR #2052 (packet S1c): R2 descriptive-norm disjunct in normativeDirectiveActions (digit-anchored, descriptiveContext-guarded, adversarial negatives for care-record prose, unrelated imperatives, and incidental norm-adjacent action words) and R3 adjacent atom-free-segment topic lending confined to the overlap clause with all other gates single-segment. Measured offline before loosening: 87 -> 78 sole-overlap rejections (46 -> 42 unique), zero protective fixture flips across the 613-test offline corpus; discriminating negatives pin cross-bullet dose mis-binding, non-adjacent synthesis, and alien-topic claims. Post-merge canary pair owner-approved. Note: request file emitted via the repo inbox schema because issues:done rejects ULID display ids (issueRowFingerprint is numeric-only). | 2026-08-17 | +| #J912J9 | issue | Decide whether a fabricated similarity of 1 on document-summary rows may earn the high confidence label a clinician reads | Answered and implemented (packet G1, PR for branch claude/g1-rag-document-context-qn9ubx). Owner decided Option B on 2026-08-17: document-summary rows keep the high confidence label, and the fabricated similarity gets its own provenance value rather than being folded into synthetic_text. Landed: document_context added to the similarity_origin union (src/lib/types.ts) and to the streamed-preview client-source validator (src/lib/answer-stream-contract.ts), and stamped in buildDocumentSummaryResults (src/lib/rag/rag-row-contracts.ts). Per the decision deriveConfidence (src/lib/rag/rag-answer-support.ts) is unchanged and still excludes only synthetic_text, and rag.ts synthetic_similarity_count still counts only synthetic_text; both are pinned by discriminating tests that go red on the rejected Option A fold. Option A (tag as synthetic_text so summaries cap at medium) recorded as rejected. docs/clinical-hazard-analysis.md H5a marks the decision implemented and names the residual: the tag closes the legibility gap, not the deeper question of whether a constant 1.0 should contribute to a confidence label, but any future gate can now discriminate the route without re-deriving provenance. No retrieval behaviour change; no canary. | 2026-08-17 | diff --git a/docs/rag-improvement/COORDINATION.md b/docs/rag-improvement/COORDINATION.md index cd5edba948..10339639f7 100644 --- a/docs/rag-improvement/COORDINATION.md +++ b/docs/rag-improvement/COORDINATION.md @@ -181,14 +181,14 @@ Always a fresh owner ask, every single time: regressions, answer gate 44/44. The first S1b post run (32038751592) was red on one non-golden case; root-caused to the finalizer gap-recovery hole → packet S1d, not S1b. - **Owner decisions:** R1 before S2; governance Option B; S1d lands before S2. -- **Wave 1 status:** S1c (#2052 + follow-ups #2063/#2065) merged, canary pair green - (32049952885 → 32052479537); S1d (#2054, merge `0bbd64fbc`) merged, **canary pending owner - approval**; S6 (#2057) merged. **Still to dispatch: G1** (`types.ts`, - `answer-stream-contract.ts`, `rag-row-contracts.ts`; no canary). S5 follow-ups queued as - inbox requests: three adversarial divergence pins (`cite-mismatched-attribution`, - `scope-other-owner-document`, `scope-guessed-chunk-id`) and the owner decision on enabling - `RAG_TELEMETRY_EXTENDED` in production. -- **Wave 2:** S2 (+S2b) once the S1d canary pair is green (S1c pair already green). **Wave 3:** S3; S7+ owner decisions. +- **Wave 1 status (2026-08-18):** all packets merged — S1c (#2052; follow-up #2063 kept), S1d + (#2054, merge `0bbd64fbc`), G1 (#2053, merge `125e98526`), S6 (#2057). Post-S1d canary run + 32097916649 (`9904fbda8`) was **red** on `agitation-im-po-route-short-terms` (extractive path, + deterministic, 5 → 0 citations); live bisect placed it on the S1c follow-up **#2065** + (condition-first `for/in` regex in `rag-claim-support.ts`), not S1d/G1. **Revert PR #2088** is + open (probe restores 5 citations; offline 614/614). A confirmation canary follows its merge and + becomes the new baseline. Reconcile D4 applied 17 requests (G1/S1c/governance rows closed). +- **Wave 2:** S2 (+S2b) once the post-#2088 confirmation canary is green. **Wave 3:** S3; S7+ owner decisions. - **Waiting on owner:** merges as PRs open; canary approvals for S1c, S1d, S2. - **Live board (artifact, owner-private):** RAG Master Plan v2 — `https://claude.ai/code/artifact/d5dba709-0df3-40e3-8a45-15997231533d`. diff --git a/docs/rag-improvement/HANDOVER.md b/docs/rag-improvement/HANDOVER.md index d01fd1c33b..ab01642215 100644 --- a/docs/rag-improvement/HANDOVER.md +++ b/docs/rag-improvement/HANDOVER.md @@ -67,25 +67,25 @@ generation-quality verdict on fallback`), merged 2026-08-13 — structured ## 2. Status table — update in every programme PR -| Packet | Scope | Branch | PR | State | Canary / evidence refs | -| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | --------------------- | ------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Guide | Programme guide | `claude/rag-plan-review-guide-vhrls9` | #1895 | Merged 2026-08-13 | docs-only | -| Handover | Multi-session handover + coordination | `claude/rag-plan-review-guide-vhrls9` | #1908 / #2024 | Merged 2026-08-13; coordination layer PR #2024 | docs-only | -| S0 | A1 phase 1: structured fallback diagnostics | `claude/lithium-generation-quality-debug-ji1vce` | #1899 | Merged 2026-08-13 | offline 93/93 focused | -| S1 | A1 phase 2: rung-1 verification-faithfulness fixes | `claude/s1-rag-mitigation-231-86c182` | #2022 | Merged 2026-08-17 (squash `2bd146eed`, landed by content) | 8 pre-fix + 5 post-fix live probes 2026-08-17; offline 583/583; canary pair run 31964560921 (baseline `8f8d111ab`) -> run 32025082010 (`2bd146eed`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 (denominator reconciled by S5; see baseline-record §3); rung-2 measurement in `docs/audit/live-drift-forensics-2026-08.md` §5 | -| S1b | A1 rung 3 (R1): pre-deadline strong routing for dosing class | `claude/s1b-rag-dosing-routing-6u1mik` | #2035 | Merged 2026-08-17 (PR #2035, merge `92f7618`) | canary pair pending: baseline run 32025082010 (`2bd146eed`) -> post-merge dispatch (owner-approved); offline 586/586 + verify:pr-local heavy scope green | -| S1c | A1 residuals R2 + R3: claim-support strictness | `claude/s1c-residuals-r2-r3-4pb1at` | #2052 | Merged 2026-08-17 (merge `b8e774bcd`; follow-ups #2063, #2065) | canary pair: baseline run 32049952885 -> post run 32052479537 (`084f63799`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 | -| S1d | A1 final-gate gap recovery: hedged cited low-confidence fast answers must recover extractively, not collapse to a citation-free `provider_source_gap` | `claude/s1d-final-gate-gap-recovery-dxgrn2` | #2054 | Merged 2026-08-17 (merge `0bbd64fbc`); canary pending (owner approval) | needs post-merge canary pair (baseline run 32039841070, `e6ad0d5db`); offline: verify:pr-local heavy green, eval:rag:offline 604/604, 6 new discriminating fixtures | -| G1 | Governance: provenance tag for document-summary rows (Option B) | `claude/g1-rag-document-context-qn9ubx` | #2053 | Implemented 2026-08-17 - awaiting review/merge | no canary by decision (provenance tag only). Offline: 58/58 across the five touched suites; `eval:rag:offline`, `check:rag:fixtures`, `verify:pr-local` - see PR body | -| S2 | A2 (+A3): composition menu + moderate length | `claude/rag-a2-composition-` | — | Blocked on the S1d canary pair (S1c pair green); dispatch when green | canary pair + `eval:answer-quality` + Gate E | -| S2b | A3: moderate length (if separate review needed) | `claude/rag-a3-length-` | — | Blocked on S2 | — | -| S3 | A4: follow-up suggestion refinement | `claude/rag-a4-follow-ups-` | — | Blocked on S2 + S2b | — | -| S4 | B0: adversarial fixtures + baseline + register | `claude/packet-s4-adversarial-fixtures-5ho5tp` | #2036 | Merged 2026-08-17 (squash `f5b093291`) | Offline only: `check:rag:adversarial-fixtures` 24 cases / 8 categories / 6 canaries; `eval:rag:offline` 24 suites, 597 tests. Baseline `scripts/fixtures/rag-adversarial-baseline.v1.json` marks the three provider-backed gates `pending_owner_run` | -| S5 | B1+B2: telemetry assessment + offline harness | `claude/s5-rag-telemetry-harness-2wvis7` | #2056 | Merged 2026-08-17 (merge `093f9340c`); post-merge canary run 32049952885 | Offline only: `eval:rag:adversarial:offline` 25/25 (24 cases + canary-free report; 3 divergences pinned in `KNOWN_DIVERGENCES`); B1 gap = `verification_latency_ms` behind `RAG_TELEMETRY_EXTENDED` (default false); canary-absence tests green; 44/44 denominator reconciled | -| S6 | B3: Docling lab benchmark | `claude/packet-s6-docling-lab-d6foa6` | #2057 | Merged 2026-08-17 (merge `5a6418636`) | Offline only: `check:docling-lab` 36 fixtures / 10 hostile / 6 canaries + Gate B template valid; `verify:pr-local` heavy plan failed:(none); contract test 20/20; legacy smoke 46 docs, 10/10 hostile contained, canary-clean report. Verdict is a separate owner dispatch of `docling-lab.yml` | -| S7+ | B4 shadow / B5 Ragas / B6 reranker / B7 DSPy | — | — | Gated — owner decision | — | -| #212 T1–T3 | Runtime row contracts (rag.ts, rag-candidate-sources.ts, src/app/api) — sibling stream sharing `src/lib/rag/**` | — | #1946 / #1981 / #2023 | Merged (T3 squash `440a34f71` 2026-08-17) | see the #212 ledger row; RAG surface complete for the cast class | -| #212 T4 | Runtime row contracts: `worker/main.ts` (11 casts) — sibling stream | `claude/ledger-212-tranche-4-worker-q3y6i4` | #2037 | Merged 2026-08-17 (squash `1726537b7`); #212 closed by reconcile PR #2045 | Governance Preflight complete; audit: 1 inbound cast (claim rows, per-row fail-soft) + 2 read-back param casts contracted, 9 outbound/interop left; closes #212 (inbox `done` queued in the PR) | +| Packet | Scope | Branch | PR | State | Canary / evidence refs | +| ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | --------------------- | ------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Guide | Programme guide | `claude/rag-plan-review-guide-vhrls9` | #1895 | Merged 2026-08-13 | docs-only | +| Handover | Multi-session handover + coordination | `claude/rag-plan-review-guide-vhrls9` | #1908 / #2024 | Merged 2026-08-13; coordination layer PR #2024 | docs-only | +| S0 | A1 phase 1: structured fallback diagnostics | `claude/lithium-generation-quality-debug-ji1vce` | #1899 | Merged 2026-08-13 | offline 93/93 focused | +| S1 | A1 phase 2: rung-1 verification-faithfulness fixes | `claude/s1-rag-mitigation-231-86c182` | #2022 | Merged 2026-08-17 (squash `2bd146eed`, landed by content) | 8 pre-fix + 5 post-fix live probes 2026-08-17; offline 583/583; canary pair run 31964560921 (baseline `8f8d111ab`) -> run 32025082010 (`2bd146eed`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 (denominator reconciled by S5; see baseline-record §3); rung-2 measurement in `docs/audit/live-drift-forensics-2026-08.md` §5 | +| S1b | A1 rung 3 (R1): pre-deadline strong routing for dosing class | `claude/s1b-rag-dosing-routing-6u1mik` | #2035 | Merged 2026-08-17 (PR #2035, merge `92f7618`) | canary pair pending: baseline run 32025082010 (`2bd146eed`) -> post-merge dispatch (owner-approved); offline 586/586 + verify:pr-local heavy scope green | +| S1c | A1 residuals R2 + R3: claim-support strictness | `claude/s1c-residuals-r2-r3-4pb1at` | #2052 | Merged 2026-08-17 (merge `b8e774bcd`; follow-up #2063 kept; follow-up #2065 reverted by PR #2088 after canary regression) | canary pair: baseline run 32049952885 -> post run 32052479537 (`084f63799`): recall 1.0/1.0, zero per-case rr regressions, answer gate 44/44 | +| S1d | A1 final-gate gap recovery: hedged cited low-confidence fast answers must recover extractively, not collapse to a citation-free `provider_source_gap` | `claude/s1d-final-gate-gap-recovery-dxgrn2` | #2054 | Merged 2026-08-17 (merge `0bbd64fbc`); landed by content; canary pair pending the #2065 revert (PR #2088) | post run 32097916649 (`9904fbda8`) RED on `agitation-im-po-route-short-terms` — bisected live to PR #2065 (S1c follow-up condition-first regex), not S1d; revert PR #2088; confirmation run owed after it merges | +| G1 | Governance: provenance tag for document-summary rows (Option B) | `claude/g1-rag-document-context-qn9ubx` | #2053 | Merged 2026-08-17 (merge `125e98526`); rows #J912J9 / #0MSNT8 closed at reconcile | no canary (no behaviour change); `document_context` tag at `types.ts` + `rag-row-contracts.ts`, deriveConfidence pinned | +| S2 | A2 (+A3): composition menu + moderate length | `claude/rag-a2-composition-` | — | Blocked on the confirmation canary after PR #2088 (revert of #2065) merges; then dispatch | canary pair + `eval:answer-quality` + Gate E | +| S2b | A3: moderate length (if separate review needed) | `claude/rag-a3-length-` | — | Blocked on S2 | — | +| S3 | A4: follow-up suggestion refinement | `claude/rag-a4-follow-ups-` | — | Blocked on S2 + S2b | — | +| S4 | B0: adversarial fixtures + baseline + register | `claude/packet-s4-adversarial-fixtures-5ho5tp` | #2036 | Merged 2026-08-17 (squash `f5b093291`) | Offline only: `check:rag:adversarial-fixtures` 24 cases / 8 categories / 6 canaries; `eval:rag:offline` 24 suites, 597 tests. Baseline `scripts/fixtures/rag-adversarial-baseline.v1.json` marks the three provider-backed gates `pending_owner_run` | +| S5 | B1+B2: telemetry assessment + offline harness | `claude/s5-rag-telemetry-harness-2wvis7` | #2056 | Merged 2026-08-17 (merge `093f9340c`); post-merge canary run 32049952885 | Offline only: `eval:rag:adversarial:offline` 25/25 (24 cases + canary-free report; 3 divergences pinned in `KNOWN_DIVERGENCES`); B1 gap = `verification_latency_ms` behind `RAG_TELEMETRY_EXTENDED` (default false); canary-absence tests green; 44/44 denominator reconciled | +| S6 | B3: Docling lab benchmark | `claude/packet-s6-docling-lab-d6foa6` | #2057 | Merged 2026-08-17 (merge `5a6418636`) | Offline only: `check:docling-lab` 36 fixtures / 10 hostile / 6 canaries + Gate B template valid; `verify:pr-local` heavy plan failed:(none); contract test 20/20; legacy smoke 46 docs, 10/10 hostile contained, canary-clean report. Verdict is a separate owner dispatch of `docling-lab.yml` | +| S7+ | B4 shadow / B5 Ragas / B6 reranker / B7 DSPy | — | — | Gated — owner decision | — | +| #212 T1–T3 | Runtime row contracts (rag.ts, rag-candidate-sources.ts, src/app/api) — sibling stream sharing `src/lib/rag/**` | — | #1946 / #1981 / #2023 | Merged (T3 squash `440a34f71` 2026-08-17) | see the #212 ledger row; RAG surface complete for the cast class | +| #212 T4 | Runtime row contracts: `worker/main.ts` (11 casts) — sibling stream | `claude/ledger-212-tranche-4-worker-q3y6i4` | #2037 | Merged 2026-08-17 (squash `1726537b7`); #212 closed by reconcile PR #2045 | Governance Preflight complete; audit: 1 inbound cast (claim rows, per-row fail-soft) + 2 read-back param casts contracted, 9 outbound/interop left; closes #212 (inbox `done` queued in the PR) | Update rule: the session that opens a packet's PR edits its row (branch, PR number, state) in the same PR. A later session updating another packet may also correct stale rows From 1a2e24f2dfedcbd449333d7d7dfe7c222f0edb2a Mon Sep 17 00:00:00 2001 From: BigSimmo <87357024+BigSimmo@users.noreply.github.com> Date: Tue, 18 Aug 2026 12:30:52 +0800 Subject: [PATCH 2/2] docs(ledger): record the D4 reconcile review Co-Authored-By: Claude Fable 5 --- ...a0a1890452f25c18d773c8c8eb05ca8de516da7f1e099e69b88.record.md | 1 + 1 file changed, 1 insertion(+) create mode 100644 docs/branch-review-records/b6d9c18765ffda0a1890452f25c18d773c8c8eb05ca8de516da7f1e099e69b88.record.md diff --git a/docs/branch-review-records/b6d9c18765ffda0a1890452f25c18d773c8c8eb05ca8de516da7f1e099e69b88.record.md b/docs/branch-review-records/b6d9c18765ffda0a1890452f25c18d773c8c8eb05ca8de516da7f1e099e69b88.record.md new file mode 100644 index 0000000000..dca25ccd2e --- /dev/null +++ b/docs/branch-review-records/b6d9c18765ffda0a1890452f25c18d773c8c8eb05ca8de516da7f1e099e69b88.record.md @@ -0,0 +1 @@ +| 2026-08-18 | claude/rag-d4-reconcile-inbox | d1d75055396cff3ea15c6ef6c45d9363440b4175 | issues:reconcile (17 requests) + HANDOVER G1/S1d/S1c/S2 rows + COORDINATION §7 (docs-only) | single fresh-base reconcile; 0 pending afterwards; G1/S1c/governance rows closed | check:outstanding-issues 0 pending/234 applied; check:ledger-write-discipline passed 981d85de0301..HEAD; verify:pr-local docs scope |