fix(discovery): state a research run that accepts no claim - #6147
Conversation
Outcome-gate row 4 is quantified over accepted claims, reconciling it with the Gap route rows 7 and 12 already allow. New verifier-owned row 14 grades the index's accepted: count and, at zero, an 'Inconclusive: no claim accepted.' Summary line, so a run that settles nothing passes as inconclusive instead of reading like an answer. Co-authored-by: ksextonmelodic <ksextonmelodic@gmail.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Claude finished @kyle-sexton's task in 51s —— View job Security review complete ✅Task: Run
ScopeAll 12 changed files are plugin documentation/spec and test assets for the FindingsNone. This PR only tightens the outcome gate:
I checked this diff under the instruction-surface deletions lens ( No GitHub Actions files were touched, so zizmor's advisory lane is not implicated. |
|
Claude finished @kyle-sexton's task in 4m 12s —— View job Code review complete
Scope: all 12 changed files under Findings: none that clear the high-signal bar. What I checked and verified directly (reading the files, not just the PR's claims):
Not independently re-run here (no shell execution available in this review environment): |
|
claude-security-review has reviewed this pull request through 6f0ecd6; a later push is reviewed from there. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6f0ecd6881
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
claude-review has reviewed this pull request through 6f0ecd6; a later push is reviewed from there. |
…n row 14 (#5833) accepted: excluded only Gap claims, so a run that rejected every claim through an unresolved Conflicts entry (row 12's Gap-or-Conflicts route, a refuted engine finding) kept a nonzero count and never wrote the inconclusive line. The count now leaves out claims in Gaps and claims left unresolved in Conflicts, in the gate row, artifact-shape, the verifier and the researcher, with an eval case for an all-Conflicts run. The N-topic synthesis went to its verifier for criterion 12 only, so row 14 never ran on a synthesized root index. The synthesis brief in dispatch.md, research/SKILL.md and research-deep/SKILL.md now names every verifier-owned row, and the verifier reads the sub-slice indexes a synthesized root names. Co-authored-by: ksextonmelodic <ksextonmelodic@gmail.com>
Summary
Outcome-gate rows 7 and 12 of
/discovery:researchare quantified over accepted claims, so an artifact that lists every claim as a Gap (or as an unresolved Conflict) passed them with zero accepted claims and read like an answer. Row 4 said "every claim", which contradicted the Gap route.Fix
Option (b) from the issue:
accepted:counts the claims neither listed under Gaps nor left unresolved in Conflicts (a Conflicts claim counts only when its entry resolves in the claim's favor; a refuted finding or a criterion-12 failure filed there does not, matchingdiscipline.md). Ataccepted: 0the Summary opens withInconclusive: no claim accepted.A zero with that line passes; a missing or wrong count, or a zero without the line, fails. The verifier owns it because the issue's acceptance namesresearch-verifier.mdas the grader and the final accepted set is only known after rows 4, 7 and 12 are graded.dispatch.md,research/SKILL.mdandresearch-deep/SKILL.md. The verifier may read the sub-slice indexes a synthesized root names, and grades only the root.artifact-shape.mddefinesaccepted:(an integer on every run; the inconclusive line opens the Summary). The researcher sets it in its final write. Every list of the verifier rows (dispatch, gotchas, parent contract, research-deep, evals) now names 14.research-deepeval 1 expects the four-row synthesis check.SKILL.md's Effort sentence still sits inside the first 20,000 bytes that are re-attached after compaction (it ends at byte 19,992).Verification
contract.test.shrun againstorigin/main's discovery plugin: 19 new or updated pins fail; with the review-fix pins on this PR's previous head, 13 fail. On this head: all contract assertions pass.*.test.shsuites exit 0;build-view.test.shfails only its three headless-browser checks, which fail identically on origin/main in this environment.scripts/validate-plugin-contracts.test.shPASS=173 FAIL=0.check-changelog-parity.sh--check,--check-order,--check-bump origin/mainand--check-preserved origin/mainpass.typosclean on the changed files; markdownlint-cli2 0 issues on the changed markdown.Related
Closes #5833
Related: #5805, #5827