Skip to content

fix(discovery): state a research run that accepts no claim - #6147

Merged
cursor[bot] merged 2 commits into
mainfrom
cursor/research-gate-zero-accepted-e44b
Oct 4, 2026
Merged

cursor[bot] merged 2 commits into
mainfrom
cursor/research-gate-zero-accepted-e44b

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Outcome-gate rows 7 and 12 of /discovery:research are quantified over accepted claims, so an artifact that lists every claim as a Gap (or as an unresolved Conflict) passed them with zero accepted claims and read like an answer. Row 4 said "every claim", which contradicted the Gap route.

Fix

Option (b) from the issue:

  • Row 4 is now quantified over accepted claims.
  • New verifier-owned row 14: the index frontmatter's accepted: counts the claims neither listed under Gaps nor left unresolved in Conflicts (a Conflicts claim counts only when its entry resolves in the claim's favor; a refuted finding or a criterion-12 failure filed there does not, matching discipline.md). At accepted: 0 the Summary opens with Inconclusive: no claim accepted. A zero with that line passes; a missing or wrong count, or a zero without the line, fails. The verifier owns it because the issue's acceptance names research-verifier.md as the grader and the final accepted set is only known after rows 4, 7 and 12 are graded.
  • On the N-topic path the synthesized root index goes to the verifier for every verifier-owned row (4, 7, 12, 14), not criterion 12 alone, in dispatch.md, research/SKILL.md and research-deep/SKILL.md. The verifier may read the sub-slice indexes a synthesized root names, and grades only the root.
  • artifact-shape.md defines accepted: (an integer on every run; the inconclusive line opens the Summary). The researcher sets it in its final write. Every list of the verifier rows (dispatch, gotchas, parent contract, research-deep, evals) now names 14.
  • Three new eval cases cover an artifact whose every claim is a Gap, one whose every claim is an unresolved Conflicts entry, and a count that includes a Gap claim. research-deep eval 1 expects the four-row synthesis check.
  • SKILL.md's Effort sentence still sits inside the first 20,000 bytes that are re-attached after compaction (it ends at byte 19,992).
  • discovery 0.28.4.

Verification

  • contract.test.sh run against origin/main's discovery plugin: 19 new or updated pins fail; with the review-fix pins on this PR's previous head, 13 fail. On this head: all contract assertions pass.
  • Eight of nine discovery *.test.sh suites exit 0; build-view.test.sh fails only its three headless-browser checks, which fail identically on origin/main in this environment. scripts/validate-plugin-contracts.test.sh PASS=173 FAIL=0.
  • check-changelog-parity.sh --check, --check-order, --check-bump origin/main and --check-preserved origin/main pass.
  • shellcheck clean; typos clean on the changed files; markdownlint-cli2 0 issues on the changed markdown.

Related

Closes #5833
Related: #5805, #5827

Open in Web Open in Cursor 

Outcome-gate row 4 is quantified over accepted claims, reconciling it with
the Gap route rows 7 and 12 already allow. New verifier-owned row 14 grades
the index's accepted: count and, at zero, an 'Inconclusive: no claim
accepted.' Summary line, so a run that settles nothing passes as
inconclusive instead of reading like an answer.

Co-authored-by: ksextonmelodic <ksextonmelodic@gmail.com>
@kyle-sexton
kyle-sexton marked this pull request as ready for review October 4, 2026 02:07
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-04T02:10:25.254053Z 6f0ecd6 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 51s —— View job


Security review complete ✅

Task: Run /review:security-review (per custom instructions) against PR #6147 at HEAD 6f0ecd6.

Scope

All 12 changed files are plugin documentation/spec and test assets for the discovery plugin's /discovery:research skill: agent prompt bodies (agents/research-verifier.md, agents/researcher.md), skill/context markdown (SKILL.md ×2, artifact-shape.md, dispatch.md, gotchas.md, reference/parent-contract.md), a bash contract-test script, an eval fixtures JSON, the plugin manifest, and the changelog. No application code, no network/file I/O, no shell execution logic, and no secret/credential handling changed.

Findings

None. This PR only tightens the outcome gate:

  • Row 4 is narrowed from "every claim" to "every accepted claim," closing a vacuous-pass loophole rather than opening one.
  • New row 14 adds a verifier-owned check (accepted: count vs. actual accepted claims, with a required Inconclusive: no claim accepted. line at zero) — a net-new control, not a removed one.
  • scripts/contract.test.sh changes are additional assert_present/assert_absent regex pins against static file content — no shell injection surface (no untrusted input reaches these assertions; patterns are hardcoded).
  • evals/evals.json additions are fixture prompts/expectations for the eval harness, not executable logic.

I checked this diff under the instruction-surface deletions lens (docs/conventions/instruction-exception-register/README.md is present in this checkout and was consulted): nothing here deletes, narrows, or softens a standing guardrail. Every change either adds a new gate criterion (row 14) or re-scopes an existing one to be more precise about what it grades — there's no case where a previously-enforced check becomes unenforced.

No GitHub Actions files were touched, so zizmor's advisory lane is not implicated.

@claude

claude Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 4m 12s —— View job


Code review complete

  • Gather context (PR diff, changed files, CLAUDE.md/REVIEW.md)
  • Run skip-gate checks (PR open, not draft, meaningful change → full review)
  • Review changed files for correctness/maintainability
  • Post findings

Scope: all 12 changed files under plugins/discovery/ (agents, skills, reference, scripts, evals), diffed against origin/main. Security findings are out of scope for this lane per REVIEW.md (claude-security-review.yml exists in this repo).

Findings: none that clear the high-signal bar.

What I checked and verified directly (reading the files, not just the PR's claims):

  • The core fix is coherent: row 4 in skills/research/SKILL.md now reads "every accepted claim," matching rows 7/12, resolving the contradiction with the Gap route the issue described.
  • New row 14's text, Owner (verifier), and FAIL action are consistent between skills/research/SKILL.md, agents/research-verifier.md's dispatch brief and "Return exactly this" block, agents/researcher.md's withheld-criteria list and verification_request.criterion string, reference/parent-contract.md, skills/research/context/dispatch.md, and skills/research/context/gotchas.md.
  • skills/research/context/artifact-shape.md's new accepted: definition ("claims the evidence table presents that the Gaps section does not list") lines up with how Conflicts-flagged claims (still in the evidence table) vs. Gap claims (excluded) are described elsewhere in SKILL.md.
  • Ran a literal grep across plugins/discovery/**/*.{md,json} for stale "rows 4, 7 and 12" / "criteria 4, 7 or 12" style references the new contract assertion (contract.test.sh section 22, assert_absent 'no stale three-row verifier count') is meant to catch — no hits outside the intentionally-scoped mentions (gotchas.md's "Criteria 4, 7 and 12 ask the run to judge..." bullet, which is still true of those three specifically and is followed by a separate clause adding criterion 14; CHANGELOG.md, excluded by the test's own design since it's a historical record).
  • Walked through both new eval cases (every-claim-a-gap-is-stated-inconclusive, accepted-count-excludes-gap-claims) against the updated gate/verifier text; the expected outcomes match row 14's pass/fail conditions as specified.
  • Confirmed the version bump (0.28.3 → 0.28.4) and CHANGELOG.md entry are present and dated, and the new contract assertions (contract.test.sh section 22) target the files/strings actually changed.

Not independently re-run here (no shell execution available in this review environment): contract.test.sh, the other *.test.sh suites, check-changelog-parity.sh, shellcheck, markdownlint. Those results are author-claimed in the PR description and unverified by me.

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

claude-security-review has reviewed this pull request through 6f0ecd6; a later push is reviewed from there.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6f0ecd6881

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/discovery/skills/research/context/artifact-shape.md Outdated
Comment thread plugins/discovery/skills/research-deep/SKILL.md
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

claude-review has reviewed this pull request through 6f0ecd6; a later push is reviewed from there.

…n row 14 (#5833)

accepted: excluded only Gap claims, so a run that rejected every claim
through an unresolved Conflicts entry (row 12's Gap-or-Conflicts route, a
refuted engine finding) kept a nonzero count and never wrote the
inconclusive line. The count now leaves out claims in Gaps and claims left
unresolved in Conflicts, in the gate row, artifact-shape, the verifier and
the researcher, with an eval case for an all-Conflicts run.

The N-topic synthesis went to its verifier for criterion 12 only, so row 14
never ran on a synthesized root index. The synthesis brief in dispatch.md,
research/SKILL.md and research-deep/SKILL.md now names every verifier-owned
row, and the verifier reads the sub-slice indexes a synthesized root names.

Co-authored-by: ksextonmelodic <ksextonmelodic@gmail.com>
@cursor
cursor Bot added this pull request to the merge queue Oct 4, 2026
Merged via the queue into main with commit dde43c8 Oct 4, 2026
38 of 39 checks passed
@cursor
cursor Bot deleted the cursor/research-gate-zero-accepted-e44b branch October 4, 2026 02:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

discovery research gate passes an artifact with zero accepted claims without saying so

2 participants