Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion plugins/discovery/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "discovery",
"version": "0.28.3",
"version": "0.28.4",
"description": "Discovery before changes: explore the local codebase, run multi-source external research, and reconstruct why a past decision was made from evidence outside the code. Each dispatches a subagent by default so the reading stays out of the main thread, with source tiers, falsification, recency gates and a coverage ledger, and persists EXPLORE.md / RESEARCH.md / INTENT.md handoff artifacts. A research sweep workflow (/discovery:research-sweep) backs deep research with adversarial claim verification.",
"author": {
"name": "Melodic Software",
Expand Down
17 changes: 17 additions & 0 deletions plugins/discovery/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,22 @@
# Changelog: discovery plugin

## [0.28.4] - 2026-10-04

### Fixed

- **A research run that accepts no claim now says so (#5833).** Outcome-gate rows 7 and 12 apply to
accepted claims only, so an artifact that listed every claim as a Gap passed them and read like an
answer. Row 4 now also applies to accepted claims only, so it no longer contradicts the Gap route.
New verifier-owned row 14 checks that the index's `accepted:` count matches the accepted claims.
The count leaves out claims listed under Gaps and claims left unresolved in Conflicts, so a run
that rejects every claim through Conflicts also counts zero. At `accepted: 0`, the Summary must
open with `Inconclusive: no claim accepted.`; with that line, a run that settles nothing passes
as inconclusive. The verifier now grades rows 4, 7, 12 and 14, and the researcher sets
`accepted:` in its final write. On the N-topic path the synthesized root index now goes to the
verifier for all four rows, not criterion 12 alone, so row 14 also runs on it. Three eval cases
cover an artifact whose every claim is a Gap, one whose every claim is an unresolved Conflicts
entry, and a count that includes a Gap claim.

## [0.28.3] - 2026-10-03

### Changed
Expand Down
16 changes: 14 additions & 2 deletions plugins/discovery/agents/research-verifier.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,8 +15,10 @@ history, and everything you need arrives in your dispatch prompt or sits on disk

- **Target**: the `RESEARCH.md` path the parent's acceptance gate printed as `index=`. Grade that
file and the sidecars and fetch log beside it, nothing else. A research index you find anywhere
else is some other run's artifact.
- **Rows**: the outcome-gate row numbers to grade, currently 4, 7 and 12. The row text lives in the
else is some other run's artifact. A synthesized slice-root index is the one exception: the
sub-slice indexes it synthesizes, inside the same slice, are the record its carried claims and
qualifiers came from, so read those too and grade only the root.
- **Rows**: the outcome-gate row numbers to grade, currently 4, 7, 12 and 14. The row text lives in the
outcome gate table of
[`${CLAUDE_PLUGIN_ROOT}/skills/research/SKILL.md`](${CLAUDE_PLUGIN_ROOT}/skills/research/SKILL.md);
Read that table and grade each named row as it is written there. Do not grade from a paraphrase in
Expand All @@ -40,6 +42,15 @@ MEDIUM or LOW and listed in the Gaps section is not an accepted claim, so it doe
A quote found at its link settles only that the quote exists; it does not show the claim follows
from it, which is the question row 12 asks.

Rows 4, 7 and 12 hold vacuously when no claim is accepted, so row 14 is what grades that case.
Count the claims that stand accepted once you have graded them, each one neither listed under Gaps
nor left unresolved in Conflicts, and compare the count with the index frontmatter's `accepted:`.
A claim recorded under Conflicts in place of acceptance, such as a row 12 failure filed there, is
not accepted; one whose Conflicts entry resolves in its favor is. A missing field or a different
number fails row 14. At zero, row 14 passes only when the Summary opens with
`Inconclusive: no claim accepted.` and names the Gaps or Conflicts that blocked one; a
zero-accepted artifact that reads as an answer fails it.

A claim at `HIGH (single source)` has no corroborator to count, so row 4 turns on its
`single_source:` reason. Judge that reason against the definition in
[`${CLAUDE_PLUGIN_ROOT}/skills/research/context/discipline.md`](${CLAUDE_PLUGIN_ROOT}/skills/research/context/discipline.md),
Expand Down Expand Up @@ -83,6 +94,7 @@ rows:
"4": pass # pass | fail: <claim id and one line>
"7": pass
"12": pass
"14": pass
verification_line: "verification: pass (research-verifier, <YYYY-MM-DD>)"
open_questions: []
```
Expand Down
18 changes: 11 additions & 7 deletions plugins/discovery/agents/researcher.md
Original file line number Diff line number Diff line change
Expand Up @@ -255,23 +255,27 @@ Write the artifact in stages:
2. Write each `RESEARCH-<section>.md` sidecar as its section settles, and update its row in the
index.
3. The final write, after the outcome gate below, replaces the marker line with
`Run status: complete`. Nothing earlier does. The parent's gate refuses an index still carrying
the marker, which is how a stop at the limit reaches the parent even when no payload does.
`Run status: complete` and sets the frontmatter's `accepted:` count, which counts the
claims in neither Gaps nor an unresolved Conflicts entry. Nothing earlier does. The parent's gate
refuses an index still carrying the marker, which is how a stop at the limit reaches the parent
even when no payload does.

A by-value `RESEARCH.md` body carries `Run status: complete`, because by-value means the work
finished; the parent writes it and grades it like any other.

## The outcome gate is split: you do not grade all of it

Run the skill's outcome gate against your own artifacts before the final write. Three criteria are
Run the skill's outcome gate against your own artifacts before the final write. Four criteria are
**not yours to render a verdict on**, because grading them means judging the quality of your own
choices, and you are the context that made them:

- the criterion requiring ≥2 **independent** corroborators per claim (a floor below criterion 7),
- the criterion requiring ≥2 **independent** corroborators per accepted claim (a floor below criterion 7),
or a `single source` reason that holds for a first-party content claim,
- the criterion requiring every accepted claim to be HIGH confidence, `HIGH (single source)`
included, and
- the criterion requiring every accepted claim to follow jointly from its cited sources.
included,
- the criterion requiring every accepted claim to follow jointly from its cited sources, and
- the criterion requiring the index's `accepted:` count to match the accepted claims, with a
zero stated as inconclusive, because the accepted set is final only once the other three hold.

The gate's Owner column is the authority; where this list and that column differ, the column wins.
Assemble the evidence those criteria need, since per-claim source URLs with their tier, publishing
Expand Down Expand Up @@ -302,7 +306,7 @@ applicability: pass # pass | fail, mirrors check-source-applicability.py
verification: pending # never anything else; you render no verdict on your own confidence
verification_request:
target: <the same path as artifact: above>
criterion: "independent corroboration with every single-publisher claim labeled and not accepted, HIGH confidence, and joint-inference validity per accepted claim"
criterion: "independent corroboration with every single-publisher claim labeled and not accepted, HIGH confidence, and joint-inference validity per accepted claim, and an accepted-claim count that states a zero"
worker: fresh-context subagent
gate_owed: "the full post-dispatch acceptance gate, not only check-dispatch-artifact.sh, check-coverage-complete.sh and check-source-applicability.py: it also owes the discovery:research-verifier dispatch and project fit. Source: the discovery plugin's skills/research/SKILL.md 'Post-dispatch acceptance gate' and reference/parent-contract.md 'Running the acceptance gate'"
open_questions:
Expand Down
2 changes: 1 addition & 1 deletion plugins/discovery/reference/parent-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -541,7 +541,7 @@ environment variable, so it is a cost defect in a worker definition.

### The verdict lane pins `opus` at `effort: high`

*Decision.* `research-verifier` grades outcome-gate rows 4, 7 and 12, the rows the producer may
*Decision.* `research-verifier` grades outcome-gate rows 4, 7, 12 and 14, the rows the producer may
not grade, so it is a verdict lane and pins `model: opus` and `effort: high`; `explorer`,
mechanical preparation, stays on `sonnet` at `effort: medium`. *Pointer:*
[docs/plugin-philosophy.md](../../../docs/plugin-philosophy.md), "Model tiers" (the verdict rule
Expand Down
94 changes: 80 additions & 14 deletions plugins/discovery/scripts/contract.test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -534,20 +534,20 @@ assert_present 'an improvised header costs criteria 12 and 13 their evidence too
'skills/research/SKILL.md' 'costs criteria 4, 6, 9, 12 and 13 their evidence'
assert_present 'the carry-forward line lists the new header fields' \
'skills/research/SKILL.md' 'Carry this much into the read:.*measures.*inference.*qualifiers'
assert_present 'the fan-out obligation sends the synthesis to a criterion-12 verifier' \
'skills/research/context/dispatch.md' '^\*\*The synthesis .*fresh verifier for criterion 12'
assert_present 'the fan-out obligation sends the synthesis to a verifier on every verifier-owned row' \
'skills/research/context/dispatch.md' '^\*\*The synthesis .*fresh verifier for every verifier-owned row\*\* \(currently rows 4, 7, 12 and 14\)'
assert_present 'the synthesis verifier also checks claims the synthesis adds' \
'skills/research/context/dispatch.md' 'a claim the synthesis adds'
assert_present 'the SKILL.md fan-out paragraph points at the synthesis criterion-12 check' \
'skills/research/SKILL.md' '^ +\*\*Fanning out over N topics.*verifier for criterion 12'
assert_present 'the verifier is briefed on rows 4, 7 and 12 by number' \
'skills/research/context/dispatch.md' 'rows 4, 7 and 12'
assert_present 'the SKILL.md fan-out paragraph sends the synthesis to the verifier-owned rows' \
'skills/research/SKILL.md' '^ +\*\*Fanning out over N topics.*fresh verifier for the verifier-owned rows'
assert_present 'the verifier is briefed on rows 4, 7, 12 and 14 by number' \
'skills/research/context/dispatch.md' 'rows 4, 7, 12 and 14'
assert_present 'the verifier brief overrides the payload criterion string' \
'skills/research/context/dispatch.md' 'verification_request\.criterion'
assert_present 'gotchas name criterion 12 among the verifier rows' \
'skills/research/context/gotchas.md' 'Criteria 4, 7 and 12'
assert_present 'evals name criterion 12 among the verifier rows' \
'skills/research/evals/evals.json' 'criteria 4, 7 or 12'
assert_present 'evals name criteria 12 and 14 among the verifier rows' \
'skills/research/evals/evals.json' 'criteria 4, 7, 12 or 14'
assert_present 'evals grade a verbatim quote attached to a claim it does not support' \
'skills/research/evals/evals.json' 'verbatim-quote-is-not-joint-inference'
status_words="$(grep -rnE -- 'CONFLICTED|UNSUPPORTED|CONFIRMED' "$PLUGIN_ROOT/skills/research" 2>/dev/null)"
Expand All @@ -561,16 +561,16 @@ assert_absent 'no stale two-row verifier count' \
'([Cc]riteri(a|on)|rows) 4 (and|or) 7([^,0-9]|$)|[Tt]wo criteria are'
assert_present 'the gate states the Owner column governs over any other enumeration' \
'skills/research/SKILL.md' 'Owner column governs over any enumeration'
assert_present 'researcher withholds three criteria' \
'agents/researcher.md' 'Three criteria are'
assert_present 'researcher withholds four criteria' \
'agents/researcher.md' 'Four criteria are'
assert_present 'researcher lists joint inference as a withheld criterion' \
'agents/researcher.md' '^- the criterion requiring every accepted claim to follow jointly'
assert_present 'researcher verification request names joint-inference validity' \
'agents/researcher.md' '^ criterion: ".*joint-inference validity'
assert_present 'research-deep lists joint inference among the verifier rows' \
'skills/research-deep/SKILL.md' 'verifier-owned rows \(independent corroboration, HIGH confidence, joint inference\)'
assert_present 'research-deep points at the synthesis criterion-12 check' \
'skills/research-deep/SKILL.md' 'synthesized root index also goes to a fresh verifier for criterion 12'
'skills/research-deep/SKILL.md' 'verifier-owned rows \(independent corroboration, HIGH confidence, joint inference, the accepted-claim count\)'
assert_present 'research-deep sends the synthesis to the verifier on every verifier-owned row' \
'skills/research-deep/SKILL.md' 'synthesized root index also goes to a fresh verifier for every verifier-owned row, the accepted-claim count included'
assert_present 'row 12 has one pass bar: the primary measures the variable and population' \
'skills/research/SKILL.md' "^\| 12 \|.*the claim's primary source measures the claim's variable and population"
assert_present 'a non-measuring corroborator is recorded, not counted' \
Expand Down Expand Up @@ -1136,7 +1136,7 @@ assert_present 'the Phase 1 gap list names criterion 7' \
assert_present 'each claim can carry subject_pool' \
'skills/research/context/artifact-shape.md' '^ {4}subject_pool: '
assert_present 'the researcher names criterion 7 beside the corroborator floor' \
'agents/researcher.md' '^- the criterion requiring ≥2 \*\*independent\*\* corroborators per claim.*criterion 7'
'agents/researcher.md' '^- the criterion requiring ≥2 \*\*independent\*\* corroborators per accepted claim.*criterion 7'
assert_present 'the researcher verification request names single-publisher labeling' \
'agents/researcher.md' '^ criterion: ".*single-publisher'
for field in pool subject_pool; do
Expand All @@ -1160,6 +1160,72 @@ assert_present 'how to invoke gives the worktree-isolation reason' \
assert_present 'research-deep fan-out says a sub-slice is not named git' \
'skills/research-deep/SKILL.md' 'Do not name a sub-slice `git`'

# ---------------------------------------------------------------------------
# 22. A run with zero accepted claims says so (#5833)
#
# Rows 4, 7 and 12 are quantified over accepted claims, so an artifact whose
# every claim is a Gap passed all three vacuously and read like an answer.
# Row 14 makes the count part of the artifact: `accepted:` in the index, and
# at zero an inconclusive line, which passes rather than fails.
# ---------------------------------------------------------------------------
assert_present 'gate row 4 is quantified over accepted claims' \
'skills/research/SKILL.md' '^\| 4 \| Every accepted claim has ≥2 INDEPENDENT'
assert_absent 'no gate row 4 quantified over every claim' \
'^\| 4 \| Every claim has'
assert_present 'gate row 14 grades the accepted count and the inconclusive line, owned by the verifier' \
'skills/research/SKILL.md' '^\| 14 \|.*`accepted:`.*`Inconclusive: no claim accepted\.`.*A zero with that line passes.*\| \*\*verifier\*\* \|'
assert_present 'the verifier dispatch block names row 14' \
'skills/research/SKILL.md' '^ +Rows: 4, 7, 12, 14"$'
assert_present 'the Summary opens with the inconclusive line at zero accepted' \
'skills/research/SKILL.md' '^1\. \*\*Summary\*\*.*`Inconclusive: no claim accepted\.`'
assert_present 'the index frontmatter carries accepted:' \
'skills/research/context/artifact-shape.md' '^\*\*`accepted:` is the number of accepted claims\*\*'
assert_present 'the verifier names row 14 among its rows' \
'agents/research-verifier.md' 'currently 4, 7, 12 and 14'
assert_present 'the verifier return block carries row 14' \
'agents/research-verifier.md' '^ "14": pass'
assert_present 'the verifier grades a zero-accepted artifact under row 14' \
'agents/research-verifier.md' 'Rows 4, 7 and 12 hold vacuously when no claim is accepted'
assert_present 'the researcher sets accepted: in its final write' \
'agents/researcher.md' 'sets the frontmatter.s `accepted:` count'
assert_present 'the parent contract names row 14 among the verifier rows' \
'reference/parent-contract.md' 'rows 4, 7, 12 and 14'
assert_absent 'no stale three-row verifier count' \
'on (rows|criteria) 4, 7 and 12[ .]|currently (rows )?4, 7 and 12|outcome-gate rows 4, 7 and 12|criteria 4, 7 or 12;|Three criteria are|Rows: 4, 7, 12"'
assert_present 'evals cover an artifact whose every claim is a Gap' \
'skills/research/evals/evals.json' '"name": "every-claim-a-gap-is-stated-inconclusive"'
assert_present 'evals cover a count that includes a Gap claim' \
'skills/research/evals/evals.json' '"name": "accepted-count-excludes-gap-claims"'

# ---------------------------------------------------------------------------
# 23. Conflicts and the synthesis both reach row 14 (#5833 review)
#
# A rejected claim may be recorded only under Conflicts (row 12's Gap-or-
# Conflicts route, a refuted engine finding), so a count that excluded only
# Gaps stayed nonzero when every claim took that route. And a synthesized root
# index went to its verifier for criterion 12 alone, so row 14 never ran on it.
# ---------------------------------------------------------------------------
assert_present 'gate row 14 excludes Gap claims and unresolved Conflicts claims from the count' \
'skills/research/SKILL.md' '^\| 14 \| The index.s `accepted:` counts the claims neither in Gaps nor unresolved in Conflicts;'
assert_present 'artifact-shape counts a Conflicts claim only when the entry resolves in its favor' \
'skills/research/context/artifact-shape.md' 'A claim recorded under Conflicts counts only when its entry resolves in the claim.s favor'
assert_present 'the verifier excludes unresolved Conflicts claims from the count it compares' \
'agents/research-verifier.md' 'nor left unresolved in Conflicts'
assert_present 'the researcher counts neither Gap nor unresolved Conflicts claims' \
'agents/researcher.md' 'claims in neither Gaps nor an unresolved Conflicts entry'
assert_absent 'no accepted count defined by the Gaps section alone' \
'counts the claims not listed under Gaps|evidence table presents that the Gaps section does not list'
assert_present 'evals cover an artifact whose every claim is rejected under Conflicts' \
'skills/research/evals/evals.json' '"name": "every-claim-in-conflicts-is-stated-inconclusive"'
assert_present 'the synthesis verifier counts the root index under row 14' \
'skills/research/context/dispatch.md' 'Row 14 counts the synthesized index.s own accepted claims'
assert_present 'the verifier reads the sub-slice indexes a synthesized root names' \
'agents/research-verifier.md' 'A synthesized slice-root index is the one exception'
assert_present 'research-deep evals send the synthesized root to rows 4, 7, 12 and 14' \
'skills/research-deep/evals/evals.json' 'synthesized slice-root RESEARCH.md goes to a fresh verifier on rows 4, 7, 12 and 14'
assert_absent 'no synthesis verifier briefed on criterion 12 alone' \
'verifier for criterion 12'

printf '\n'
if [[ "$fails" -eq 0 ]]; then
printf 'All contract assertions passed.\n'
Expand Down
Loading
Loading