Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/specs/provenance-capability-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,8 @@ Evidence-gated tiers, fixed mapping (S4-adopted):

- fingerprint-confirmed: a matched span above the separation rule against an identity-checked
fetched source. Fix-eligible; the only tier that reaches the relay for copy findings.
- vendored-snapshot: source read from a committed snapshot because the live fetch was
unavailable or failed. Human flag, report-only, never fix-eligible.
- source-fetched-similar: source fetched, similarity below the deterministic rule, judges say
copy. Human flag, report-only.
- llm-suspected: no lexical evidence possible (paraphrase, summary). Report-only, permanently.
Expand Down
3 changes: 2 additions & 1 deletion docs/specs/provenance-type-inventory.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ Contracts and shapes the implementation binds to. Working plugin name: `provenan
| Value | Evidence gate | Reaches relay | Fix-eligible |
|---|---|---|---|
| `fingerprint-confirmed` | Matched span above the separation rule against an identity-checked fetched source | Yes | Yes |
| `vendored-snapshot` | Source read from a committed snapshot because the live fetch was unavailable or failed; the finding records `source.route: vendored-snapshot` and names the snapshot path, its declared upstream ref and its sync date | No (human report) | No |
| `source-fetched-similar` | Source fetched; below the deterministic rule; unanimous judge verdict STANDS | No (human report) | No |
| `llm-suspected` | No lexical evidence possible (paraphrase, summary) | No (human report) | No |
| `not-found` | Budgets exhausted without a source; every searched surface named | No (human report) | No |
Expand Down Expand Up @@ -133,7 +134,7 @@ silently collapsed to the documented default.
| `<name>/audit/rule-stamp-expired` | A four-part record whose as-of date exceeds the configured window (fired values: date, window, days over) | CRITICAL fails identically. IMPORTANT's degradation limb matches with a named trigger: the record's currency ceiling has lapsed, and the first reader acting on the stamped claim without the re-fetch the convention requires acts on an assertion nobody has re-derived. | IMPORTANT | No. The repair is re-deriving the record against its live basis, a judgment the relay surfaces, never applies |
| `<name>/audit/rule-trigger-less-stamp` | Repo-override only: a dated stamp whose surface states no recheck trigger | The stated-rule limb directly: the consuming repo that enables this check has adopted the upstream-drift required parts, and a trigger-less stamp violates part 4. Portable default stays off because the fleet's stamp forms are not uniformly greppable and a guessing gate converts signal to noise. | IMPORTANT | No. Writing the missing trigger is a judgment about what observable event guards the claim |

Judgment verdicts (`source-fetched-similar`, `llm-suspected`, split rubric outcomes) have NO
Judgment verdicts (`vendored-snapshot`, `source-fetched-similar`, `llm-suspected`, split rubric outcomes) have NO
rows: they never reach the relay (the ai-slop V1 boundary, restated in the Brief). The
fail-safe direction holds structurally: the deterministic rules have no withholding verdicts,
and every LLM uncertainty falls toward a report-only tier, never toward silence; each tier is
Expand Down
2 changes: 1 addition & 1 deletion plugins/attribution/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
"name": "attribution",
"version": "0.6.4",
"version": "0.7.0",
"description": "Finds prose in tracked markdown that restates content an external source owns (vendor docs, blogs, articles) without adequate attribution, confirms the source, and refactors the copy into a pointer, a citation, or a dated stamped record. Documentation attribution, not software supply chain. Nomination and judgment are LLM work; the scripts do only reasoning-free work (corpus scoping, breadcrumb extraction, stamp expiry, fingerprint compare of two concrete texts). Read-only audit by default; explicit fix and sweep actions apply dispositions behind a semantic-diff guard and live pointer verification. Findings conform to the detector-findings convention.",
"author": {
"name": "Melodic Software",
Expand Down
94 changes: 94 additions & 0 deletions plugins/attribution/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,99 @@
# Changelog

## [0.7.0] - 2026-09-29

### Added

- **`vendored-snapshot` is an evidence tier of its own.** `reference/rubric.md` "Tier mapping"
gains a row between `fingerprint-confirmed` and `source-fetched-similar`: the source was read
from a committed snapshot because the live fetch was unavailable or failed, the finding records
`source.route: vendored-snapshot` and names the snapshot path, its declared upstream ref and its
sync date, and it never reaches the relay and is never fix-eligible. The 0.4.0 rule that a
snapshot basis caps at `source-fetched-similar` and borrows that tier is replaced in
`reference/source-fetch.md` and `SKILL.md`, and the README and the two `docs/specs/provenance-*`
tier lists carry the new row. The fix-eligibility rule itself is unchanged.

`emit-findings.sh` recognizes the name as a withheld judgment verdict. Before, a finding
declaring it and carrying no rule id printed verbatim into `## Unparsed`, tier name and payload
included, and one paired with a copy rule id was counted as "not relay-eligible" rather than as a
judgment finding. Both now take the withheld path and are counted under "judgment findings". The
suite pins both shapes.

**The rubric stays at version 4.** Its header rule names carve-outs and criteria, and a tier row
is neither: judges grade carve-outs and criteria before any tier is mapped, no grade changes, and
no golden case declares this tier, so no recorded measurement is invalidated. The row joins
version 4 before any measurement is pinned to that version (Refs #3465).

- **`scripts/sweep-ledger.sh` keeps the sweep ledger.** `init`, `close <file>`, `spend <n>`,
`cache-add`, `cache-check`, and `status` manage `.work/<topic-slug>/sweep-ledger.md` in the
current checkout, so a resumed sweep restores its closures, its running fetch spend and its
source cache instead of relying on a hand-kept file. `init` gives the sweep an id and records the
checkout it started in; every call refuses, with exit 3, a ledger that names another checkout or
none, so a copy carried elsewhere is a new sweep, not a resume. `close` refuses an entry missing
the file, the dispositions or any of the four guard outcomes, and a file already closed, and
stamps the running spend on each closure. `spend` sums across separate invocations.
`cache-check` reports a hit for re-validation with its recorded hash and fetch time, never as
something to reuse. `status` prints the sweep id, the closed files, spend against
`corpus_fetch_ceiling` (read through the config layers) and the cache size, and exits 1 once
spend reaches the ceiling. With no ledger in the checkout it says
`no ledger here: this is a new sweep (no closures, no spend, no cache)`.

**The script checks an entry's shape and the spend's arithmetic, not whether a disposition is
right.** It cannot know how many findings a file had, so every field remains the run's own
claim. `SKILL.md` "Sweep", `reference/dispositions.md` "Sweep closure" and
`reference/source-fetch.md` "Budgets, caching, and stopping" no longer say the ledger is written
by hand. The 0.5.1 entry recording that no machinery existed is left as recorded. The suite is
`scripts/sweep-ledger.test.sh` (#5353, Refs #3465).

### Changed

- **`fix` and `sweep` apply dispositions to hand-written markdown only.** The generated-output
paragraph in `reference/dispositions.md` was a flagged gap and is now the rule: a file whose head
carries a generated-output marker is not edited, its findings are reported and routed to the
human, and each names the generator's input as the fix site. The exclusion is the marker, never a
list of files, and no script enforces the check. `SKILL.md` "Sweep" states the same in one
sentence (Refs #3465).
- **`SKILL.md`'s description no longer enumerates the tier names**, so a tier added later does not
leave it stale.
- **Golden-set re-score against rubric version 4: the deterministic layer was re-run, the judgment
panel was not.** `fingerprint.mjs compare` over all ten `case.md` and `source.md` pairs reproduces
every figure the fixtures record, and the separation rule fires on the same six cases and stays
silent on the same four:

| Case | Containment | Jaccard | Longest span (words) | Rule fires |
|---|---|---|---|---|
| `c01` | 0.643 | 0.336 | 76 | yes |
| `c02` | 0.713 | 0.477 | 59 | yes |
| `c03` | 0.436 | 0.208 | 22 | yes |
| `c04` | 0.507 | 0.325 | 68 | yes |
| `c05` | 0.000 | 0.000 | 0 | no |
| `c06` | 0.039 | 0.014 | 7 | no |
| `c07` | 0.000 | 0.000 | 0 | no |
| `c08` | 0.570 | 0.312 | 22 | yes |
| `c09` | 0.413 | 0.178 | 10 | yes |
| `c10` | 0.000 | 0.000 | 0 | no |

**Not run: the blind judgment panel.** Version 4 is scored by the three-judge panel per case the
0.4.0 re-score used, thirty independent judges that see the candidate, the fetched source, the
containing file and the rubric and never `expected.json` or another judge's verdict. This run had
no subagent tool, and it had read every `expected.json` before any grading, so an inline grade
would be neither blind nor a panel. No tp, fp, fn or tn, and no precision or recall, is therefore
pinned to version 4. The version-3 table (8 tp, 0 fp, 0 fn, 2 tn) stays superseded and is not
restated as a version-4 claim. Running the panel and recording its table is what remains of
#5354.

**No verdict moved, and none could be measured as moving.** Version 4 differs from version 3 in
carve-out 5 and in the tier table's `vendored-snapshot` row, which joined version 4 before its
first measurement, so no version-4 figure predates it. All ten cases are ordinary local
notes, none is a distilling surface with a Sources section, so the carve-out 5 qualifier has
nothing to act on, and no case declares the new tier. That reading predicts every expected
verdict holds; it is a prediction from the fixtures, not a panel result, and no `expected.json`
was edited.

**Every class stays below `min_n_per_class` 10:** near-verbatim n = 5, verbatim n = 2, paraphrase
n = 1, hard negatives n = 2 (10 cases). The class-size gate therefore holds whatever the panel
returns, and no class is fix-eligible (Refs #3465).

## [0.6.4] - 2026-09-29

### Fixed
Expand Down
6 changes: 4 additions & 2 deletions plugins/attribution/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,7 @@ Evidence tiers are discrete and evidence-gated, never verbalized probabilities:
| Tier | Evidence | Fix-eligible |
|---|---|---|
| `fingerprint-confirmed` | matched span above the separation rule against an identity-checked source | yes |
| `vendored-snapshot` | source read from a committed snapshot because the live fetch failed | no, human report |
| `source-fetched-similar` | source fetched, below the deterministic rule, judges unanimous | no, human report |
| `llm-suspected` | no lexical evidence is possible (paraphrase, summary) | no, human report |
| `not-found` | budgets exhausted; every searched surface is named | no, human report |
Expand Down Expand Up @@ -122,10 +123,11 @@ prints one warning naming it and the `attribution` file name to rename it to.
## Prerequisites

- **bash** for `list-corpus.sh`, `extract-breadcrumbs.sh`, `check-stamps.sh`,
`emit-findings.sh`, and `score-golden.sh`.
`emit-findings.sh`, `score-golden.sh`, and `sweep-ledger.sh`.
- **Node** for `fingerprint.mjs`, the one module with real data structures.
- **Web fetch** for source confirmation. Without it, the audit still runs and reports, but every
finding that would have been verified stops at `llm-suspected` and nothing is fix-eligible.
finding that would have been verified stops at `llm-suspected`, or at `vendored-snapshot` where
an in-repo snapshot is the only basis, and nothing is fix-eligible.
- **Web search**, optional. It is the enrichment branch used only when no breadcrumb names a
candidate source. Without it the audit degrades to breadcrumb-only resolution: passages whose
source is already cited nearby still reach `fingerprint-confirmed`, and the rest land on
Expand Down
Loading
Loading