diff --git a/docs/specs/provenance-capability-matrix.md b/docs/specs/provenance-capability-matrix.md index 36c8c9ad92..56c30b732b 100644 --- a/docs/specs/provenance-capability-matrix.md +++ b/docs/specs/provenance-capability-matrix.md @@ -143,6 +143,8 @@ Evidence-gated tiers, fixed mapping (S4-adopted): - fingerprint-confirmed: a matched span above the separation rule against an identity-checked fetched source. Fix-eligible; the only tier that reaches the relay for copy findings. +- vendored-snapshot: source read from a committed snapshot because the live fetch was + unavailable or failed. Human flag, report-only, never fix-eligible. - source-fetched-similar: source fetched, similarity below the deterministic rule, judges say copy. Human flag, report-only. - llm-suspected: no lexical evidence possible (paraphrase, summary). Report-only, permanently. diff --git a/docs/specs/provenance-type-inventory.md b/docs/specs/provenance-type-inventory.md index 8364eb6d86..ccd52d99ce 100644 --- a/docs/specs/provenance-type-inventory.md +++ b/docs/specs/provenance-type-inventory.md @@ -10,6 +10,7 @@ Contracts and shapes the implementation binds to. Working plugin name: `provenan | Value | Evidence gate | Reaches relay | Fix-eligible | |---|---|---|---| | `fingerprint-confirmed` | Matched span above the separation rule against an identity-checked fetched source | Yes | Yes | +| `vendored-snapshot` | Source read from a committed snapshot because the live fetch was unavailable or failed; the finding records `source.route: vendored-snapshot` and names the snapshot path, its declared upstream ref and its sync date | No (human report) | No | | `source-fetched-similar` | Source fetched; below the deterministic rule; unanimous judge verdict STANDS | No (human report) | No | | `llm-suspected` | No lexical evidence possible (paraphrase, summary) | No (human report) | No | | `not-found` | Budgets exhausted without a source; every searched surface named | No (human report) | No | @@ -133,7 +134,7 @@ silently collapsed to the documented default. | `/audit/rule-stamp-expired` | A four-part record whose as-of date exceeds the configured window (fired values: date, window, days over) | CRITICAL fails identically. IMPORTANT's degradation limb matches with a named trigger: the record's currency ceiling has lapsed, and the first reader acting on the stamped claim without the re-fetch the convention requires acts on an assertion nobody has re-derived. | IMPORTANT | No. The repair is re-deriving the record against its live basis, a judgment the relay surfaces, never applies | | `/audit/rule-trigger-less-stamp` | Repo-override only: a dated stamp whose surface states no recheck trigger | The stated-rule limb directly: the consuming repo that enables this check has adopted the upstream-drift required parts, and a trigger-less stamp violates part 4. Portable default stays off because the fleet's stamp forms are not uniformly greppable and a guessing gate converts signal to noise. | IMPORTANT | No. Writing the missing trigger is a judgment about what observable event guards the claim | -Judgment verdicts (`source-fetched-similar`, `llm-suspected`, split rubric outcomes) have NO +Judgment verdicts (`vendored-snapshot`, `source-fetched-similar`, `llm-suspected`, split rubric outcomes) have NO rows: they never reach the relay (the ai-slop V1 boundary, restated in the Brief). The fail-safe direction holds structurally: the deterministic rules have no withholding verdicts, and every LLM uncertainty falls toward a report-only tier, never toward silence; each tier is diff --git a/plugins/attribution/.claude-plugin/plugin.json b/plugins/attribution/.claude-plugin/plugin.json index d88ccf95af..3b52c16f5d 100644 --- a/plugins/attribution/.claude-plugin/plugin.json +++ b/plugins/attribution/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "attribution", - "version": "0.6.4", + "version": "0.7.0", "description": "Finds prose in tracked markdown that restates content an external source owns (vendor docs, blogs, articles) without adequate attribution, confirms the source, and refactors the copy into a pointer, a citation, or a dated stamped record. Documentation attribution, not software supply chain. Nomination and judgment are LLM work; the scripts do only reasoning-free work (corpus scoping, breadcrumb extraction, stamp expiry, fingerprint compare of two concrete texts). Read-only audit by default; explicit fix and sweep actions apply dispositions behind a semantic-diff guard and live pointer verification. Findings conform to the detector-findings convention.", "author": { "name": "Melodic Software", diff --git a/plugins/attribution/CHANGELOG.md b/plugins/attribution/CHANGELOG.md index d1148fd589..bda14e8067 100644 --- a/plugins/attribution/CHANGELOG.md +++ b/plugins/attribution/CHANGELOG.md @@ -1,5 +1,99 @@ # Changelog +## [0.7.0] - 2026-09-29 + +### Added + +- **`vendored-snapshot` is an evidence tier of its own.** `reference/rubric.md` "Tier mapping" + gains a row between `fingerprint-confirmed` and `source-fetched-similar`: the source was read + from a committed snapshot because the live fetch was unavailable or failed, the finding records + `source.route: vendored-snapshot` and names the snapshot path, its declared upstream ref and its + sync date, and it never reaches the relay and is never fix-eligible. The 0.4.0 rule that a + snapshot basis caps at `source-fetched-similar` and borrows that tier is replaced in + `reference/source-fetch.md` and `SKILL.md`, and the README and the two `docs/specs/provenance-*` + tier lists carry the new row. The fix-eligibility rule itself is unchanged. + + `emit-findings.sh` recognizes the name as a withheld judgment verdict. Before, a finding + declaring it and carrying no rule id printed verbatim into `## Unparsed`, tier name and payload + included, and one paired with a copy rule id was counted as "not relay-eligible" rather than as a + judgment finding. Both now take the withheld path and are counted under "judgment findings". The + suite pins both shapes. + + **The rubric stays at version 4.** Its header rule names carve-outs and criteria, and a tier row + is neither: judges grade carve-outs and criteria before any tier is mapped, no grade changes, and + no golden case declares this tier, so no recorded measurement is invalidated. The row joins + version 4 before any measurement is pinned to that version (Refs #3465). + +- **`scripts/sweep-ledger.sh` keeps the sweep ledger.** `init`, `close `, `spend `, + `cache-add`, `cache-check`, and `status` manage `.work//sweep-ledger.md` in the + current checkout, so a resumed sweep restores its closures, its running fetch spend and its + source cache instead of relying on a hand-kept file. `init` gives the sweep an id and records the + checkout it started in; every call refuses, with exit 3, a ledger that names another checkout or + none, so a copy carried elsewhere is a new sweep, not a resume. `close` refuses an entry missing + the file, the dispositions or any of the four guard outcomes, and a file already closed, and + stamps the running spend on each closure. `spend` sums across separate invocations. + `cache-check` reports a hit for re-validation with its recorded hash and fetch time, never as + something to reuse. `status` prints the sweep id, the closed files, spend against + `corpus_fetch_ceiling` (read through the config layers) and the cache size, and exits 1 once + spend reaches the ceiling. With no ledger in the checkout it says + `no ledger here: this is a new sweep (no closures, no spend, no cache)`. + + **The script checks an entry's shape and the spend's arithmetic, not whether a disposition is + right.** It cannot know how many findings a file had, so every field remains the run's own + claim. `SKILL.md` "Sweep", `reference/dispositions.md` "Sweep closure" and + `reference/source-fetch.md` "Budgets, caching, and stopping" no longer say the ledger is written + by hand. The 0.5.1 entry recording that no machinery existed is left as recorded. The suite is + `scripts/sweep-ledger.test.sh` (#5353, Refs #3465). + +### Changed + +- **`fix` and `sweep` apply dispositions to hand-written markdown only.** The generated-output + paragraph in `reference/dispositions.md` was a flagged gap and is now the rule: a file whose head + carries a generated-output marker is not edited, its findings are reported and routed to the + human, and each names the generator's input as the fix site. The exclusion is the marker, never a + list of files, and no script enforces the check. `SKILL.md` "Sweep" states the same in one + sentence (Refs #3465). +- **`SKILL.md`'s description no longer enumerates the tier names**, so a tier added later does not + leave it stale. +- **Golden-set re-score against rubric version 4: the deterministic layer was re-run, the judgment + panel was not.** `fingerprint.mjs compare` over all ten `case.md` and `source.md` pairs reproduces + every figure the fixtures record, and the separation rule fires on the same six cases and stays + silent on the same four: + + | Case | Containment | Jaccard | Longest span (words) | Rule fires | + |---|---|---|---|---| + | `c01` | 0.643 | 0.336 | 76 | yes | + | `c02` | 0.713 | 0.477 | 59 | yes | + | `c03` | 0.436 | 0.208 | 22 | yes | + | `c04` | 0.507 | 0.325 | 68 | yes | + | `c05` | 0.000 | 0.000 | 0 | no | + | `c06` | 0.039 | 0.014 | 7 | no | + | `c07` | 0.000 | 0.000 | 0 | no | + | `c08` | 0.570 | 0.312 | 22 | yes | + | `c09` | 0.413 | 0.178 | 10 | yes | + | `c10` | 0.000 | 0.000 | 0 | no | + + **Not run: the blind judgment panel.** Version 4 is scored by the three-judge panel per case the + 0.4.0 re-score used, thirty independent judges that see the candidate, the fetched source, the + containing file and the rubric and never `expected.json` or another judge's verdict. This run had + no subagent tool, and it had read every `expected.json` before any grading, so an inline grade + would be neither blind nor a panel. No tp, fp, fn or tn, and no precision or recall, is therefore + pinned to version 4. The version-3 table (8 tp, 0 fp, 0 fn, 2 tn) stays superseded and is not + restated as a version-4 claim. Running the panel and recording its table is what remains of + #5354. + + **No verdict moved, and none could be measured as moving.** Version 4 differs from version 3 in + carve-out 5 and in the tier table's `vendored-snapshot` row, which joined version 4 before its + first measurement, so no version-4 figure predates it. All ten cases are ordinary local + notes, none is a distilling surface with a Sources section, so the carve-out 5 qualifier has + nothing to act on, and no case declares the new tier. That reading predicts every expected + verdict holds; it is a prediction from the fixtures, not a panel result, and no `expected.json` + was edited. + + **Every class stays below `min_n_per_class` 10:** near-verbatim n = 5, verbatim n = 2, paraphrase + n = 1, hard negatives n = 2 (10 cases). The class-size gate therefore holds whatever the panel + returns, and no class is fix-eligible (Refs #3465). + ## [0.6.4] - 2026-09-29 ### Fixed diff --git a/plugins/attribution/README.md b/plugins/attribution/README.md index 7138276abd..8fe1fba26d 100644 --- a/plugins/attribution/README.md +++ b/plugins/attribution/README.md @@ -41,6 +41,7 @@ Evidence tiers are discrete and evidence-gated, never verbalized probabilities: | Tier | Evidence | Fix-eligible | |---|---|---| | `fingerprint-confirmed` | matched span above the separation rule against an identity-checked source | yes | +| `vendored-snapshot` | source read from a committed snapshot because the live fetch failed | no, human report | | `source-fetched-similar` | source fetched, below the deterministic rule, judges unanimous | no, human report | | `llm-suspected` | no lexical evidence is possible (paraphrase, summary) | no, human report | | `not-found` | budgets exhausted; every searched surface is named | no, human report | @@ -122,10 +123,11 @@ prints one warning naming it and the `attribution` file name to rename it to. ## Prerequisites - **bash** for `list-corpus.sh`, `extract-breadcrumbs.sh`, `check-stamps.sh`, - `emit-findings.sh`, and `score-golden.sh`. + `emit-findings.sh`, `score-golden.sh`, and `sweep-ledger.sh`. - **Node** for `fingerprint.mjs`, the one module with real data structures. - **Web fetch** for source confirmation. Without it, the audit still runs and reports, but every - finding that would have been verified stops at `llm-suspected` and nothing is fix-eligible. + finding that would have been verified stops at `llm-suspected`, or at `vendored-snapshot` where + an in-repo snapshot is the only basis, and nothing is fix-eligible. - **Web search**, optional. It is the enrichment branch used only when no breadcrumb names a candidate source. Without it the audit degrades to breadcrumb-only resolution: passages whose source is already cited nearby still reach `fingerprint-confirmed`, and the rest land on diff --git a/plugins/attribution/skills/audit/SKILL.md b/plugins/attribution/skills/audit/SKILL.md index c0613dc747..e51024630d 100644 --- a/plugins/attribution/skills/audit/SKILL.md +++ b/plugins/attribution/skills/audit/SKILL.md @@ -1,9 +1,9 @@ --- -description: "Audit tracked markdown for prose restating content an external source owns without a pointer or a stamped record, and convert copies into links, quoted citations, or four-part stamped records. Breadcrumb-first: citations in or near a passage are the first confirm targets; budgeted search runs only when no breadcrumb exists. Findings carry evidence-gated tiers (fingerprint-confirmed, source-fetched-similar, llm-suspected, not-found); only fingerprint-confirmed copies are fix-eligible. Also flags verification stamps past their expiry window. Use when: 'find copied content', 'is this copied from the docs', 'check our docs for copied text', 'replace copies with links', 'find stale verification stamps', 'audit provenance', 'where did this paragraph come from', or before publishing prose that restates an upstream page. Read-only by default; explicit 'fix' applies dispositions behind a semantic-diff guard and live pointer checks, and 'sweep' adds per-file closure. Empty target audits tracked markdown." +description: "Audit tracked markdown for prose restating content an external source owns without a pointer or a stamped record, and convert copies into links, quoted citations, or four-part stamped records. Breadcrumb-first: citations in or near a passage are the first confirm targets; budgeted search runs only when no breadcrumb exists. Findings carry evidence-gated tiers; only fingerprint-confirmed copies are fix-eligible. Also flags verification stamps past their expiry window. Use when: 'find copied content', 'is this copied from the docs', 'check our docs for copied text', 'replace copies with links', 'find stale verification stamps', 'audit provenance', 'where did this paragraph come from', or before publishing prose that restates an upstream page. Read-only by default; explicit 'fix' applies dispositions behind a semantic-diff guard and live pointer checks, and 'sweep' adds per-file closure. Empty target audits tracked markdown." argument-hint: "[audit|fix|sweep] [target]" user-invocable: true disable-model-invocation: false -allowed-tools: ["Bash(${CLAUDE_SKILL_DIR}/scripts/list-corpus.sh:*)", "Bash(\"${CLAUDE_SKILL_DIR}/scripts/list-corpus.sh\":*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/extract-breadcrumbs.sh:*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/check-stamps.sh:*)", "Bash(\"${CLAUDE_SKILL_DIR}/scripts/check-stamps.sh\":*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/emit-findings.sh:*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/score-golden.sh:*)", "Bash(node ${CLAUDE_SKILL_DIR}/scripts/fingerprint.mjs:*)", "Bash(git:*)", "Bash(jq:*)", "Bash(grep:*)", "Bash(head:*)", "Bash(wc:*)"] +allowed-tools: ["Bash(${CLAUDE_SKILL_DIR}/scripts/list-corpus.sh:*)", "Bash(\"${CLAUDE_SKILL_DIR}/scripts/list-corpus.sh\":*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/extract-breadcrumbs.sh:*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/check-stamps.sh:*)", "Bash(\"${CLAUDE_SKILL_DIR}/scripts/check-stamps.sh\":*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/emit-findings.sh:*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/score-golden.sh:*)", "Bash(${CLAUDE_SKILL_DIR}/scripts/sweep-ledger.sh:*)", "Bash(node ${CLAUDE_SKILL_DIR}/scripts/fingerprint.mjs:*)", "Bash(git:*)", "Bash(jq:*)", "Bash(grep:*)", "Bash(head:*)", "Bash(wc:*)"] shell: bash metadata: workflow-stage: anytime @@ -124,8 +124,8 @@ texts, file composition); every judgment about whether a passage is a copy is mo 9. **Map the tier**, by fixed rule from the evidence, never from a judge's confidence. A paraphrase can never be `fingerprint-confirmed`: no lexical evidence is possible for one, and unanimity does not manufacture any. A finding whose only basis is an in-repo vendored - snapshot, reached because every live fetch failed, caps at `source-fetched-similar` and is - never fix-eligible; the full rule is in + snapshot, reached because every live fetch failed, caps at the report-only `vendored-snapshot` + tier and is never fix-eligible; the full rule is in [`reference/source-fetch.md`](reference/source-fetch.md). When `accuracy.review_agents` > 0, run the review pass over STANDS verdicts; a veto never reassigns a tier, it forces `leave-with-reason`. @@ -183,24 +183,31 @@ an explicit neutral outcome**, never when the interesting ones are done. Write e the sweep ledger at `.work//sweep-ledger.md` in the run's memory slice, so an interrupted sweep resumes without re-deciding closed files and the closure count is a fact rather than a memory. The entry's required fields are in -[`reference/dispositions.md`](reference/dispositions.md) "Sweep closure". +[`reference/dispositions.md`](reference/dispositions.md) "Sweep closure". Like `fix`, it applies +dispositions to hand-written markdown only: a file whose head carries a generated-output marker is +reported and routed to the human, and its finding names the generator's input as the fix site. -**Nothing writes or reads that ledger for you.** No script in this plugin creates it, parses it, -or checks an entry for completeness. It is a file the run keeps by hand, and every resume rule -below holds only as far as the run kept it honestly. +**`${CLAUDE_SKILL_DIR}/scripts/sweep-ledger.sh --topic ` keeps that ledger:** `init` +(creates it under a sweep id, or reports that it exists on a resume), `close `, `spend `, +`cache-add`, `cache-check`, and `status`, which lists the closed files a resume skips. It checks an entry's shape and the spend's arithmetic, and nothing +more. It cannot tell whether a disposition is right or whether every finding in a file is +accounted for, so each field stays the run's own claim. **The fetch ceiling and the response cache are scoped to the sweep, not to one invocation.** -`corpus_fetch_ceiling` is spent across the whole sweep, so carry the running spend into the -ledger beside each closure and, on resume, read it back and continue from that number instead of -starting again at zero. The cache is per-sweep for the same reason: record which sources the -sweep holds and when each was fetched, and on resume re-validate an entry before you reuse it, -because a page fetched before the interruption may have changed since. Reusing an entry unseen -means reporting on a body nobody in this sweep read. +`corpus_fetch_ceiling` is spent across the whole sweep. Record each batch of fetches with `spend` +as you make them: `close` stamps the running total on every closure, and on a resume `status` +reads the total back, so you continue from it instead of starting again at zero, and it exits +non-zero once spend reaches the ceiling. The cache is per-sweep for the same reason: `cache-add` +each fetched source with its sha256, and on a resume `cache-check` reports the entry for +re-validation, never as something to reuse. Fetch it again, compare the hash, and spend that +fetch, because a page fetched before the interruption may have changed since. Reusing an entry +unseen means reporting on a body nobody in this sweep read. **The ledger is checkout-local.** It lives under this checkout's `.work/` and is never tracked, -so no other checkout can see it. A sweep resumed where the ledger is not is a new sweep: it -carries no closures, no spend, and no cache, and it says so in its report rather than presenting -itself as a continuation. +so no other checkout can see it. A sweep resumed where the ledger is not is a new sweep: `status` +there says so, it carries no closures, no spend, and no cache, and the report says so rather than +presenting itself as a continuation. The ledger names the sweep id and the checkout it started +in, so a copy carried to another checkout is refused (exit 3) rather than resumed. ## Configuration @@ -226,8 +233,8 @@ fired on an identifier, a test runner exiting non-zero without failing. - **Does not fix on bare invocation.** Mutation rides only the explicit `fix` or `sweep` argument. -- **Does not put judgment verdicts in the findings file.** `source-fetched-similar`, - `llm-suspected`, and `not-found` reach the human report only. They have no crosswalk row to +- **Does not put judgment verdicts in the findings file.** `vendored-snapshot`, + `source-fetched-similar`, `llm-suspected`, and `not-found` reach the human report only. They have no crosswalk row to look a tier up from, and a relay row is an instruction to a remediation surface. - **Does not treat a missing source as evidence.** `not-found` names every surface checked and concludes nothing about the passage. `scripts/emit-findings.sh` refuses a sidecar whose diff --git a/plugins/attribution/skills/audit/context/persist-findings.md b/plugins/attribution/skills/audit/context/persist-findings.md index ff83254fa0..d643eeaa42 100644 --- a/plugins/attribution/skills/audit/context/persist-findings.md +++ b/plugins/attribution/skills/audit/context/persist-findings.md @@ -65,8 +65,8 @@ says" below. ## The relay boundary, and why the script enforces it **Only fingerprint-confirmed copy findings and the two deterministic stamp rules enter the -file.** Judgment verdicts go to the human report only: `source-fetched-similar`, -`llm-suspected`, and the neutral outcome `not-found`. They have no crosswalk row to look a tier up +file.** Judgment verdicts go to the human report only: `vendored-snapshot`, +`source-fetched-similar`, `llm-suspected`, and the neutral outcome `not-found`. They have no crosswalk row to look a tier up from, and a relay row is an instruction to a remediation surface, not a place to record a suspicion. @@ -204,9 +204,9 @@ already gives a verdict name spelled in a `note`. It has to be: a `verdict.tier` llm-suspected nomination was overruled" is a review note, and withholding the fingerprint-confirmed copy that carries it is the same drop as reading a `tier` key at any depth. -**Five names, and one reader for every question about a WELL-FORMED record.** The three withheld -verdicts, counting both spellings of the neutral one, plus `fingerprint-confirmed`, the one tier -a copy finding may be relayed on. The searched-surfaces refusal, the withhold predicate and the +**One reader for every question about a WELL-FORMED record.** It knows the withheld verdicts, +counting both spellings of the neutral one, plus `fingerprint-confirmed`, the one tier a copy +finding may be relayed on. The searched-surfaces refusal, the withhold predicate and the eligibility test all ask that one reader. A record that is not an object is the stated exception: it has no declared tier for any of them to read, so the boundary withholds it on a verdict name appearing anywhere inside it and the schema check never runs on it. Refusing a whole sidecar diff --git a/plugins/attribution/skills/audit/reference/dispositions.md b/plugins/attribution/skills/audit/reference/dispositions.md index e477f2c67c..14f433215b 100644 --- a/plugins/attribution/skills/audit/reference/dispositions.md +++ b/plugins/attribution/skills/audit/reference/dispositions.md @@ -83,12 +83,14 @@ to the human with the guard's own reason. 4. **Carve-outs re-checked at edit time.** A passage inside a quotation context, a conforming stamped record, or a vendored tree is not edited, even if a finding reached this point. -One gap is flagged here rather than guarded: a corpus file can be the rendered output of a -generator whose source of record lives outside the markdown corpus, in a data file whose -rendering carries a never-hand-edit marker in its own header. A fix applied to such a rendering edits a file its -own header forbids editing, and the next regeneration overwrites it. Check the file head for a -generated-output marker before applying any disposition; when one is present, route the finding -to the human and name the generator's input as the real fix site. No script enforces this check. +**`fix` and `sweep` apply dispositions to hand-written markdown only.** A corpus file can be the +rendered output of a generator whose source of record lives outside the markdown corpus, in a +data file, and whose rendering carries a never-hand-edit marker in its own header. A fix applied +to such a rendering edits a file its own header forbids editing, and the next regeneration +overwrites it. Check the file head for a generated-output marker before applying any +disposition. A file that carries one is not edited: its findings are reported and routed to the +human, and each names the generator's input as the fix site. The exclusion is the marker, never +a list of files. No script enforces this check. ## The demotion path when a pointer later dies @@ -162,22 +164,29 @@ neutral outcome**, never when the interesting ones are done. Record each closure ledger with its dispositions and guard outcomes, so an interrupted sweep resumes without re-deciding files it already closed, and so the closure count is a fact rather than a memory. -The ledger is `.work//sweep-ledger.md`, and it is prose the run writes by hand. No -script creates it, reads it back, or checks that an entry is complete, so an entry is worth -exactly what the run put in it. Each closure carries five things, and an entry missing any of -them cannot support a resume: +The ledger is `.work//sweep-ledger.md`, and `scripts/sweep-ledger.sh` writes and +reads it. `close` refuses an entry missing any of the first three fields below and a file already +closed. The script checks an entry's shape and the arithmetic of the spend, never whether a +disposition is right: it cannot know how many findings a file had, so "one per finding" and the +truth of each guard outcome stay the run's own claim. Each closure carries five things, and an +entry missing any of them cannot support a resume: -1. **The file** that closed, by repo-relative path. +1. **The file** that closed, by repo-relative path (`close `). 2. **The dispositions** applied in it, one per finding, `leave-with-reason` and - `neutral-not-found` included. -3. **The guard outcomes** for that file: pointer liveness, the semantic-diff verdict, the - in-span check, and the carve-out re-check. + `neutral-not-found` included (`--dispositions`). +3. **The guard outcomes** for that file, one flag each: pointer liveness (`--pointer-liveness`), + the semantic-diff verdict (`--semantic-diff`), the in-span check (`--in-span`), and the + carve-out re-check (`--carve-out`). Any non-empty text passes, `n/a` included, because a + file that only leaves findings has no pointer to check. 4. **The fetches spent** against `corpus_fetch_ceiling`, as a running total for the sweep rather - than for the file. -5. **The cache entries** the sweep holds: each source URL with the time it was fetched, so a - resume can re-validate an entry instead of reusing it unseen. + than for the file. The run adds each batch with `spend `, and `close` stamps the running + total on the entry. +5. **The cache entries** the sweep holds: each source URL with its sha256 and the time it was + fetched (`cache-add`), so a resume can re-validate an entry (`cache-check`) instead of reusing + it unseen. Fields 4 and 5 are what make the ceiling and the cache per-sweep rather than per-invocation. The -ledger is checkout-local and never tracked; `SKILL.md` "Sweep" carries what a resume does with +ledger opens with the sweep's id and the checkout it started in (`init` writes both), and every +call refuses a ledger that names another checkout. The ledger is checkout-local and never tracked; `SKILL.md` "Sweep" carries what a resume does with these fields, and why a sweep resumed in a different checkout is a new sweep rather than a continuation of this one. diff --git a/plugins/attribution/skills/audit/reference/rubric.md b/plugins/attribution/skills/audit/reference/rubric.md index 9586a54c64..1a53dde947 100644 --- a/plugins/attribution/skills/audit/reference/rubric.md +++ b/plugins/attribution/skills/audit/reference/rubric.md @@ -262,6 +262,7 @@ prose**, and a unanimous panel does not upgrade one. | Tier | Evidence gate | Reaches relay | Fix-eligible | |---|---|---|---| | `fingerprint-confirmed` | A matched span above the separation rule, against an identity-checked fetched source | Yes | Yes | +| `vendored-snapshot` | The source was read from a committed snapshot because the live fetch was unavailable or failed; the finding records `source.route: vendored-snapshot` and names the snapshot path, its declared upstream ref and its sync date; a paraphrase or summary stays `llm-suspected` | No, human report | No | | `source-fetched-similar` | Source fetched, below the deterministic rule, unanimous STANDS | No, human report | No | | `llm-suspected` | No lexical evidence is possible (paraphrase, summary) | No, human report | No | | `not-found` | Budgets exhausted with no source; every searched surface named | No, human report | No | diff --git a/plugins/attribution/skills/audit/reference/source-fetch.md b/plugins/attribution/skills/audit/reference/source-fetch.md index 50e47d6fc4..6790dcda2d 100644 --- a/plugins/attribution/skills/audit/reference/source-fetch.md +++ b/plugins/attribution/skills/audit/reference/source-fetch.md @@ -90,17 +90,15 @@ upstream says now, and drift since that date is exactly what this audit exists t The rule, stated so no judge has to improvise it: -- **The finding caps at `source-fetched-similar` and is never fix-eligible**, whatever the - fingerprint module reports. `fingerprint-confirmed` requires an identity-checked live fetch - because fix eligibility rests on current upstream state, which a snapshot cannot establish. - Stale evidence licenses no edit. +- **The finding caps at `vendored-snapshot`, a report-only tier, and is never fix-eligible**, + whatever the fingerprint module reports. `fingerprint-confirmed` requires an identity-checked + live fetch because fix eligibility rests on current upstream state, which a snapshot cannot + establish. Stale evidence licenses no edit. A paraphrase or summary stays `llm-suspected`, as + it does against a live source. The tier has its own row in `reference/rubric.md` "Tier + mapping". - **Record `source.route: vendored-snapshot`** and name, in the finding: the snapshot path, its declared upstream ref, its sync date, and each live fetch that failed with how it failed. The human report must be able to say the basis was a committed copy, not a live read. -- **The tier is borrowed knowingly.** `source-fetched-similar` is worded for a fetched source; a - snapshot basis is admitted under it as the strongest report-only tier, and the recorded route - is what keeps the report honest about the difference. This rule caps the tier and leaves the - table in `reference/rubric.md` unchanged. - **The follow-up is human.** Recommend re-running the candidate when upstream is reachable again or the snapshot re-syncs; do not hold the finding open waiting for either. @@ -147,10 +145,14 @@ loops rather than to save money. All are config keys (`.claude/attribution.json` **Under `sweep`, both are scoped to the sweep rather than to one invocation.** The ceiling is spent across the whole sweep, so a resumed sweep restores its spend from the sweep ledger instead of starting again at zero, and the cache is likewise the sweep's: a resume re-validates an entry -before reusing it, because a page fetched before the interruption may have changed since. Neither -happens on its own. The ledger at `.work//sweep-ledger.md` is prose the run keeps by -hand, no script writes or reads it, and it is checkout-local, so a sweep resumed in a different -checkout has no spend and no cache to restore and is a new sweep. `SKILL.md` "Sweep" and +before reusing it, because a page fetched before the interruption may have changed since. The +ledger at `.work//sweep-ledger.md` is where both live, and `scripts/sweep-ledger.sh` +keeps it: `spend` adds fetches to the total, `cache-add` and `cache-check` record and look up a +source, and `status` exits non-zero once the total reaches `corpus_fetch_ceiling`. The run still +has to call them, and the script checks arithmetic, never whether a fetch was worth spending. The +ledger is checkout-local, so a sweep resumed in a different checkout has no spend and no cache to +restore and is a new sweep, and a ledger copied there is refused. The cache entries are the run's +own record, and `emit-findings.sh` never reads the ledger. `SKILL.md` "Sweep" and `reference/dispositions.md` "Sweep closure" carry the resume rules and the entry's fields. Exhausting a budget produces the neutral outcome, not a failure and not a negative verdict: diff --git a/plugins/attribution/skills/audit/scripts/emit-findings.sh b/plugins/attribution/skills/audit/scripts/emit-findings.sh index 0613215b86..b095e8c37f 100755 --- a/plugins/attribution/skills/audit/scripts/emit-findings.sh +++ b/plugins/attribution/skills/audit/scripts/emit-findings.sh @@ -16,8 +16,8 @@ # # The RELAY BOUNDARY is enforced here, not upstream. Only fingerprint-confirmed # copies and the two deterministic stamp rules may reach a findings file; -# judgment verdicts (source-fetched-similar, llm-suspected, not-found) stay in -# the human report. They are counted in `## Surfaces` rather than dropped, and +# judgment verdicts (vendored-snapshot, source-fetched-similar, llm-suspected, +# not-found) stay in the human report. They are counted in `## Surfaces` rather than dropped, and # their tier names are deliberately NOT printed — the findings file is the # apply relay's input, and a tier name in it invites a consumer to act on a # verdict this producer withheld on purpose. @@ -177,8 +177,8 @@ fi # note, and withholding the fingerprint-confirmed copy carrying it is the drop above # wearing an allowlisted key. # -# Scoped to THESE FOUR NAMES on purpose — the three withheld verdicts, and the one -# tier a copy finding may be relayed on. A tier naming none of them is a tier this +# Scoped to THESE NAMES on purpose: the withheld verdicts, and the one tier a copy +# finding may be relayed on. A tier naming none of them is a tier this # producer neither withheld nor can relay. # shellcheck disable=SC2016 # `$name` is a jq parameter, not a shell expansion. TIER_DEFS=' @@ -253,7 +253,7 @@ def names_in: # name too many can only withhold a record; recognizing one too few relays a judgment # verdict. Do not narrow this to one name. def is_verdict_name: - . == "source-fetched-similar" or . == "llm-suspected" + . == "vendored-snapshot" or . == "source-fetched-similar" or . == "llm-suspected" or . == "not-found" or . == "source-not-identified"; def is_neutral_name: . == "not-found" or . == "source-not-identified"; diff --git a/plugins/attribution/skills/audit/scripts/emit-findings.test.sh b/plugins/attribution/skills/audit/scripts/emit-findings.test.sh index 2662913e44..364e04f090 100755 --- a/plugins/attribution/skills/audit/scripts/emit-findings.test.sh +++ b/plugins/attribution/skills/audit/scripts/emit-findings.test.sh @@ -380,10 +380,17 @@ write_report withheld-llm.json '{ {"tier": "llm-suspected", "file": "c.md", "excerpt": "LEAKCANARY-LLM"} ] }' +write_report withheld-vs.json '{ + "counts": {"files": 2}, + "findings": [ + {"tier": "vendored-snapshot", "file": "e.md", "excerpt": "LEAKCANARY-VS"} + ] +}' for wcase in "nf:not-found:LEAKCANARY.example" \ "sfs:source-fetched-similar:LEAKCANARY-SFS" \ - "llm:llm-suspected:LEAKCANARY-LLM"; do + "llm:llm-suspected:LEAKCANARY-LLM" \ + "vs:vendored-snapshot:LEAKCANARY-VS"; do IFS=: read -r wname wtier wcanary <<<"$wcase" WOUT="$OUTDIR/withheld-$wname.md" run --report "$REPORTS/withheld-$wname.json" --out "$WOUT" >/dev/null 2>&1 @@ -395,6 +402,31 @@ for wcase in "nf:not-found:LEAKCANARY.example" \ assert_contains "a rule-less $wtier finding is counted, not dropped" "$WBODY" "1 judgment" done +# Paired with a copy rule id, a vendored-snapshot verdict is still a judgment +# verdict: it is counted as one, not as a copy declaring no relayable tier. +write_report withheld-vs-rule.json '{ + "counts": {"files": 2}, + "findings": [ + {"rule": "attribution/audit/rule-verbatim-copy", "tier": "vendored-snapshot", + "file": "e.md", "excerpt": "LEAKCANARY-VSRULE", + "source": {"route": "vendored-snapshot"}} + ] +}' +WVSR="$OUTDIR/withheld-vs-rule.md" +run --report "$REPORTS/withheld-vs-rule.json" --out "$WVSR" >/dev/null 2>&1 +assert_exit "a rule-paired vendored-snapshot finding exits 0" "$?" "0" +VSR_BODY="$(cat "$WVSR")" +assert_eq "a rule-paired vendored-snapshot finding is no relay row" \ + "$(grep -c '^| [0-9]' "$WVSR")" "0" +assert_not_contains "a rule-paired vendored-snapshot tier name never reaches the file" \ + "$VSR_BODY" "vendored-snapshot" +assert_not_contains "a rule-paired vendored-snapshot payload never reaches the file" \ + "$VSR_BODY" "LEAKCANARY-VSRULE" +assert_contains "a rule-paired vendored-snapshot finding is counted as judgment" \ + "$VSR_BODY" "1 judgment" +assert_not_contains "a rule-paired vendored-snapshot finding is not counted as ineligible" \ + "$VSR_BODY" "Not relay-eligible" + # Casing is not a way around the boundary: the tier is matched case-folded. write_report withheld-cased.json '{ "counts": {"files": 2}, diff --git a/plugins/attribution/skills/audit/scripts/sweep-ledger.sh b/plugins/attribution/skills/audit/scripts/sweep-ledger.sh new file mode 100755 index 0000000000..065bc0bf33 --- /dev/null +++ b/plugins/attribution/skills/audit/scripts/sweep-ledger.sh @@ -0,0 +1,329 @@ +#!/usr/bin/env bash +# Keep a sweep's closure ledger, running fetch spend and source cache in one file. +# +# sweep-ledger.sh --topic SLUG init +# sweep-ledger.sh --topic SLUG close FILE --dispositions T --pointer-liveness T +# --semantic-diff T --in-span T --carve-out T +# sweep-ledger.sh --topic SLUG spend N +# sweep-ledger.sh --topic SLUG cache-add URL SHA256 +# sweep-ledger.sh --topic SLUG cache-check URL +# sweep-ledger.sh --topic SLUG status +# sweep-ledger.sh [--topic SLUG] --show-config +# +# Reasoning-free (Brief constraint C1): this script checks that an entry has the +# fields the ledger needs and that the spend adds up. Whether a disposition is +# RIGHT, or whether a file's findings are all accounted for, is not knowable from +# an entry and is never asserted here; every field is the run's own claim. +# +# The ledger is `.work/SLUG/sweep-ledger.md` under the checkout's toplevel, so a +# worktree keeps its own and no other checkout sees it. `init` writes the sweep's +# id and the checkout it started in on the first line after the header, and every +# call refuses a ledger that does not name this checkout, so a copy carried to +# another checkout is not resumed as the same sweep. The script reads and +# writes that one file, plus the config cascade for `budgets.corpus_fetch_ceiling`. +# It is append-only: the running spend is the sum of the `spend` lines, and the +# newest `cache` line for a URL wins. +# +# Contract: reference/dispositions.md "Sweep closure". +# Exit: 0 on success, 1 when `status` finds spend at or over the ceiling, 2 on +# usage or input error (a `close` missing a field included), 3 when the ledger's +# state refuses the call (no ledger for a write, a file already closed, a ledger +# from another checkout), 4 when +# `status` needs jq to read the config layers and it is absent, 5 when the ledger +# could not be written. +set -uo pipefail + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +if [[ ! -r "$SCRIPT_DIR/lib.sh" ]]; then + echo "sweep-ledger.sh: cannot read $SCRIPT_DIR/lib.sh" >&2 + for arg in "$@"; do + if [[ "$arg" == "--show-config" ]]; then + echo "detector unavailable" + exit 0 + fi + done + exit 2 +fi +# shellcheck source=lib.sh +source "$SCRIPT_DIR/lib.sh" + +usage() { + cat <<'EOF' +sweep-ledger.sh: the sweep ledger, holding closures, fetch spend and the source cache. + +Usage: + sweep-ledger.sh --topic SLUG init + sweep-ledger.sh --topic SLUG close FILE --dispositions T --pointer-liveness T + --semantic-diff T --in-span T --carve-out T + sweep-ledger.sh --topic SLUG spend N + sweep-ledger.sh --topic SLUG cache-add URL SHA256 + sweep-ledger.sh --topic SLUG cache-check URL + sweep-ledger.sh --topic SLUG status + sweep-ledger.sh [--topic SLUG] --show-config + + init create the ledger with a sweep id, or report that it exists (a resume) + close record a closed file; refuses a missing field or a file already closed + spend add N fetches to the running total + cache-add record a fetched source with its hash and fetch time + cache-check look a source up; a hit is reported for re-validation, never as reusable + status sweep id, closed files, spend against corpus_fetch_ceiling, cache size; + exits 1 when spend is at or over the ceiling + +Values may not contain a newline or '|'. A ledger is checkout-local: none here means +a new sweep, and a ledger copied in from another checkout is refused (exit 3). +EOF +} + +die() { + echo "sweep-ledger.sh: $2" >&2 + exit "$1" +} + +SHOW_CONFIG=0 +ARGS=() +declare -A OPT=() + +while [[ $# -gt 0 ]]; do + case "$1" in + --topic | --dispositions | --pointer-liveness | --semantic-diff | --in-span | --carve-out) + require_opt_value "sweep-ledger.sh" "$@" + OPT["${1#--}"]="$2" + shift 2 + ;; + --show-config) + SHOW_CONFIG=1 + shift + ;; + --help | -h) + usage + exit 0 + ;; + -*) + echo "sweep-ledger.sh: unknown argument: $1" >&2 + usage >&2 + exit 2 + ;; + *) + ARGS+=("$1") + shift + ;; + esac +done + +TOPIC="${OPT[topic]:-}" +CMD="${ARGS[0]:-}" +ARGS=(${ARGS[@]+"${ARGS[@]:1}"}) + +# --- Roots and config ------------------------------------------------------------ +# +# The ledger lives in the checkout; the config is read from CLAUDE_PROJECT_DIR, +# falling back to the checkout, as the sibling scripts do. + +REPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)" +CONFIG_ROOT="${CLAUDE_PROJECT_DIR:-$REPO_ROOT}" +cfg_layers_init "$CONFIG_ROOT" + +HAVE_JQ=1 +command -v jq >/dev/null 2>&1 || HAVE_JQ=0 +if [[ "$HAVE_JQ" -eq 0 && "${#CFG_LAYERS[@]}" -gt 0 ]]; then + echo "sweep-ledger.sh: jq not found; config layers present but unread, using defaults" >&2 +fi + +# The last layer that defines a whole-number ceiling wins. Not read through a +# command substitution: the layer's name has to reach --show-config and status. +CEILING=200 +CEILING_FROM="" +if [[ "$HAVE_JQ" -eq 1 ]]; then + for layer in ${CFG_LAYERS[@]+"${CFG_LAYERS[@]}"}; do + v="$(jq -r '.budgets.corpus_fetch_ceiling // empty' "$layer" 2>/dev/null)" || continue + v="${v//$'\r'/}" + if [[ "$v" =~ ^[0-9]+$ ]]; then + CEILING="$((10#$v))" + CEILING_FROM="$layer" + fi + done +fi +CEILING_LABEL="(bundled default)" +[[ -z "$CEILING_FROM" ]] || CEILING_LABEL="(from $CEILING_FROM)" + +if [[ "$SHOW_CONFIG" -eq 1 ]]; then + cfg_layers_print + echo "Ledger: $REPO_ROOT/.work/${TOPIC:-}/sweep-ledger.md" + echo "Effective: corpus_fetch_ceiling=$CEILING $CEILING_LABEL" + exit 0 +fi + +# --- Inputs ---------------------------------------------------------------------- + +[[ -n "$CMD" ]] || { + usage >&2 + exit 2 +} +[[ "$TOPIC" =~ ^[A-Za-z0-9][A-Za-z0-9._-]*$ ]] || + die 2 "--topic is required and takes a slug of letters, digits, '.', '_' and '-'" + +LEDGER="$REPO_ROOT/.work/$TOPIC/sweep-ledger.md" +NO_LEDGER="no ledger here: this is a new sweep (no closures, no spend, no cache)" + +# want : the subcommand's positional arguments. +want() { + [[ "${#ARGS[@]}" -eq "$1" ]] || die 2 "$CMD takes $2" +} + +# plain : a ledger line is one line of ' | ' separated fields. +plain() { + [[ "$2" != *[$'\n\r|']* ]] || die 2 "$1 may not contain a newline or '|'" +} + +append() { + printf '%s\n' "$1" >>"$LEDGER" || die 5 "cannot write $LEDGER" +} + +SPEND=0 +SWEEP_ID="" +SWEEP_ROOT="" +SWEEP_AT="" +declare -A CLOSED=() +declare -A CACHE=() + +# load_ledger: one pass over the ledger, filling SWEEP_*, SPEND, CLOSED (file set) +# and CACHE (URL to "sha256 | fetch time", the newest line winning). A ledger that +# does not record this checkout is refused: a copy carried here is not this sweep. +load_ledger() { + local line rest n + while IFS= read -r line || [[ -n "$line" ]]; do + line="${line%$'\r'}" + case "$line" in + "- sweep: "*) + rest="${line#- sweep: }" + SWEEP_ID="${rest%% | checkout: *}" + rest="${rest#* | checkout: }" + SWEEP_ROOT="${rest%% | started: *}" + SWEEP_AT="${rest#* | started: }" + ;; + "- spend: "*) + n="${line#- spend: }" + [[ "$n" =~ ^[0-9]+$ ]] && SPEND=$((SPEND + 10#$n)) + ;; + "- closed: "*) + rest="${line#- closed: }" + CLOSED["${rest%% | *}"]=1 + ;; + "- cache: "*) + rest="${line#- cache: }" + CACHE["${rest%% | *}"]="${rest#* | }" + ;; + *) ;; + esac + done <"$LEDGER" + [[ "$SWEEP_ROOT" == "$REPO_ROOT" ]] || + die 3 "$LEDGER is not this checkout's sweep (${SWEEP_ID:-no sweep line}, started in ${SWEEP_ROOT:-no checkout recorded}): this is a new sweep; move the file aside to start one here" +} + +# need_ledger: every subcommand but init and a bare status refuses without one. +need_ledger() { + [[ -f "$LEDGER" ]] || die 3 "$NO_LEDGER; run init to start one" + load_ledger +} + +now() { date -u +%Y-%m-%dT%H:%M:%SZ; } + +case "$CMD" in +init) + want 0 "no arguments" + if [[ -f "$LEDGER" ]]; then + load_ledger + echo "ledger exists: $LEDGER (a resume of sweep $SWEEP_ID; run status)" + else + mkdir -p "$(dirname "$LEDGER")" || die 5 "cannot create $(dirname "$LEDGER")" + at="$(now)" + SWEEP_ID="$TOPIC-${at//[-:]/}" + printf '%s\n' \ + "# Sweep ledger" \ + "" \ + "Checkout-local and never tracked. One appended line per event: the running spend is the" \ + "sum of the \`spend\` lines, and the newest \`cache\` line for a URL wins." \ + "" \ + "- sweep: $SWEEP_ID | checkout: $REPO_ROOT | started: $at" >"$LEDGER" || die 5 "cannot write $LEDGER" + echo "created: $LEDGER (sweep $SWEEP_ID)" + fi + ;; +close) + want 1 "one argument, the repo-relative file that closed" + file="${ARGS[0]#./}" + [[ -n "$file" ]] || die 2 "close takes a repo-relative path, not an empty one" + [[ "$file" != /* ]] || die 2 "close takes a repo-relative path, not $file" + plain file "$file" + missing="" + entry="- closed: $file" + for field in dispositions pointer-liveness semantic-diff in-span carve-out; do + if [[ -z "${OPT[$field]:-}" ]]; then + missing+=" --$field" + else + plain "--$field" "${OPT[$field]}" + entry+=" | $field: ${OPT[$field]}" + fi + done + [[ -z "$missing" ]] || die 2 "close refused for $file, missing:$missing" + need_ledger + [[ -z "${CLOSED[$file]+x}" ]] || die 3 "already closed: $file" + append "$entry | spent: $SPEND" + echo "closed: $file ($((${#CLOSED[@]} + 1)) closed, spend $SPEND)" + ;; +spend) + want 1 "one argument, a whole number of fetches" + n="${ARGS[0]}" + [[ "$n" =~ ^[0-9]+$ && "${#n}" -le 9 ]] || die 2 "spend takes a whole number of fetches (at most 9 digits), not $n" + need_ledger + n=$((10#$n)) + append "- spend: $n" + echo "spend: +$n, $((SPEND + n)) so far" + ;; +cache-add) + want 2 "two arguments, a URL and its sha256" + url="${ARGS[0]}" + sha="${ARGS[1]}" + [[ "$url" =~ ^https?://[^[:space:]\|]+$ ]] || die 2 "cache-add takes an http(s) URL with no whitespace or '|', not $url" + [[ "${#sha}" -eq 64 && "$sha" =~ ^[0-9a-f]+$ ]] || die 2 "cache-add takes a lowercase hex sha256 (64 characters)" + need_ledger + at="$(now)" + append "- cache: $url | $sha | $at" + echo "cached: $url at $at" + ;; +cache-check) + want 1 "one argument, a URL" + url="${ARGS[0]}" + need_ledger + entry="${CACHE[$url]-}" + if [[ -z "$entry" ]]; then + echo "not cached: $url (fetch it, spend the fetch, then cache-add it)" + else + echo "re-validate: $url was fetched at ${entry#* | } with sha256 ${entry%% | *}; fetch it again and compare the hash before reusing it, spend that fetch, then cache-add the result" + fi + ;; +status) + want 0 "no arguments" + if [[ ! -f "$LEDGER" ]]; then + echo "$NO_LEDGER" + exit 0 + fi + [[ "$HAVE_JQ" -eq 1 || "${#CFG_LAYERS[@]}" -eq 0 ]] || + die 4 "jq is required to read corpus_fetch_ceiling from the config layers" + load_ledger + echo "ledger: $LEDGER" + echo "sweep: $SWEEP_ID (started $SWEEP_AT)" + echo "closed files: ${#CLOSED[@]}" + [[ "${#CLOSED[@]}" -eq 0 ]] || printf ' %s\n' "${!CLOSED[@]}" | sort + echo "spend: $SPEND of $CEILING corpus_fetch_ceiling $CEILING_LABEL" + echo "cache entries: ${#CACHE[@]}" + if [[ "$SPEND" -ge "$CEILING" ]]; then + echo "at or over corpus_fetch_ceiling: stop fetching; a candidate still unresolved takes the budget-exhausted neutral outcome" + exit 1 + fi + ;; +*) + echo "sweep-ledger.sh: unknown subcommand: $CMD" >&2 + usage >&2 + exit 2 + ;; +esac diff --git a/plugins/attribution/skills/audit/scripts/sweep-ledger.test.sh b/plugins/attribution/skills/audit/scripts/sweep-ledger.test.sh new file mode 100755 index 0000000000..906859af18 --- /dev/null +++ b/plugins/attribution/skills/audit/scripts/sweep-ledger.test.sh @@ -0,0 +1,333 @@ +#!/usr/bin/env bash +# Self-contained tests for sweep-ledger.sh. Fixtures are built inline in a tmpdir. +# Per the shell-test-helpers convention, assertion helpers are local. +# +# Every case runs from inside a throwaway git repository, so the ledger lands in +# that repository's .work/ and never in this checkout's. The load-bearing cases are +# the ones a resumed sweep leans on: spend adding up across separate invocations, +# the ceiling stopping a sweep, a second checkout seeing a new sweep, and a cache +# hit coming back as a re-validation rather than as something reusable. +set -uo pipefail + +unset GIT_DIR GIT_WORK_TREE GIT_CONFIG + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +LEDGER_SH="$SCRIPT_DIR/sweep-ledger.sh" +TEST_TMPDIR="$(mktemp -d)" +trap 'rm -rf "$TEST_TMPDIR"' EXIT + +if ! command -v jq >/dev/null 2>&1; then + echo "SKIP: jq not installed (the ceiling is read from JSON config)" >&2 + exit 0 +fi +if ! command -v git >/dev/null 2>&1; then + echo "SKIP: git not installed (the ledger root is the checkout toplevel)" >&2 + exit 0 +fi + +export HOME="$TEST_TMPDIR/home" +export CLAUDE_PROJECT_DIR="$TEST_TMPDIR/config" +mkdir -p "$HOME" "$CLAUDE_PROJECT_DIR/.claude" +git init -q "$CLAUDE_PROJECT_DIR" >/dev/null 2>&1 + +FAILED=0 +CASE_NUM=0 + +pass() { + CASE_NUM=$((CASE_NUM + 1)) + printf 'PASS: %s\n' "$1" +} +fail() { + CASE_NUM=$((CASE_NUM + 1)) + FAILED=$((FAILED + 1)) + printf 'FAIL: %s\n expected: %s\n actual: %s\n' "$1" "$2" "$3" >&2 +} +assert_exit() { + if [[ "$2" == "$3" ]]; then pass "$1"; else fail "$1" "exit $3" "exit $2"; fi +} +assert_eq() { + if [[ "$2" == "$3" ]]; then pass "$1"; else fail "$1" "$3" "$2"; fi +} +assert_contains() { + case "$2" in + *"$3"*) pass "$1" ;; + *) fail "$1" "contains: $3" "$2" ;; + esac +} +assert_not_contains() { + case "$2" in + *"$3"*) fail "$1" "does not contain: $3" "$2" ;; + *) pass "$1" ;; + esac +} + +# --- Fixtures -------------------------------------------------------------------- + +REPO="$TEST_TMPDIR/repo" +OTHER="$TEST_TMPDIR/other-checkout" +for d in "$REPO" "$OTHER"; do + mkdir -p "$d" + git -C "$d" init -q + printf '%s\n' '.work/' >"$d/.git/info/exclude" +done +# The script names the ledger from git's own spelling of the toplevel, which is not +# the mktemp spelling on macOS (/private/var) or Git for Windows (C:/...). +REPO="$(git -C "$REPO" rev-parse --show-toplevel)" +OTHER="$(git -C "$OTHER" rev-parse --show-toplevel)" +LEDGER="$REPO/.work/t1/sweep-ledger.md" +OTHER_LEDGER_DIR="$OTHER/.work/t1" + +# run : the script from inside the first checkout, stdout and stderr together. +run() { (cd "$REPO" && bash "$LEDGER_SH" "$@" 2>&1); } +run_other() { (cd "$OTHER" && bash "$LEDGER_SH" "$@" 2>&1); } + +CLOSE_FIELDS=(--dispositions "convert-to-pointer x1, leave-with-reason x1" + --pointer-liveness "live" --semantic-diff "no loss" --in-span "inside" --carve-out "unchanged") + +URL="https://example.com/docs/page" +SHA_A="$(printf 'a%.0s' {1..64})" +SHA_B="$(printf 'b%.0s' {1..64})" + +# --- Usage ----------------------------------------------------------------------- + +OUT="$(bash "$LEDGER_SH" --help 2>&1)" +assert_exit "--help exits 0" "$?" "0" +assert_contains "--help names the script" "$OUT" "sweep-ledger.sh" + +OUT="$(run status)" +assert_exit "a missing --topic exits 2" "$?" "2" +assert_contains "a missing --topic says what it takes" "$OUT" "--topic is required" + +OUT="$(run --topic ../escape init)" +assert_exit "a slug with a path separator exits 2" "$?" "2" +assert_eq "a refused slug writes nothing outside .work" "$([[ -e "$REPO/../escape" ]] && echo yes || echo no)" "no" + +OUT="$(run --topic t1 frobnicate)" +assert_exit "an unknown subcommand exits 2" "$?" "2" + +# --- No ledger: a new sweep ------------------------------------------------------ + +OUT="$(run --topic t1 status)" +assert_exit "status with no ledger exits 0" "$?" "0" +assert_eq "status with no ledger reports a new sweep" "$OUT" \ + "no ledger here: this is a new sweep (no closures, no spend, no cache)" +assert_eq "status with no ledger creates nothing" "$([[ -e "$REPO/.work" ]] && echo yes || echo no)" "no" + +OUT="$(run --topic t1 spend 1)" +assert_exit "spend with no ledger is refused, exit 3" "$?" "3" +assert_contains "spend with no ledger says it is a new sweep" "$OUT" "this is a new sweep" + +# --- init ------------------------------------------------------------------------ + +OUT="$(run --topic t1 init)" +assert_exit "init exits 0" "$?" "0" +assert_contains "init reports the ledger it created" "$OUT" "created: $LEDGER" +assert_eq "init creates the ledger under the checkout's .work" "$([[ -f "$LEDGER" ]] && echo yes || echo no)" "yes" +assert_contains "init records a sweep id" "$(cat "$LEDGER")" "- sweep: t1-20" +assert_contains "init records the checkout the sweep started in" "$(cat "$LEDGER")" "| checkout: $REPO | started: 20" +SWEEP_LINE="$(grep '^- sweep: ' "$LEDGER")" +SWEEP_ID_T1="${SWEEP_LINE#- sweep: }" +SWEEP_ID_T1="${SWEEP_ID_T1%% | *}" + +run --topic t1 spend 5 >/dev/null +OUT="$(run --topic t1 init)" +assert_exit "a second init exits 0" "$?" "0" +assert_contains "a second init reports the ledger already exists" "$OUT" "ledger exists" +assert_contains "a second init names the sweep it resumes" "$OUT" "a resume of sweep t1-20" +assert_eq "a second init keeps the sweep line" "$(grep '^- sweep: ' "$LEDGER")" "$SWEEP_LINE" +assert_contains "status names the sweep" "$(run --topic t1 status)" "sweep: t1-20" +assert_contains "a second init leaves the recorded spend alone" "$(run --topic t1 status)" "spend: 5 of" + +# --- close ----------------------------------------------------------------------- + +OUT="$(run --topic t1 close docs/a.md --dispositions "leave-with-reason x1" --pointer-liveness live)" +assert_exit "a close missing fields exits 2" "$?" "2" +assert_contains "a close missing fields names the semantic-diff field" "$OUT" "--semantic-diff" +assert_contains "a close missing fields names the in-span field" "$OUT" "--in-span" +assert_contains "a close missing fields names the carve-out field" "$OUT" "--carve-out" +assert_contains "a refused close records nothing" "$(run --topic t1 status)" "closed files: 0" + +OUT="$(run --topic t1 close docs/a.md "${CLOSE_FIELDS[@]}" --carve-out "")" +assert_exit "an empty field value exits 2" "$?" "2" + +OUT="$(run --topic t1 close /abs/a.md "${CLOSE_FIELDS[@]}")" +assert_exit "an absolute path is refused" "$?" "2" + +OUT="$(run --topic t1 close "" "${CLOSE_FIELDS[@]}")" +assert_exit "an empty path exits 2" "$?" "2" +OUT="$(run --topic t1 close ./ "${CLOSE_FIELDS[@]}")" +assert_exit "a bare ./ path exits 2" "$?" "2" +assert_contains "an empty-path close records nothing" "$(run --topic t1 status)" "closed files: 0" + +OUT="$(run --topic t1 close docs/a.md --dispositions "one | two" --pointer-liveness live --semantic-diff ok --in-span ok --carve-out ok)" +assert_exit "a value holding a pipe is refused" "$?" "2" + +OUT="$(run --topic t1 close docs/a.md "${CLOSE_FIELDS[@]}")" +assert_exit "a complete close exits 0" "$?" "0" +assert_contains "a complete close reports the count" "$OUT" "1 closed" +assert_contains "the closure carries the running spend, not the file's" "$(cat "$LEDGER")" "| spent: 5" +assert_contains "the closure carries the file" "$(cat "$LEDGER")" "- closed: docs/a.md | dispositions: convert-to-pointer x1, leave-with-reason x1" + +OUT="$(run --topic t1 close docs/a.md "${CLOSE_FIELDS[@]}")" +assert_exit "closing a closed file exits 3" "$?" "3" +assert_contains "closing a closed file names it" "$OUT" "already closed: docs/a.md" + +OUT="$(run --topic t1 close ./docs/a.md "${CLOSE_FIELDS[@]}")" +assert_exit "a ./ spelling of a closed file is still a duplicate" "$?" "3" +assert_contains "the duplicate is not recorded" "$(run --topic t1 status)" "closed files: 1" + +run --topic t1 close docs/b.md "${CLOSE_FIELDS[@]}" >/dev/null +OUT="$(run --topic t1 status)" +assert_contains "a second file counts" "$OUT" "closed files: 2" +assert_contains "status lists the first closed file for a resume to skip" "$OUT" " docs/a.md" +assert_contains "status lists the second closed file for a resume to skip" "$OUT" " docs/b.md" + +# --- spend: accumulation across invocations (the resume case) -------------------- + +OUT="$(run --topic t1 spend 7)" +assert_exit "spend exits 0" "$?" "0" +assert_contains "spend adds to the running total" "$OUT" "12 so far" + +run --topic t1 init >/dev/null +OUT="$(run --topic t1 spend 010)" +assert_contains "a leading zero is decimal, not octal" "$OUT" "22 so far" +assert_contains "status reads the summed spend back" "$(run --topic t1 status)" "spend: 22 of 200" + +OUT="$(run --topic t1 spend many)" +assert_exit "a non-numeric spend exits 2" "$?" "2" +OUT="$(run --topic t1 spend -3)" +assert_exit "a negative spend exits 2" "$?" "2" +assert_contains "refused spend leaves the total alone" "$(run --topic t1 status)" "spend: 22 of 200" + +# --- ceiling ----------------------------------------------------------------------- + +OUT="$(run --topic t1 status)" +assert_exit "status under the default ceiling exits 0" "$?" "0" +assert_contains "the default ceiling is attributed to the defaults" "$OUT" "corpus_fetch_ceiling (bundled default)" + +printf '%s\n' '{"budgets": {"corpus_fetch_ceiling": 30}}' >"$CLAUDE_PROJECT_DIR/.claude/attribution.json" +OUT="$(run --topic t1 status)" +assert_exit "status under a configured ceiling exits 0" "$?" "0" +assert_contains "the configured ceiling applies" "$OUT" "spend: 22 of 30" +assert_contains "the configured ceiling names its layer" "$OUT" "(from $CLAUDE_PROJECT_DIR/.claude/attribution.json)" + +run --topic t1 spend 7 >/dev/null +OUT="$(run --topic t1 status)" +assert_exit "one under the ceiling still exits 0" "$?" "0" + +run --topic t1 spend 1 >/dev/null +OUT="$(run --topic t1 status)" +assert_exit "spend at the ceiling exits 1" "$?" "1" +assert_contains "spend at the ceiling says to stop" "$OUT" "at or over corpus_fetch_ceiling" + +run --topic t1 spend 4 >/dev/null +OUT="$(run --topic t1 status)" +assert_exit "spend over the ceiling exits 1" "$?" "1" +assert_contains "spend over the ceiling is reported as spent" "$OUT" "spend: 34 of 30" + +# A team layer's ceiling is refined by the overlay: the later layer wins. +printf '%s\n' '{"budgets": {"corpus_fetch_ceiling": 500}}' >"$CLAUDE_PROJECT_DIR/.claude/attribution.local.json" +OUT="$(run --topic t1 status)" +assert_exit "a later layer raising the ceiling clears the stop" "$?" "0" +assert_contains "the later layer supplies the ceiling" "$OUT" "spend: 34 of 500" +rm -f "$CLAUDE_PROJECT_DIR/.claude/attribution.json" "$CLAUDE_PROJECT_DIR/.claude/attribution.local.json" + +# --- cache ------------------------------------------------------------------------- + +OUT="$(run --topic t1 cache-check "$URL")" +assert_exit "cache-check on a miss exits 0" "$?" "0" +assert_contains "a miss says so and says to fetch" "$OUT" "not cached: $URL" + +OUT="$(run --topic t1 cache-add "$URL" "$SHA_A")" +assert_exit "cache-add exits 0" "$?" "0" +assert_contains "cache-add reports the fetch time" "$OUT" "cached: $URL at 20" + +OUT="$(run --topic t1 cache-check "$URL")" +assert_exit "cache-check on a hit exits 0" "$?" "0" +assert_contains "a hit must be re-validated" "$OUT" "re-validate: $URL" +assert_contains "a hit carries the recorded hash" "$OUT" "sha256 $SHA_A" +assert_contains "a hit carries the fetch time" "$OUT" "was fetched at 20" +assert_not_contains "a hit is never offered as reusable" "$OUT" "reusable" + +run --topic t1 cache-add "$URL" "$SHA_B" >/dev/null +assert_contains "the newest entry for a URL wins" "$(run --topic t1 cache-check "$URL")" "sha256 $SHA_B" +assert_contains "a re-added URL is one cache entry" "$(run --topic t1 status)" "cache entries: 1" + +OUT="$(run --topic t1 cache-add "$URL" "not-a-hash")" +assert_exit "a malformed hash exits 2" "$?" "2" +OUT="$(run --topic t1 cache-add "$URL" "${SHA_A^^}")" +assert_exit "an uppercase hash exits 2" "$?" "2" +OUT="$(run --topic t1 cache-add "ftp://example.com/x" "$SHA_A")" +assert_exit "a non-http URL exits 2" "$?" "2" +OUT="$(run --topic t1 cache-add "https://example.com/a b" "$SHA_A")" +assert_exit "a URL with whitespace exits 2" "$?" "2" + +# --- checkout-local ---------------------------------------------------------------- + +OUT="$(run_other --topic t1 status)" +assert_exit "status in another checkout exits 0" "$?" "0" +assert_eq "another checkout sees a new sweep, not this one's ledger" "$OUT" \ + "no ledger here: this is a new sweep (no closures, no spend, no cache)" +OUT="$(run_other --topic t1 cache-check "$URL")" +assert_exit "another checkout has no cache to look in" "$?" "3" + +# A ledger file copied into another checkout is refused, not resumed as the same +# sweep: it names the checkout it started in. +mkdir -p "$OTHER_LEDGER_DIR" +cp "$LEDGER" "$OTHER_LEDGER_DIR/sweep-ledger.md" +BEFORE="$(cat "$OTHER_LEDGER_DIR/sweep-ledger.md")" +OUT="$(run_other --topic t1 status)" +assert_exit "status on a copied ledger exits 3" "$?" "3" +assert_contains "a copied ledger is called a new sweep" "$OUT" "this is a new sweep" +assert_contains "a copied ledger names its sweep and the checkout it started in" "$OUT" "($SWEEP_ID_T1, started in $REPO)" +OUT="$(run_other --topic t1 spend 1)" +assert_exit "spend on a copied ledger exits 3" "$?" "3" +OUT="$(run_other --topic t1 close docs/c.md "${CLOSE_FIELDS[@]}")" +assert_exit "close on a copied ledger exits 3" "$?" "3" +OUT="$(run_other --topic t1 cache-check "$URL")" +assert_exit "cache-check on a copied ledger exits 3, the cache is not reused" "$?" "3" +OUT="$(run_other --topic t1 init)" +assert_exit "init on a copied ledger exits 3, it is not adopted as a resume" "$?" "3" +assert_eq "a refused copied ledger is left unchanged" "$(cat "$OTHER_LEDGER_DIR/sweep-ledger.md")" "$BEFORE" + +# A ledger with no sweep line (one not started by init) records no checkout and is +# refused the same way. +mkdir -p "$REPO/.work/hand" +printf '%s\n' '# notes kept by hand' >"$REPO/.work/hand/sweep-ledger.md" +OUT="$(run --topic hand status)" +assert_exit "a ledger with no sweep line exits 3" "$?" "3" +assert_contains "a ledger with no sweep line says so" "$OUT" "no sweep line" + +# Nothing but the ignored .work/ tree is written: no tracked or untracked file +# appears in the checkout. +assert_eq "the ledger is the only thing written, under the ignored .work/" \ + "$(git -C "$REPO" status --porcelain --ignored)" "!! .work/" + +# --- --show-config ----------------------------------------------------------------- + +OUT="$(run --show-config)" +assert_exit "--show-config exits 0" "$?" "0" +assert_contains "--show-config names the ledger path" "$OUT" "$REPO/.work//sweep-ledger.md" +assert_contains "--show-config prints the effective ceiling" "$OUT" "corpus_fetch_ceiling=200 (bundled default)" + +printf '%s\n' '{"budgets": {"corpus_fetch_ceiling": 90}}' >"$CLAUDE_PROJECT_DIR/.claude/attribution.json" +OUT="$(run --topic t1 --show-config)" +assert_contains "--show-config attributes a configured ceiling to its layer" "$OUT" \ + "corpus_fetch_ceiling=90 (from $CLAUDE_PROJECT_DIR/.claude/attribution.json)" +assert_contains "--show-config names the ledger for the topic" "$OUT" "$LEDGER" +rm -f "$CLAUDE_PROJECT_DIR/.claude/attribution.json" + +# The script alone, with no lib.sh beside it: the --show-config probe reads +# "detector unavailable", never an empty configuration. +ALONE="$TEST_TMPDIR/alone" +mkdir -p "$ALONE" +cp "$LEDGER_SH" "$ALONE/sweep-ledger.sh" +OUT_NL="$(bash "$ALONE/sweep-ledger.sh" --show-config 2>/dev/null)" +assert_exit "--show-config without lib.sh exits 0" "$?" "0" +assert_eq "--show-config without lib.sh prints the unavailable marker" "$OUT_NL" "detector unavailable" +ERR_NL="$(bash "$ALONE/sweep-ledger.sh" --topic t1 status 2>&1 >/dev/null)" +assert_exit "a real run without lib.sh exits 2" "$?" "2" +assert_contains "a real run without lib.sh explains itself on stderr" "$ERR_NL" "cannot read" + +printf '\nPassed: %s Failed: %s\n' "$((CASE_NUM - FAILED))" "$FAILED" +[[ "$FAILED" -eq 0 ]]