Skip to content

feat(attribution): script the sweep ledger, add the vendored-snapshot tier, record the rubric v4 re-score - #5435

Merged
kyle-sexton merged 12 commits into
mainfrom
feat/3465-attribution-sweep-ledger-and-rescore
Sep 30, 2026
Merged

kyle-sexton merged 12 commits into
mainfrom
feat/3465-attribution-sweep-ledger-and-rescore

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Closes #5353

Refs #3465, Refs #5354

Summary

Delivers two slices of the #3465 Brief. The sweep-ledger machinery (#5353) is complete and closes here. The golden-set re-score (#5354) is only partly done: the deterministic layer is recorded, the blind judgment panel was not run, so no precision or recall is pinned to rubric v4. #5354 and #3465 stay open.

Slice C (#5351, the searched-surfaces "presence only" wording) shipped in #5406 (attribution 0.6.3), which merged first. This PR carries none of that wording: the design-threads spec and source-fetch.md take main's text, and the README sentence and the second CHANGELOG entry this branch had are dropped. The alternative (this PR replacing #5406) is moot now that #5406 has merged: B2 belongs to #5406 alone, so nothing is recorded twice.

Coordination: draft #5439 also takes attribution 0.7.0 and rewrites the rubric.md header (copy v4 plus restated-fact v1) and adds golden cases c11 to c24. Whichever of the two merges second must rebump, and must re-check the "10 cases" and "rubric v4" claims in the 0.7.0 CHANGELOG entry.

Fix

  • Vendored-snapshot evidence tier in reference/rubric.md, source-fetch.md, SKILL.md, the README and the provenance specs; emit-findings.sh withholds it as a judgment verdict. The rubric stays at version 4, and the CHANGELOG says the row joins version 4 before any measurement is pinned to it. Paraphrase and summary stay at llm-suspected under a snapshot basis.
  • scripts/sweep-ledger.sh (init, close, spend, cache-add, cache-check, status). init writes a sweep id and the checkout it started in. Every call refuses a ledger that names another checkout or none (exit 3), so a copy carried elsewhere is a new sweep, not a resume. status prints the sweep id and lists the closed files a resume skips. The docs no longer say the ledger is written by hand.
  • Golden-set re-score record: fingerprint figures reproduced for all ten cases; the panel is recorded as not run (attribution audit: add the vendored-snapshot tier row, then re-score all 10 golden cases #5354).
  • attribution 0.7.0 with its CHANGELOG entry, above main's 0.6.3.

Verification

  • sweep-ledger.test.sh: 99 passed, 0 failed. Disabling the checkout check fails 10 of them.
  • All seven plugins/attribution/skills/audit/scripts/*.test.sh suites pass (emit-findings.test.sh 400).
  • scripts/check-changelog-parity.sh --check --check-order, scripts/validate-plugins.sh, markdownlint-cli2 on the six touched markdown files, shellcheck on both scripts, and check-shell-portability.sh (paths and --awk-probe): clean.
  • attribution audit: build sweep-ledger machinery for the per-sweep fetch ceiling, cache and checkout-local ledger #5353 acceptance criteria against sweep-ledger.test.sh:
    • Spend carries over a resume: "a second init leaves the recorded spend alone" and "spend adds to the running total", each call a separate process.
    • A closed file is skipped and the count is read from the ledger: "status lists ... closed file for a resume to skip", "closing a closed file exits 3", "closed files: 2".
    • A stale or mismatched cache entry is not reused: a hit is reported "re-validate", never "reusable"; the newest entry for a URL wins; a copied ledger's cache is refused. The script never judges freshness, the run compares the hash.
    • A ledger from another checkout is not resumed: a copy placed in a second checkout is refused by status, spend, close, cache-check and init, and left unchanged.
    • Only the ignored .work/ tree is written: git status --porcelain --ignored is exactly !! .work/.
    • Docs state the script and the checkout-local rule: SKILL.md "Sweep", dispositions.md "Sweep closure", source-fetch.md.
    • No Python is touched, so the ruff wrapper does not apply.

Related

#3465 (Brief), #5353 (closed here), #5354 (stays open for the judgment panel), #5351 and #5406 (slice C, shipped there), #5439 (also takes 0.7.0), #5288.

🤖 Generated with Claude Code

kyle-sexton and others added 7 commits September 29, 2026 14:13
…ed files and the searched-surfaces claim

The audit's tier table gains a vendored-snapshot row between
fingerprint-confirmed and source-fetched-similar: report-only, never
fix-eligible, recorded as source.route vendored-snapshot with the snapshot
path, declared upstream ref and sync date. source-fetch.md and SKILL.md no
longer describe a borrowed source-fetched-similar tier, and emit-findings.sh
withholds the new name as a judgment verdict; the suite pins a rule-less and
a rule-paired finding. The rubric stays at version 4: a tier row is neither a
carve-out nor a criterion.

fix and sweep now apply dispositions to hand-written markdown only. A file
carrying a generated-output marker is reported and routed to the human, and
its finding names the generator's input as the fix site.

The searched-surfaces listing is stated as checked for presence only in the
design-threads spec and the README; no live doc claims it validates anything.

Refs #3465

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… a snapshot basis

A snapshot basis caps at vendored-snapshot rather than mapping to it, so a
paraphrase or summary compared against a snapshot still lands on
llm-suspected and the tier table gives one answer per finding. The README's
no-web-fetch line names the snapshot case as well.

Refs #3465

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…hand

sweep-ledger.sh manages .work/<topic-slug>/sweep-ledger.md in the current
checkout: init, close <file>, spend <n>, cache-add, cache-check and status.
close refuses a missing field and a file already closed, spend sums across
invocations, cache-check reports a hit for re-validation rather than reuse,
and status exits 1 once spend reaches corpus_fetch_ceiling. With no ledger it
reports a new sweep.

SKILL.md "Sweep", dispositions.md "Sweep closure" and source-fetch.md no
longer say the ledger is written by hand. The script checks entry shape and
arithmetic, not whether a disposition is right.

Refs #3465

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…toplevel

The script names the ledger from git rev-parse --show-toplevel, which differs
from the mktemp spelling on macOS and Git for Windows, so the three path
assertions in sweep-ledger.test.sh would fail there. Resolve the fixture root
through git before building the expected paths.

Refs #3465

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nd what was not run

The deterministic layer reproduces every recorded fingerprint figure on all ten
cases. The blind judgment panel was not run, so no precision or recall is pinned
to version 4. Every class stays below min_n_per_class 10.

Refs #3465

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

PR body contract — issue linkage

This PR body conforms to the issue-linkage contract. Nothing to do.

kyle-sexton and others added 2 commits September 29, 2026 16:39
…-sweep-ledger-and-rescore

Slice C (#5351) shipped in #5406 (attribution 0.6.3), so this branch carries none of the
searched-surfaces wording: the design-threads spec and source-fetch.md take main's text, the
README drops its own copy of the sentence, and the 0.7.0 CHANGELOG entry drops its duplicate
bullet. The 0.6.3 entry stays below 0.7.0 unchanged. The rubric-version note now says the
vendored-snapshot row joins version 4 before its first measurement.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…pied from another checkout

init writes a sweep id and the checkout it started in; every call refuses, with exit 3, a
ledger that names another checkout or none, so a copy carried elsewhere is a new sweep rather
than a resume. status prints the sweep id and lists the closed files a resume skips, as the
CHANGELOG already said it did. The suite adds a copied-ledger case (status, spend, close,
cache-check and init all refuse, the file is left unchanged), a hand-kept ledger with no sweep
line, and an assertion that only the ignored .work/ tree is written.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 29, 2026 21:00
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T21:04:06.499390Z 64e2aba Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 64e2abab36

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/attribution/skills/audit/scripts/sweep-ledger.sh
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 3m 50s —— View job


Security review — /review:security-review

  • Gather context (PR diff, changed files)
  • Review scripts/sweep-ledger.sh (new, 328 lines) for injection / path / trust-boundary issues
  • Review scripts/sweep-ledger.test.sh for test harness issues
  • Review scripts/emit-findings.sh diff (vendored-snapshot tier withholding)
  • Review doc/spec changes for instruction-surface deletions
  • Post final findings

No security issues found. Reviewed at 64e2aba.

Scope: the two new/changed shell scripts (sweep-ledger.sh, emit-findings.sh + tests) and the doc/spec changes that add the vendored-snapshot evidence tier.

What I checked and why it's clean:

  • sweep-ledger.sh input validation. --topic is constrained to ^[A-Za-z0-9][A-Za-z0-9._-]*$ (sweep-ledger.sh#L162) — no /, so the ledger path $REPO_ROOT/.work/$TOPIC/sweep-ledger.md can't escape .work/. cache-add validates the URL scheme/charset and requires a 64-char lowercase-hex sha256 (L285-286).
  • close FILE accepts ../ sequences (only a leading ./ is stripped, and there's no .. component check) — but file is only ever written as an opaque text field in the ledger and used as a bash associative-array key for a duplicate check; it's never used to open, read, or write a filesystem path, so there's no traversal to exploit.
  • No eval, no unquoted/unsafe expansions, no command substitution built from user-controlled strings. All values flow into the ledger via printf '%s\n' ... >>"$LEDGER" and are read back with plain string parsing (case/parameter expansion), never sourced or eval'd.
  • emit-findings.sh — the diff only adds vendored-snapshot to the existing is_verdict_name allowlist so it's withheld from the findings-file relay like the other judgment verdicts. The relay-boundary logic itself (tier reader, ## Unparsed fallback, jq program) is unchanged and was not reviewed for new issues, since nothing there changed structurally.
  • Instruction-surface lens (this skill's check for deleted/narrowed guardrails): the doc edits in dispositions.md, SKILL.md, source-fetch.md, and persist-findings.md rephrase or strengthen existing rules (e.g., the generated-output carve-out goes from "flagged, unenforced" to a stated mandatory exclusion) — none removes or weakens a Gate 0-class protection (irreversible-action, secret-handling, data-integrity, external-publication, legal-compliance, agent-authority) per docs/conventions/instruction-exception-register/README.md.
  • GitHub Actions files are not touched by this PR, so no Actions-hardening findings apply here.

An unset expansion or a bare ./ normalized to an empty path, which the
ledger recorded as a closed entry that later broke status.

Refs #3465

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 2m 0s —— View job


Claude is reviewing this PR

  • Gather context (PR body, diff, comments)
  • Invoke /review:code-review and follow it exactly
  • Post inline findings
  • Post summary

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

Keeps attribution 0.7.0 above main's 0.6.4 and puts the 0.7.0 changelog entry on top. The
sweep-ledger suite initializes its config root as a git repo, which the config-root classifier
now requires before it reads a project layer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude encountered an error after 0s —— View job


I'll analyze this and get back to you.

@kyle-sexton
kyle-sexton merged commit 8ac77ad into main Sep 30, 2026
17 of 19 checks passed
@kyle-sexton
kyle-sexton deleted the feat/3465-attribution-sweep-ledger-and-rescore branch September 30, 2026 02:18
kyle-sexton added a commit that referenced this pull request Sep 30, 2026
… golden set (#5597)

Closes #5354

## Summary

The vendored-snapshot tier row is on main (#5435), and #5435 recorded
the golden-set re-score against rubric version 4 as the deterministic
layer only, because it had no subagent tool and had read
`expected.json`. This PR runs the missing part, the blind judgment
panel, and records the result.

Result against rubric version 4: 8 tp / 0 fp / 0 fn / 2 tn, precision
1.00, recall 1.00. Thirty judges (three per case) were unanimous on
every case, and no verdict moved. Every class stays below
`min_n_per_class` 10 (verbatim 2, near-verbatim 5, paraphrase 1, hard
negatives 2), so no class becomes fix-eligible.

## Fix

- `plugins/attribution/CHANGELOG.md`: new 0.7.1 entry with the method,
the per-case panel table, the `score-golden.sh` output, per-class
precision and recall with n, the class-size confirmation, and the two
differences of route or extent (below).
- `plugins/attribution/.claude-plugin/plugin.json`: 0.7.0 to 0.7.1.
- `scripts/changelog-no-op-bumps.txt`: names `attribution@0.7.1`,
because the change set ships no file beyond the CHANGELOG and
`--check-bump` otherwise refuses the bump.

Method: each judge was a separate `claude -p` process on model `opus`
with no tools, hooks, skills or MCP, in a neutral directory. Cases were
relabelled at random, the source's fixture scaffolding paragraph and
local path were dropped, and no `expected.json` was opened until all
thirty raw outputs were saved. The combine rule was fixed before
scoring. Tier was mapped from `fingerprint.mjs` output, never from a
judge.

Two differences recorded, neither a moved verdict:

- `c02`: the panel returned one finding (lines 5-8) where the fixture
records two (5-8 and 12-14). Every judge's reasoning names 12-14 as a
second copy; the return format allowed one span per judge.
- `c07`: cleared at carve-out 4, not at C1 as the fixture's notes
explain it. The rubric's order of evaluation makes them exclusive, and
the recorded answer (no finding) holds.

Limits stated in the entry: no nomination pass ran (the whole case body
was the candidate), and the return format glossed the class names
because the rubric defines none.

`expected.json` and the rubric are unchanged.

## Verification

- `bash plugins/attribution/skills/audit/scripts/score-golden.test.sh`:
41 passed, 0 failed.
- `node plugins/attribution/skills/audit/scripts/fingerprint.test.mjs`:
40 passed, 0 failed.
- `score-golden.sh --golden
plugins/attribution/skills/audit/evals/fixtures/golden --actual <panel
findings>` with `cases_run` declaring all ten: overall
`{"scored":10,"tp":8,"fp":0,"fn":0,"tn":2,"precision":1,"recall":1}`,
all four class gates `report-only`.
- `bash scripts/check-changelog-parity.sh --check`, `--check-order` and
`--check-bump origin/main`: pass.
- `bash scripts/validate-plugins.sh`: all manifests and the catalog
validated.
- `npx markdownlint-cli2 plugins/attribution/CHANGELOG.md`: 0 issues.
- `git diff --stat origin/main` shows three files and no
`expected.json`.

## Related

- Follows #5435 (vendored-snapshot row, sweep ledger, deterministic
re-score).
- #3465 stays open; this PR does not close it.
- Draft #5439 grows the golden set to c24 and also touches attribution;
whichever of the two merges second rebumps and re-checks the "10 cases"
claim in this entry.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

attribution audit: build sweep-ledger machinery for the per-sweep fetch ceiling, cache and checkout-local ledger

1 participant