Part of #6074
Paths below are relative to plugins/session-flow/skills/audit-sessions/ unless they start with plugins/. Line numbers are on origin/main 9671ece.
Problem
E5. tools.denials routes to a skill that addresses a different problem (medium). reference/sweep-rules.json:87-96 routes tools.denials to config-change / /fewer-permission-prompts. scripts/collect.py:368-369 records denials per toolDenialKind (stored at collect.py:654), but scripts/sweep.py:45-47 sums the kinds, so the finding cannot say which kind dominates. /fewer-permission-prompts adds allow rules (https://code.claude.com/docs/en/commands: "add a prioritized allowlist ... to reduce permission prompts"); the permissions docs say "A blocking hook also takes precedence over allow rules." and "An allow rule can't carve an exception out of a deny rule".
R1. subagents.count routes to a skill about standing instructions (medium). sweep-rules.json:97-106 routes it to unhobble-experiment / /harness-config:unhobble, whose purpose is "Bare-baseline experiment: strip a repo's standing instructions, log stumbles against the bare model". Its description has no subagent, fan-out or delegation text, and no rationale for the route is recorded (only sweep-rules.json:102 and scripts/tests/test_sweep.py:21 name it; PR #5925 and issue #5818 have no "unhobble" mention). A plausible unstated rationale may exist, so the route is undocumented, not provably wrong.
I4. The two unchecked drift rows need checks the report does not name (low). sweep.py:126-127 prints "drift not checked: no canary feeds this metric" for tools.denials and subagents.count. A plain presence canary on toolDenialKind would misfire only in a version with no denial at all, and subagents.count counts subagent files (collect.py:666), so no key canary can feed it.
Q2. The skill catalog is passed as hundreds of shell arguments (low). SKILL.md:87-94 tells the model to pass every skill in the session listing to --catalog (sweep.py:307, nargs +), but skill_label (sweep.py:96-99) only checks the five suggested_skill values in the rules file.
Evidence
Verified this pass: the lines above at 9671ece.
Measured by the audit on one Windows 11 machine, 2026-10-03 (not re-measured here):
- E5 store denial kinds: permission-rule 2,432; automode-blocked 190; user-rejected 144; automode-unavailable 56; automode-parsing-error 3; cancelled 2; interrupted 1. Of 2,582 permission-rule denials in transcripts still on disk, 2,420 carry the PreToolUse hook-block notice and 113 the deny-rule wording. The worst session had 123 hook blocks and 0 deny-rule denials. An allowlist reduces neither. The
toolDenialKind values are undocumented; their meaning beyond the notice wording is inferred.
- R1 worst sessions: 173, 110 and 109 subagents.
- I4: 16 sessions had Agent or Task tool calls but zero subagent files.
- Q2: the evidence run passed 273 names to
--catalog.
Proposed approach
E5:
- Render the per-kind split in each
tools.denials finding (markdown and JSON) and name the dominant kind.
- Replace the
/fewer-permission-prompts route. A rule has one suggested_skill. Recommended (judgment for the maintainer): /session-flow:retro (why the model kept trying blocked or unwanted calls), which fits hook blocks and human declines; name /harness-config:draft-auto-mode-rules beside the automode-blocked kind in the split. Not /harness-config:audit-permission-state as the single route: it fits only the 113 deny-rule denials.
- Record the kind reading's basis with an As of date and a recheck trigger beside the rule, since the kinds are undocumented.
R1: either record in the subagents.count basis why instruction ablation is the intended response, or route to /harness-ops:observability ("cross-session trends, cost, hooks"), matching tokens.sub (sweep-rules.json:23). Not /harness-ops:audit-performance: its engine never reads transcripts.
I4:
- Add a
denial_kind:<value> census key in collect.py's census key builder, following effort_value (collect.py:229), so a new or renamed kind surfaces as a "new" drift change; have a canary feed tools.denials from it.
- Add an expect-absent invariant (style of
invariant:usage_split, canaries.json:13) counting sessions with Agent or Task calls and zero subagent files, feeding subagents.count.
- Use the same invariant-canary mechanism as the metric-accuracy child's S2 canaries.
Q2: pass --catalog only the suggested skills present in the session's listing (the five in sweep-rules.json, updated for any route E5 or R1 changes), and update SKILL.md Step 2 and evals/evals.json:8 and :13 to match. Alternative (scope review): drop the catalog step and have the model check the five names itself.
Files: reference/sweep-rules.json, reference/canaries.json, scripts/sweep.py, scripts/collect.py, scripts/census.py if the invariant needs support there, SKILL.md, evals/evals.json, scripts/tests/test_sweep.py, scripts/tests/test_drift.py, CHANGELOG and plugin.json.
Acceptance criteria
Constraints and gotchas
Context
Source: local handoff item 20261003-070338-session-flow-audit-sessions-accuracy-and-coverage.md (findings E5, R1, I4, Q2). Related: #5985, #6033, #4261 (guard rejections costing retry turns).
Part of #6074
Paths below are relative to
plugins/session-flow/skills/audit-sessions/unless they start withplugins/. Line numbers are onorigin/main9671ece.Problem
E5.
tools.denialsroutes to a skill that addresses a different problem (medium).reference/sweep-rules.json:87-96routestools.denialstoconfig-change//fewer-permission-prompts.scripts/collect.py:368-369records denials pertoolDenialKind(stored at collect.py:654), butscripts/sweep.py:45-47sums the kinds, so the finding cannot say which kind dominates./fewer-permission-promptsadds allow rules (https://code.claude.com/docs/en/commands: "add a prioritized allowlist ... to reduce permission prompts"); the permissions docs say "A blocking hook also takes precedence over allow rules." and "An allow rule can't carve an exception out of a deny rule".R1.
subagents.countroutes to a skill about standing instructions (medium). sweep-rules.json:97-106 routes it tounhobble-experiment//harness-config:unhobble, whose purpose is "Bare-baseline experiment: strip a repo's standing instructions, log stumbles against the bare model". Its description has no subagent, fan-out or delegation text, and no rationale for the route is recorded (only sweep-rules.json:102 andscripts/tests/test_sweep.py:21name it; PR #5925 and issue #5818 have no "unhobble" mention). A plausible unstated rationale may exist, so the route is undocumented, not provably wrong.I4. The two unchecked drift rows need checks the report does not name (low). sweep.py:126-127 prints "drift not checked: no canary feeds this metric" for
tools.denialsandsubagents.count. A plain presence canary ontoolDenialKindwould misfire only in a version with no denial at all, andsubagents.countcounts subagent files (collect.py:666), so no key canary can feed it.Q2. The skill catalog is passed as hundreds of shell arguments (low).
SKILL.md:87-94tells the model to pass every skill in the session listing to--catalog(sweep.py:307, nargs+), butskill_label(sweep.py:96-99) only checks the fivesuggested_skillvalues in the rules file.Evidence
Verified this pass: the lines above at 9671ece.
Measured by the audit on one Windows 11 machine, 2026-10-03 (not re-measured here):
toolDenialKindvalues are undocumented; their meaning beyond the notice wording is inferred.--catalog.Proposed approach
E5:
tools.denialsfinding (markdown and JSON) and name the dominant kind./fewer-permission-promptsroute. A rule has onesuggested_skill. Recommended (judgment for the maintainer):/session-flow:retro(why the model kept trying blocked or unwanted calls), which fits hook blocks and human declines; name/harness-config:draft-auto-mode-rulesbeside the automode-blocked kind in the split. Not/harness-config:audit-permission-stateas the single route: it fits only the 113 deny-rule denials.R1: either record in the
subagents.countbasis why instruction ablation is the intended response, or route to/harness-ops:observability("cross-session trends, cost, hooks"), matchingtokens.sub(sweep-rules.json:23). Not/harness-ops:audit-performance: its engine never reads transcripts.I4:
denial_kind:<value>census key in collect.py's census key builder, followingeffort_value(collect.py:229), so a new or renamed kind surfaces as a "new" drift change; have a canary feedtools.denialsfrom it.invariant:usage_split, canaries.json:13) counting sessions with Agent or Task calls and zero subagent files, feedingsubagents.count.Q2: pass
--catalogonly the suggested skills present in the session's listing (the five in sweep-rules.json, updated for any route E5 or R1 changes), and update SKILL.md Step 2 and evals/evals.json:8 and :13 to match. Alternative (scope review): drop the catalog step and have the model check the five names itself.Files:
reference/sweep-rules.json,reference/canaries.json,scripts/sweep.py,scripts/collect.py,scripts/census.pyif the invariant needs support there,SKILL.md,evals/evals.json,scripts/tests/test_sweep.py,scripts/tests/test_drift.py, CHANGELOG and plugin.json.Acceptance criteria
tools.denialsrule'ssuggested_skillis not/fewer-permission-prompts.tools.denialsfinding, in markdown and JSON, shows the count per denial kind and names the dominant kind.tools.denialsrule's basis records how the kinds were read, with an As of date.subagents.countbasis states why/harness-config:unhobbleis the intended response, or itssuggested_skillis no longer/harness-config:unhobble.unhobble-experimentif no rule uses it).toolDenialKindvalue shows that value as a "new" change incollect.py driftoutput.--catalog.Constraints and gotchas
skill_labelrenders a missing skill as(not installed)(sweep.py:99). Since fix: gate cross-plugin routing on enabled, not installed #5985, routes in this repo are gated on "enabled", and/harness-config:*and/harness-ops:*belong to plugins that shipdefaultEnabled: false; consider "(not enabled)" while touching this./harness-ops:audit-performanceread recorded hook durations; thehooks.stop_msroute is handled in the metric-accuracy child, not here.Context
Source: local handoff item 20261003-070338-session-flow-audit-sessions-accuracy-and-coverage.md (findings E5, R1, I4, Q2). Related: #5985, #6033, #4261 (guard rejections costing retry turns).