diff --git a/.claude/rules/skill-bodies-state-current-rules.md b/.claude/rules/skill-bodies-state-current-rules.md index 443f53f751..ffac3247c9 100644 --- a/.claude/rules/skill-bodies-state-current-rules.md +++ b/.claude/rules/skill-bodies-state-current-rules.md @@ -1,5 +1,5 @@ --- -description: "Skill and agent bodies carry a four-part verification record for any volatile specific they restate, and name their successor in a `## Next` section; read before editing any skill body" +description: "Skill and agent bodies point at the live upstream source for any volatile specific instead of restating it, recorded as pointer, as-of date and recheck trigger, and name their successor in a `## Next` section; read before editing any skill body" paths: - "plugins/*/skills/**" - "plugins/*/agents/**" @@ -7,12 +7,13 @@ paths: # Skill bodies state current rules -A pointer to an external upstream source (an official doc page, an upstream issue) is required -when a skill or agent body restates a volatile specific it cannot defer to at read time, recorded -as the four-part verification record the -[upstream-drift convention](../../docs/conventions/upstream-drift/README.md) defines: claim, -basis, as-of date, recheck trigger. A dated verification with a trigger is the correct form; an -undated claim is the defect. +A skill or agent body that depends on a volatile upstream specific (an official doc page, an +upstream issue) never restates it, quoted or paraphrased. It states our decision in our own words +and records where to read the specific live, in the links-only record the +[upstream-drift convention](../../docs/conventions/upstream-drift/README.md#required-parts) +defines: pointer to the exact section, as-of date, recheck trigger. A body that needs the specific +at run time fetches it from the pointer. Restated upstream text is the defect, and so is a pointer +with no as-of date or no trigger. ## Successor sections diff --git a/AGENTS.md b/AGENTS.md index 20f6e399ba..fe2f97f755 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -46,7 +46,7 @@ and its content is not already in context, read the file directly. |---|---|---| | `.claude/rules/mod-authoring.md` | `plugins/*/hooks/**, plugins/*/types/**` | Mods stay deferred under ADR 0035: no plugin gains a `modules` key until its five go criteria pass; when they do, load the built-in `plugin-authoring` skill and the upstream mods docs first | | `.claude/rules/ruff-pin.md` | `**/*.py` | Python linting runs through the pinned ruff wrapper, never a bare ruff on PATH | -| `.claude/rules/skill-bodies-state-current-rules.md` | `plugins/*/skills/**, plugins/*/agents/**` | Skill and agent bodies carry a four-part verification record for any volatile specific they restate, and name their successor in a `## Next` section; read before editing any skill body | +| `.claude/rules/skill-bodies-state-current-rules.md` | `plugins/*/skills/**, plugins/*/agents/**` | Skill and agent bodies point at the live upstream source for any volatile specific instead of restating it, recorded as pointer, as-of date and recheck trigger, and name their successor in a `## Next` section; read before editing any skill body | | `plugins/attribution/skills/audit/AGENTS.md` | `plugins/attribution/skills/audit/**` | Editing the attribution audit skill: contributor conventions | | `plugins/autonomy/AGENTS.md` | `plugins/autonomy/**` | autonomy plugin: contributor conventions | | `plugins/machine-health/skills/audit/AGENTS.md` | `plugins/machine-health/skills/audit/**` | machine-health audit skill: contributor conventions | diff --git a/docs/adr/0038-restore-the-claude-review-lanes-on-every-push.md b/docs/adr/0038-restore-the-claude-review-lanes-on-every-push.md index 832803b533..b79a65ac4c 100644 --- a/docs/adr/0038-restore-the-claude-review-lanes-on-every-push.md +++ b/docs/adr/0038-restore-the-claude-review-lanes-on-every-push.md @@ -14,7 +14,7 @@ not mean a review blocks a merge. This rule is the operator's. Anthropic does not document review on every pull request as a requirement. Its launch post says it runs Code Review "on nearly every PR at Anthropic" -(, fetched 2026-09-24); the product docs set no default +(correlate with , fetched 2026-09-24); the product docs set no default trigger, leaving an Owner to pick once, every push, or manual per repository (, fetched 2026-09-24); and Boris Cherny's Steps of AI Adoption says "Automated code review and security review are on by default" diff --git a/docs/conventions/detector-findings/CHANGELOG.md b/docs/conventions/detector-findings/CHANGELOG.md index d8f1455506..39442421aa 100644 --- a/docs/conventions/detector-findings/CHANGELOG.md +++ b/docs/conventions/detector-findings/CHANGELOG.md @@ -4,6 +4,14 @@ Notable changes to the detector-findings contract (SemVer). Changing a producer- the coexistence obligations, or an enforceability verdict is a major bump; additive guidance or a new adopter row is a minor bump; docs-only clarification is a patch. +## [3.6.2] - 2026-10-01 + +**Patch, docs-only.** The `attribution/audit/rule-stamp-expired` and +`attribution/audit/rule-restated-upstream-fact` crosswalk rows and the +`harness-config/audit-instructions/rule-trigger-less-stamp` row describe the stamped record by its +current shape (pointer, as-of date, observable recheck trigger beside the decision) instead of the +claim, basis, as-of and trigger shape. No producer-owned field's rule, tier, coexistence obligation or enforceability verdict moves. + ## [3.6.1] - 2026-10-01 **Patch, docs-only.** The rationale of `mutation-testing/audit/rule-survivor-productive` and diff --git a/docs/conventions/detector-findings/README.md b/docs/conventions/detector-findings/README.md index 7068309560..d18064b97e 100644 --- a/docs/conventions/detector-findings/README.md +++ b/docs/conventions/detector-findings/README.md @@ -247,14 +247,14 @@ side. | harness-config/audit-instructions/rule-blanket-tool-default | A blanket tool default ("default to using/running/calling", "if in doubt, use", "always use", "use even when") on an instruction line, under the identical body-scope fences and counted declines as the emphasis rule above. Mechanical phrase-list selection, case-folding, no withholding verdict. | The same walk, on the same mechanism and the same official source, which is why the two rules share a tier: a blanket default is the second arm of one defect: prompting written against undertriggering that no longer exists. CRITICAL fails identically (the phrasing determines no wrong result). IMPORTANT's degradation limb matches with the same nameable trigger, sharpened by the guide stating the consequence outright: "Instructions like 'If in doubt, use [tool]' will cause overtriggering". So the cost is a named behavioral one, not a register preference, and SUGGESTION is never reached. | IMPORTANT | No, contained to `Location`, but the repair replaces the blanket with the targeted condition it stood in for; recovering that condition is judgment, and the instruction itself is kept, never deleted | | harness-config/audit-instructions/rule-description-restatement | An H2 section whose every content unit is recoverable from the file's own `description` (the capability sentence; Use-when / Not-for is stripped before comparison), **body-scoped**: the scanner never points at frontmatter, and the writer independently declines any body row quoting a `'trigger phrase'` from the file's own `description` / `when_to_use`. Selection is mechanical token-containment with a length floor, no withholding verdict; a section with one unique content token is not a finding (threshold: wholly recoverable; the fired shape travels in the `Finding` cell). | CRITICAL fails every limb: restated prose computes nothing, so no concrete input, caller, or subsequent otherwise-correct change is made to produce a wrong, unsafe, or absent result. The description is already in context, and the body copy is a second load of the same bytes. IMPORTANT's degradation-with-a-named-trigger limb then matches: the trigger is the next session that pays the listing `description` *and* the body restatement for the same fact, spending tokens on content the model already has loaded. SUGGESTION is never reached. Keeping both copies is not a working alternative, it is the defect. | IMPORTANT | No, contained to `Location`, but the repair is a **cut of the body only**: the description, `when_to_use`, and every quoted trigger phrase must survive verbatim, because check-skill.sh check 3 hard-FAILs a dropped trigger phrase. Deciding that a near-match is still wholly recoverable is a rewrite judgment the relay surfaces. | | harness-config/audit-instructions/rule-sibling-restatement | An H2 section whose every content unit is recoverable from a sibling H2 section of the same file, under the identical body-scope fences and counted declines as the description-restatement rule. Footer headings (`Cross-references`, `Sources`, `History`, `External authority`, `Recheck triggers`) are sources for the comparison and are never themselves a finding. | The same walk on the same mechanism, a second load of bytes already in the file, which is why the two I29 rules share a tier. CRITICAL fails identically. IMPORTANT's degradation limb matches with the same nameable trigger, sharpened by the sibling being three lines above the copy (the #3122 morning-brief shape): the reader pays twice for one fact inside one invocation. SUGGESTION is never reached. | IMPORTANT | No, contained to `Location`, and the same body-only cut: the source section stays; only the restating copy is removed. | -| harness-config/audit-instructions/rule-trigger-less-stamp | A claim carrying an as-of date ("verified 2026-07-13", "as of 2.1.240") with no observable event that would re-open it, selected by a model lane and emitted through `emit-findings.sh --from-lane` (threshold: three of the four record parts present, the recheck trigger absent; the fired shape travels in the `Finding` cell). The catalog tiers it `mechanical`, but no scanner seeds it, so selection is the lane's, and **every withholding boundary requires its evidence PRESENT**: a trigger living in an owner record requires the site to name that record, the history exemption requires the file to be a CHANGELOG or an ADR, and date-as-data requires the date to sit in a data position such as a release table. A candidate showing none of them emits. Body-scoped under the identical fences and counted declines as the rules above, and `Confidence` is omitted because a judgment selected the row. | CRITICAL fails every limb: a stamp missing its trigger makes no input, caller, or subsequent otherwise-correct change produce a wrong, unsafe, or absent result. IMPORTANT's stated-rule limb then matches directly: this fleet adopted the four-part verification record (claim, basis, as-of date, recheck trigger) in writing, in the upstream-drift convention and `.claude/rules/skill-bodies-state-current-rules.md`, so a trigger-less stamp violates a rule the fleet states rather than one phrasing among several. This is the walk `attribution/audit/rule-trigger-less-stamp` makes on the same record. SUGGESTION is never reached. | IMPORTANT | No, contained to `Location`, but writing the missing trigger is a judgment about which observable event obliges re-derivation, and the stamp itself is kept | +| harness-config/audit-instructions/rule-trigger-less-stamp | A claim carrying an as-of date ("verified 2026-07-13", "as of 2.1.240") with no observable event that would re-open it, selected by a model lane and emitted through `emit-findings.sh --from-lane` (threshold: three of the four record parts present, the recheck trigger absent; the fired shape travels in the `Finding` cell). The catalog tiers it `mechanical`, but no scanner seeds it, so selection is the lane's, and **every withholding boundary requires its evidence PRESENT**: a trigger living in an owner record requires the site to name that record, the history exemption requires the file to be a CHANGELOG or an ADR, and date-as-data requires the date to sit in a data position such as a release table. A candidate showing none of them emits. Body-scoped under the identical fences and counted declines as the rules above, and `Confidence` is omitted because a judgment selected the row. | CRITICAL fails every limb: a stamp missing its trigger makes no input, caller, or subsequent otherwise-correct change produce a wrong, unsafe, or absent result. IMPORTANT's stated-rule limb then matches directly: this fleet adopted the upstream-drift record (a pointer, an as-of date, and an observable recheck trigger beside the decision) in writing, in the upstream-drift convention and `.claude/rules/skill-bodies-state-current-rules.md`, so a trigger-less stamp violates a rule the fleet states rather than one phrasing among several. This is the walk `attribution/audit/rule-trigger-less-stamp` makes on the same record. SUGGESTION is never reached. | IMPORTANT | No, contained to `Location`, but writing the missing trigger is a judgment about which observable event obliges re-derivation, and the stamp itself is kept | | harness-config/audit-instructions/rule-migration-relative-phrasing | Migration-relative phrasing ("now works differently", "no longer", "instead of the old", "since the change") in a file on the surfaces [`criteria.md`](../../../plugins/harness-config/skills/audit-instructions/reference/criteria.md) `### I31` defines, describing a diff against a version the reader never saw, selected by a model lane and emitted through `--from-lane` (the fired shape travels in the `Finding` cell). Behavioral and unseeded, so selection is a judgment, and **every withholding boundary requires its evidence PRESENT**: the structural-contrast exemption requires the sentence to name two current mechanisms, the history exemption requires a CHANGELOG or an ADR, and the four-part-record exemption requires all four parts with a trigger naming the prior state. A candidate showing none of them emits. A row outside the surfaces `### I31` defines is declined by the writer (`reason=outside-rule-surfaces`) and counted: that is where the remedy is offered, not a verdict on the text. Body-scoped under the identical fences; `Confidence` omitted. | CRITICAL fails every limb: the phrasing computes nothing, so no input, caller, or subsequent otherwise-correct change is made to produce a wrong, unsafe, or absent result. IMPORTANT's degradation-with-a-named-trigger limb matches: the trigger is the first session that loads the spoke, where "no longer counts" states only the negation of a prior state the reader never loaded, so the current rule must be inferred from half a diff. That is a named cost, not a register preference, so SUGGESTION is never reached. | IMPORTANT | No, the remediation can lie outside `Location`'s file: the sentence is restated at `Location`, and any history worth keeping moves to the owning plugin's `CHANGELOG.md` or an ADR, which every emitted row names in `Action` | | harness-config/audit-instructions/rule-route-to-absent-skill | A `/:` or `:` reference in routing text whose target has no `plugins//skills//SKILL.md`, or a routing sentence naming a plugin where a skill is required, selected by a model lane and emitted through `--from-lane` (the named target travels in the `Finding` cell). The rule has two arms, from [`criteria.md`](../../../plugins/harness-config/skills/audit-instructions/reference/criteria.md) `### I32`: the marketplace arm (`Location` under `plugins/`), which is the reference described above, and the user and project arm, which `### I32` defines. The tier follows the arm. Unseeded, so selection is the lane's, and **every withholding boundary requires its evidence PRESENT**: the bundled-skill exemption requires the reference to be marked bundled, the capability-by-class exemption requires the sentence to carry no binding, and the example exemption requires a fence around it. A candidate showing none of them emits. **Body-scoped**: the catalog row reaches descriptions and frontmatter routing clauses, but the relay is body-scoped, so a frontmatter-located row is declined by the writer (`reason=frontmatter`) and counted, never silently dropped, and stays in the human report (at least one measured row sits on a `SKILL.md` description line). `Confidence` omitted. | **Marketplace arm.** CRITICAL's none-at-all limb matches, and the input is nameable: a request matching the routing clause, in a session that follows the route, invokes a skill that does not exist, so the routed path produces no result at all. Unlike an emphasis or restatement row this is not a shift in how likely a behavior is: every session that follows the route reaches the same absent target. First match wins there, so the IMPORTANT and SUGGESTION tests are not reached. **User and project arm.** CRITICAL's limb is not established, for the reason `### I32` gives for this arm's `warning` severity: one session's skill listing is not every session's, so a skill absent from the listing the run saw can be present where the surface is read, and the evidence in hand names no session whose route reaches nothing. `### I32` sets this arm at `warning`, the catalog severity of `### I31` too, whose row is IMPORTANT; `criteria.md` is the authority for the tier, and this row states no argument beyond the reason it gives. | CRITICAL on the marketplace arm (`Location` under `plugins/`), IMPORTANT on the user and project arm | No, contained to `Location`, but choosing which existing skill the route means, or rephrasing it by capability class, is a judgment about intent; the routing sentence is kept and repointed | | harness-config/audit-instructions/rule-spoke-self-description | The opener of a spoke on the surfaces [`criteria.md`](../../../plugins/harness-config/skills/audit-instructions/reference/criteria.md) `### I33` defines, describing its own role or loading ("this file is read by step 3", "the hub links here") instead of stating its content, one finding per spoke keyed on the opener sentence's excerpt anchor, selected by a model lane and emitted through `--from-lane` (the fired shape travels in the `Finding` cell). Behavioral and unseeded, so selection is a judgment, and **every withholding boundary requires its evidence PRESENT**: the scope-note exemption requires one line bounding the file's subject, the index exemption requires the text to be the hub's own index table, and the frontmatter exemption requires a frontmatter position, which the writer also re-fences. A candidate showing none of them emits. A row outside the surfaces `### I33` defines is declined (`reason=outside-rule-surfaces`) and counted. `Confidence` omitted. | CRITICAL fails every limb: a self-description computes nothing and routes nothing. IMPORTANT fails: the catalog row is house guidance that no fleet rule adopts in writing, and its cost has no nameable degradation trigger, since the content below the opener loads and is followed either way. SUGGESTION holds: a spoke with and without the self-description both work, and the finding is a preference for spending the opener on content, which matches the catalog row's `info` severity. | SUGGESTION | No, the remediation can lie outside `Location`'s file: the opener is deleted at `Location`, and when the hub's index row does not already carry the loading condition it is added there, the hub every emitted row names in `Action` | | attribution/audit/rule-verbatim-copy | A fingerprint-confirmed matched span (fired values: containment, span words, source URL, identity check). Deterministic confirmation over a resolved source, no withholding verdict; the copy-class judgment verdicts around it (`source-fetched-similar`, `llm-suspected`, split rubric outcomes) have no rows because they never reach the relay (a restated-fact finding is a different class and reaches it only under its own rule, `rule-restated-upstream-fact` below), and every one of them falls toward a report-only human surface rather than toward silence. | CRITICAL fails every limb: copied prose computes nothing, so there is no input, caller, or subsequent otherwise-correct change the copy makes produce a wrong result, an unsafe one, or none at all. IMPORTANT matches twice over: the stated-rule limb (the org standard `documentation-and-citations.md` says prefer citing and fetching at read time over storing a snapshot, so a retained copy violates a rule the org already adopted in writing) and the degradation limb with a named trigger (the upstream page's next content change strands the local copy; the first reader trusting the stale copy acts on drifted facts under this repo's authority). SUGGESTION is never reached. | IMPORTANT | No, remediated by `/attribution:audit fix`. Choosing the disposition, clearing the semantic-diff guard, and verifying pointer liveness before writing are producer-owned | -| attribution/audit/rule-stamp-expired | A four-part record whose as-of date exceeds the configured window (fired values: date, window, days over). Deterministic date arithmetic, no withholding verdict; a stamp whose date form does not parse is declined with the form named and counted on the run's own surface rather than folded into the clean count. | CRITICAL fails on the same walk as the copy rule: a lapsed stamp computes nothing, so no input, caller, or subsequent otherwise-correct change is made to produce a wrong result by it. IMPORTANT's degradation limb matches with a named trigger: the record's currency ceiling has lapsed, and the first reader acting on the stamped claim without the re-fetch the convention requires acts on an assertion nobody has re-derived. Because that limb matches, SUGGESTION is never reached. | IMPORTANT | No, the repair is re-deriving the record against its live basis and restamping, or replacing the restatement with a pointer, and which of the two applies is a judgment the relay surfaces rather than applies | +| attribution/audit/rule-stamp-expired | A stamped record (pointer, as-of date, observable recheck trigger) whose as-of date exceeds the configured window (fired values: date, window, days over). Deterministic date arithmetic, no withholding verdict; a stamp whose date form does not parse is declined with the form named and counted on the run's own surface rather than folded into the clean count. | CRITICAL fails on the same walk as the copy rule: a lapsed stamp computes nothing, so no input, caller, or subsequent otherwise-correct change is made to produce a wrong result by it. IMPORTANT's degradation limb matches with a named trigger: the record's currency ceiling has lapsed, and the first reader acting on the stamped claim without the re-fetch the convention requires acts on an assertion nobody has re-derived. Because that limb matches, SUGGESTION is never reached. | IMPORTANT | No, the repair is re-deriving the record against its pointer and restamping, or replacing the restatement with a pointer, and which of the two applies is a judgment the relay surfaces rather than applies | | attribution/audit/rule-trigger-less-stamp | Repo-override only: a dated stamp whose surface states no recheck trigger. The portable default is OFF, and a repo turns it on by adopting the upstream-drift required parts; over the enabled corpus selection is mechanical, with no withholding verdict. | CRITICAL fails first: a stamp missing its trigger produces no wrong result, unsafe result, or absent one. IMPORTANT then matches on the stated-rule limb directly: the consuming repo that enables this check has adopted the upstream-drift required parts, and a trigger-less stamp violates part 4, a rule that repo adopted in writing. So SUGGESTION is never reached. The portable default stays off because the fleet's stamp forms are not uniformly greppable and a guessing gate converts signal to noise; that is a selection decision, and it is stated here so it is not read as part of the tier argument. | IMPORTANT | No, writing the missing trigger is a judgment about which observable event obliges re-derivation (upstream-drift required part 4) | -| attribution/audit/rule-restated-upstream-fact | A passage restating a fact an external source owns (a constant, default, version pin, field list, or the semantics of a named external product) with neither a pointer at the point of use nor a whole four-part record. Selection is a model panel's judgment, so the relay requires a **declared outcome**: the finding's class is `restated-fact`, the panel graded it unanimous STANDS under the restated-fact rubric, and the refutation pass (a fresh-context adversary told to default to refute) returned SURVIVES. A split panel, a REFUTED finding, an UNKNOWN grade, and a restated-fact finding declaring no such outcome are not emitted: they stay on the human report and are counted in `## Surfaces`. The declared outcome decides, never the evidence tier: a restated-fact finding maps by fixed rule to `source-fetched-similar`, `llm-suspected` or `not-found` and is never `fingerprint-confirmed`, and no tier name reaches the file (fired values: the panel size and the refutation outcome, plus the source URL when one was fetched). `Confidence` is omitted because a judgment selected the row. | CRITICAL fails every limb: restated prose computes nothing, so no input, caller, or subsequent otherwise-correct change is made to produce a wrong, unsafe, or absent result by the sentence existing. IMPORTANT's degradation limb matches with a named trigger and holds in any repository: the owner's next change to the restated default, limit, pin, or field list leaves the local statement wrong with nothing in the repository recording that it may be, and the first reader who acts on it acts on a fact nobody re-derived. IMPORTANT's stated-rule limb matches too where the rule is adopted: this repository's upstream-drift convention and `.claude/rules/skill-bodies-state-current-rules.md` require a pointer or a four-part record for a volatile specific a body restates. SUGGESTION is never reached. **The fail-safe direction is argued as the copy row argues it.** Admission test 2 guards against an unresolved judgment silently withholding a finding, and this rule's selection does not do that: every non-emitting outcome (a carve-out, a split, a REFUTED, an UNKNOWN) stays on the human report and is counted in `## Surfaces`, so nothing falls toward silence. What separates it from the deterministic rows is that the relay is the narrower surface: the refutation pass defaults to refute and treats an open question as REFUTED, so what reaches the relay is what an adversary tried to break and could not. A carve-out declines a candidate only on evidence the rubric requires to be present (all four parts of a conforming record, or the distilling surface's own attribution line quoted), so a carve-out whose evidence is absent leaves the candidate with the panel. | IMPORTANT | No, report-only: neither `/attribution:audit fix` nor `sweep` reaches this class, so the producer names no fix. Which disposition applies (a pointer at the point of use, or a four-part record with a real observable recheck trigger) turns on whether the surface must work offline and on which event obliges re-derivation, both judgments the relay surfaces rather than applies | +| attribution/audit/rule-restated-upstream-fact | A passage restating a fact an external source owns (a constant, default, version pin, field list, or the semantics of a named external product) with neither a pointer at the point of use nor a whole stamped record (pointer, as-of date, observable recheck trigger). Selection is a model panel's judgment, so the relay requires a **declared outcome**: the finding's class is `restated-fact`, the panel graded it unanimous STANDS under the restated-fact rubric, and the refutation pass (a fresh-context adversary told to default to refute) returned SURVIVES. A split panel, a REFUTED finding, an UNKNOWN grade, and a restated-fact finding declaring no such outcome are not emitted: they stay on the human report and are counted in `## Surfaces`. The declared outcome decides, never the evidence tier: a restated-fact finding maps by fixed rule to `source-fetched-similar`, `llm-suspected` or `not-found` and is never `fingerprint-confirmed`, and no tier name reaches the file (fired values: the panel size and the refutation outcome, plus the source URL when one was fetched). `Confidence` is omitted because a judgment selected the row. | CRITICAL fails every limb: restated prose computes nothing, so no input, caller, or subsequent otherwise-correct change is made to produce a wrong, unsafe, or absent result by the sentence existing. IMPORTANT's degradation limb matches with a named trigger and holds in any repository: the owner's next change to the restated default, limit, pin, or field list leaves the local statement wrong with nothing in the repository recording that it may be, and the first reader who acts on it acts on a fact nobody re-derived. IMPORTANT's stated-rule limb matches too where the rule is adopted: this repository's upstream-drift convention and `.claude/rules/skill-bodies-state-current-rules.md` require a pointer or a whole stamped record for a volatile specific a body depends on. SUGGESTION is never reached. **The fail-safe direction is argued as the copy row argues it.** Admission test 2 guards against an unresolved judgment silently withholding a finding, and this rule's selection does not do that: every non-emitting outcome (a carve-out, a split, a REFUTED, an UNKNOWN) stays on the human report and is counted in `## Surfaces`, so nothing falls toward silence. What separates it from the deterministic rows is that the relay is the narrower surface: the refutation pass defaults to refute and treats an open question as REFUTED, so what reaches the relay is what an adversary tried to break and could not. A carve-out declines a candidate only on evidence the rubric requires to be present (every part of a conforming record, or the distilling surface's own attribution line quoted), so a carve-out whose evidence is absent leaves the candidate with the panel. | IMPORTANT | No, report-only: neither `/attribution:audit fix` nor `sweep` reaches this class, so the producer names no fix. Which disposition applies (a pointer at the point of use, or a whole stamped record with a real observable recheck trigger) turns on whether the surface must work offline and on which event obliges re-derivation, both judgments the relay surfaces rather than applies | | docs-hygiene/audit-noise/rule-negation-without-positive | An **imperative** prohibition (`never`, `do not`, `don't`, `avoid`, `must not`, `should not`) opening the sentence, with no positive alternative stated in that sentence. Soft-wrapped sentences are accumulated across paragraph lines before classification; a finding is attributed to the first physical line of the triggering sentence. The imperative-only gate is what keeps the rule usable, and its effect is measured: without it the rule fired 1053 times on an 85-file sample (99% of all findings). Descriptive prose, an already-paired mid-sentence cue, and a table row are all out of scope. **Body-scoped**: `detect.sh` never leaves frontmatter, fenced code, exempt sections or opt-out-marked content, and the writer independently re-fences frontmatter and declines any body line quoting a `'trigger phrase'` that appears in the file's own `description` / `when_to_use`. Selection is per SENTENCE, on the backtick-unwrapped accumulated paragraph, and case-folded (threshold: any occurrence; the fired prohibition travels in the `Finding` cell). The other eight shapes stay line-scoped. **Decline evidence, three classes, each counted per shape in `## Surfaces`:** a body row whose line number falls inside frontmatter the writer re-fenced itself (`reason=frontmatter`); a body row quoting a `'trigger phrase'` present in the file's own `description` or `when_to_use` (`reason=quoted-trigger-phrase`); and a candidate the skill's model judgment lane dismissed on one of the grounds its `SKILL.md` enumerates under "Dismissal grounds the judgment pass may use", which is applied before the writer sees the row (`reason=judgment-lane-dismissal`). The first two are recomputed by the writer over its own input rather than trusted from the caller; the third is the model's and is why this producer is not fully mechanical end to end, only in its scanner. | CRITICAL fails every limb: instruction prose computes nothing, so no concrete input, caller, or subsequent otherwise-correct change is made to produce a wrong, unsafe, or absent result. A prohibition without its positive leaves the target under-specified rather than determined wrong. IMPORTANT's **stated-rule** limb then matches directly, and does so without needing the degradation limb: `docs-hygiene:write-for-agents` "Prompt the positive" is a rule this fleet already adopted in writing, "Write what to do, not what to avoid", so a bare prohibition is a surface violating a stated rule rather than one phrasing among several that all work. Official guidance supplies the same mechanism (*"Do not use markdown"* → *"Your response should be composed of smoothly flowing prose paragraphs"*). SUGGESTION's catch-all is therefore never reached. | IMPORTANT | No, contained to `Location`, but the repair rewrites to the positive target the prohibition implies, and recovering that target is a rewrite judgment: the constraint must survive while its framing changes. Matches the disposition of both `audit-instructions` rules for the same reason | | docs-hygiene/audit-noise/rule-negation-hard-guardrail | **Non-emitting.** A prohibition whose sentence also carries a safety-critical marker (`secret`, `credential`, `token`, `password`, `api key`, `force-push`, `--force`, `rm -rf`, `destructive`, `irreversible`, `data loss`, `production`, `security`, `vulnerab`, `rewrite history`), a hard guardrail whose constraint a positive form cannot carry, where the prohibition IS the correct shape. | **Boundary ground, not a tier test**: the contract's "Findings that never reach a relay". This disposition is argued from the Boundary because a tier test can only ever return a tier, and the claim here is that the candidate is not a defect at all: the sibling `write-for-agents` rule itself preserves the negation "when the positive form genuinely loses the constraint". Reaching for a tier test to justify the non-emission would look argued while arguing nothing. **Fail-safe direction:** the carve-out requires its marker to be PRESENT on the sentence, so absence of that evidence selects the emitting rule above. An unresolved judgment can never withhold. The same holds for the paired-positive boundary (`instead`, `rather than`, `prefer`, `in place of`, `in favor of`) and the worked-example boundary (a `->` / `→` demonstration): each requires positive evidence, and both were checked rather than only the one easier to argue. | *(never reaches the relay)* | Not applicable, no row, and suppressed from the human report too, so one candidate carries one disposition on every surface this producer emits to | diff --git a/docs/conventions/hook-config-delivery/README.md b/docs/conventions/hook-config-delivery/README.md index d148f7ff17..6571844b90 100644 --- a/docs/conventions/hook-config-delivery/README.md +++ b/docs/conventions/hook-config-delivery/README.md @@ -26,28 +26,30 @@ with each plugin's own docs. ## Verified upstream behavior (version-pinned) -Each row carries its own verified version and date in the **Verified** column. Facts 1-8 were -verified on **Claude Code 2.1.218**: doc-stated facts re-fetched from the live official docs on -2026-07-24, behavioral facts proven by a controlled fresh-session probe (isolated -`claude -p --plugin-dir` runs with positive controls) on 2026-07-23. Facts 9-12 were measured on -**Claude Code 2.1.283** by a sandbox probe on 2026-09-27 (fixture `CLAUDE_CONFIG_DIR`, `HOME` and -`USERPROFILE` under a scratch directory, positive control of an empty plugin list before any write). -Existing state is not evidence of its own correctness: **recheck these facts** (triggers at the end -of this doc) before extending the matrix or relying on a row in new work. - -| # | Fact | Basis | Verified | +Each row carries its own Claude Code version and date in the **As of** column. A doc-stated row +records what we rely on, in our words, and points at the section that states it; a probed row +records what our own probe observed. Rows 1-8 date from **Claude Code 2.1.218**: doc-stated rows +re-read from the live official docs on 2026-07-24, behavioral rows proven by a controlled +fresh-session probe (isolated `claude -p --plugin-dir` runs with positive controls) on 2026-07-23. +Rows 9-12 were measured on **Claude Code 2.1.283** by a sandbox probe on 2026-09-27 (fixture +`CLAUDE_CONFIG_DIR`, `HOME` and `USERPROFILE` under a scratch directory, positive control of an +empty plugin list before any write). Existing state is not evidence of its own correctness: +**recheck these facts** (triggers at the end of this doc) before extending the matrix or relying on +a row in new work. + +| # | What we rely on | Pointer | As of | |---|---|---|---| -| 1 | Plugin `hooks.json` hooks receive `${user_config.KEY}` in **exec form only**, substituted into `command` and each `args` element as a plain string; a shell-form command referencing it fails with an error instead of running (since 2.1.207) | doc-stated ([hooks](https://code.claude.com/docs/en/hooks)) | 2.1.218, 2026-07-24 | -| 2 | Configured values are exported to hook processes as `CLAUDE_PLUGIN_OPTION_` (key uppercased) | doc-stated ([plugins-reference](https://code.claude.com/docs/en/plugins-reference#user-configuration)); scope narrowed by fact 4 | 2.1.218, 2026-07-24 | -| 3 | The declared `default` field ("Value used when the user provides nothing") is in the schema but **implemented for neither argv substitution nor env export**. An unset-but-defaulted `${user_config.*}` argv token **silently drops the entire hook entry**. It is not passed literally and not empty-substituted; the same unset key exports **no** env var | proven (probe T1); upstream [#46477](https://github.com/anthropics/claude-code/issues/46477) closed not-planned, [#39455](https://github.com/anthropics/claude-code/issues/39455) open, [#39827](https://github.com/anthropics/claude-code/issues/39827) closed not-planned; undocumented | 2.1.218, 2026-07-23 | +| 1 | `${user_config.KEY}` reaches a plugin `hooks.json` hook only in **exec form**, as a plain string in `command` and each `args` element; we never write it in a shell-form command | doc-stated ([Exec form and shell form](https://code.claude.com/docs/en/hooks#exec-form-and-shell-form)) | 2.1.218, 2026-07-24 | +| 2 | A hook process reads a configured value from `CLAUDE_PLUGIN_OPTION_` (key uppercased) | doc-stated ([User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration)); scope narrowed by fact 4 | 2.1.218, 2026-07-24 | +| 3 | The declared `default` field is in the schema but **implemented for neither argv substitution nor env export**. An unset-but-defaulted `${user_config.*}` argv token **silently drops the entire hook entry**. It is not passed literally and not empty-substituted; the same unset key exports **no** env var | proven (probe T1); upstream [#46477](https://github.com/anthropics/claude-code/issues/46477) closed not-planned, [#39455](https://github.com/anthropics/claude-code/issues/39455) open, [#39827](https://github.com/anthropics/claude-code/issues/39827) closed not-planned; undocumented | 2.1.218, 2026-07-23 | | 4 | Tamper split on the env channel: for a **configured** key, harness injection overwrites a repo `.claude/settings.json` `env` block (injection wins); for an **unconfigured** key nothing is injected and the repo's `env` block freely populates `CLAUDE_PLUGIN_OPTION_`. Env carries no provenance, so a hook cannot tell the two apart | proven (probe T2/T2b); undocumented | 2.1.218, 2026-07-23 | -| 5 | `pluginConfigs` is written to user settings and read back from **user settings, the `--settings` flag, and managed settings only**; entries in a project's `.claude/settings.json` / `.claude/settings.local.json` are ignored (since 2.1.207) | doc-stated ([plugins-reference](https://code.claude.com/docs/en/plugins-reference#user-configuration)) | 2.1.218, 2026-07-24 | +| 5 | We treat `pluginConfigs` as living in **user settings, the `--settings` flag, and managed settings only**, and an entry in a project's `.claude/settings.json` / `.claude/settings.local.json` as ignored | doc-stated ([User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration)) | 2.1.218, 2026-07-24 | | 6 | Skill- and agent-frontmatter hooks receive **neither** the argv substitution nor `CLAUDE_PLUGIN_OPTION_*` | evidence-strong (probe + field repro); CC docs silent | 2.1.218, 2026-07-23 | -| 7 | Skill/agent **body** `${user_config.KEY}` substitutes into model-visible content, non-sensitive values only | doc-stated ([plugins-reference](https://code.claude.com/docs/en/plugins-reference#user-configuration)) | 2.1.218, 2026-07-24 | -| 8 | Sensitive values are stored in the OS keychain (or `~/.claude/.credentials.json`), never in `settings.json` | doc-stated ([plugins-reference](https://code.claude.com/docs/en/plugins-reference#user-configuration)) | 2.1.218, 2026-07-24 | +| 7 | We use skill/agent **body** `${user_config.KEY}` only as model-visible content, and only for non-sensitive values | doc-stated ([User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration)) | 2.1.218, 2026-07-24 | +| 8 | We never expect a sensitive value in `settings.json`, so no settings reader can see one | doc-stated ([User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration)) | 2.1.218, 2026-07-24 | | 9 | `claude plugin install --config` writes `pluginConfigs` to **user settings whatever `-s` says**; `-s project` / `-s local` governs only the install record and the `enabledPlugins` entry in that scope's settings file. Exit 0, no warning about the scope. A consequence of fact 5, so no setup recipe may document `-s project --config` as a per-repo value | proven (sandbox probe, Claude Code 2.1.283, 2026-09-27) | 2.1.283, 2026-09-27 | -| 10 | `--config` validates at write time but **never fails the command**: a wrong-type boolean and an undeclared key are rejected with a warning and exit 0; a prose-enumerated string and a non-existent `directory` path are stored without a warning. A declared `options` fixed list is enforced (a value outside it is rejected with a warning, exit 0); a plugin declaring `options` cannot load on Claude Code before v2.1.271. So an in-consumer fallback is mandatory for any key without `options`, and a headless caller reads the CLI output, not the exit code | proven (sandbox probe, Claude Code 2.1.283, 2026-09-27); `options` doc-stated ([plugins-reference](https://code.claude.com/docs/en/plugins-reference), "Limit a field to fixed options", fetched 2026-09-27) and probed | 2.1.283, 2026-09-27 | -| 11 | There is **no CLI path to unset a key**: `claude plugin --help` lists `details`, `disable`, `enable`, `eval`, `help`, `init`, `install`, `list`, `marketplace`, `prune`, `tag`, `uninstall`, `update`, `validate` (no `config` or `configure`), and `--config KEY=` with an empty value is rejected ("Omit the flag to leave ... unset"). A key using the declare-no-default idiom (for example work-items `work_dispatch_concurrency_cap`) cannot be cleared from the CLI once set: clearing it means hand-editing user settings or uninstalling, which drops every option | proven (sandbox probe, Claude Code 2.1.283, 2026-09-27) | 2.1.283, 2026-09-27 | +| 10 | `--config` validates at write time but **never fails the command**: a wrong-type boolean and an undeclared key are rejected with a warning and exit 0; a prose-enumerated string and a non-existent `directory` path are stored without a warning. A declared `options` fixed list is enforced (a value outside it is rejected with a warning, exit 0); a plugin declaring `options` needs Claude Code v2.1.271 or later. So an in-consumer fallback is mandatory for any key without `options`, and a headless caller reads the CLI output, not the exit code | proven (sandbox probe, Claude Code 2.1.283, 2026-09-27); `options` doc-stated ([Limit a field to fixed options](https://code.claude.com/docs/en/plugins-reference#limit-a-field-to-fixed-options), read 2026-09-27) and probed | 2.1.283, 2026-09-27 | +| 11 | There is **no CLI path to unset a key**: `claude plugin --help` lists `details`, `disable`, `enable`, `eval`, `help`, `init`, `install`, `list`, `marketplace`, `prune`, `tag`, `uninstall`, `update`, `validate` (no `config` or `configure`), and `--config KEY=` with an empty value is rejected with a message to omit the flag instead. A key using the declare-no-default idiom (for example work-items `work_dispatch_concurrency_cap`) cannot be cleared from the CLI once set: clearing it means hand-editing user settings or uninstalling, which drops every option | proven (sandbox probe, Claude Code 2.1.283, 2026-09-27) | 2.1.283, 2026-09-27 | | 12 | Rerunning `install --config` on a plugin already installed at the same scope is a pure config write: the install record is byte-identical and only the named option keys change. Measured for a `string` option at `user` and `project` scope and for `boolean` and `directory` options at `local` scope | proven (sandbox probe, Claude Code 2.1.283, 2026-09-27); the reconfigure guidance built on it is owned by [plugin-reconfiguration's verified-version record](../plugin-reconfiguration/README.md#verified-version-record) | 2.1.283, 2026-09-27 | ## The channels @@ -139,8 +141,9 @@ unproven. A ratified D adoption (after the Open-gaps probe) is recorded in ## Open gaps - **G-required (channel D's premise).** That `required:true` forces a prompt and so removes the - unset case is inferred from the schema (`required`: "validation fails when the field is empty") - and upstream discussion, not doc-stated and not yet probed: the verification probe declared + unset case is inferred from the schema's `required` field (see + [User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration)) and + upstream discussion, not doc-stated and not yet probed: the verification probe declared optional keys only. Cheap to settle. Add a `required:true` key to the probe plugin and rerun the unset-key test. Until then D stays in the matrix as unproven and the CI gate has no allowlist entries. @@ -161,7 +164,7 @@ authority it points to. ## Recheck triggers -Stamp-and-trigger discipline: [upstream-drift](../upstream-drift/README.md). Recheck the facts +Record shape and trigger discipline: [upstream-drift](../upstream-drift/README.md). Recheck the facts table (and re-derive the decision rule) when any of these fires: - A Claude Code CHANGELOG entry touches `userConfig` substitution, the `default` field, or diff --git a/docs/conventions/hook-observability/README.md b/docs/conventions/hook-observability/README.md index 7eb626b1bb..32504d2e89 100644 --- a/docs/conventions/hook-observability/README.md +++ b/docs/conventions/hook-observability/README.md @@ -6,10 +6,9 @@ telemetry envelope. The [plugin philosophy](../../plugin-philosophy.md) owns the advisory-versus-blocking, fail-open-versus-closed. This doc owns which of the three surfaces a given situation uses and how each is shaped. -Grounded against the official Claude Code hooks reference -(, fetched 2026-08-10). Every field name, cap, and timing -claim below is sourced from that fetch, not from training-data recall, per this repo's own -research-verification discipline. +The rules below are this convention's decisions. Where one depends on hook behavior Claude Code +owns, it points at the section of the [hooks reference](https://code.claude.com/docs/en/hooks) to +read live, with the date it was checked and the event that sends someone back. ## The three surfaces @@ -26,8 +25,10 @@ A static field on a `hooks.json` **handler object**, sibling of `type`/`command` } ``` -Displayed as the UI spinner label while the hook process runs. **A hook script never emits this. -There is no runtime JSON output field by this name.** **Rollout status: near-complete.** As of +It labels the spinner while the hook runs. **A hook script never emits this: it is configuration, +not a runtime output field.** Pointer: for the field, see +. As of: 2026-10-01. Recheck trigger: that +table moves the field or a runtime output field takes the name. **Rollout status: near-complete.** As of 2026-07-23, 30 of the 31 wired `type: "command"` handlers across the fleet's 15 hook-bearing plugins declare `statusMessage`; the sole remaining holdout is `plugins/disk-hygiene/hooks/hooks.json`. Tracked against @@ -40,14 +41,16 @@ telemetry..."`), not a generic `"Running hook..."`. ### 2. `systemMessage`: user-visible, scoped by who can act on the content -A JSON output field (`hookSpecificOutput` sibling) that Claude Code reads on every exit code, exit 2 -included ("Claude Code still reads any valid JSON output on stdout", hooks reference, Exit code 2, -fetched 2026-09-27), 10,000-character cap (an overflow to a -file, not a truncation; see [Output caps](#output-caps-stated-by-the-reference)), shown to the -user immediately. Composed via `hook::emit_channels` / `hook::emit_skip_notice` -(`lib/hook-utils.sh`) alongside `additionalContext` in one JSON document. Claude Code parses a -hook's entire stdout as a single document, so a hook with both agent-channel content and a -pending notice must compose them there, never `printf` twice. +A JSON output field (`hookSpecificOutput` sibling) addressed to the user. We compose it on any exit +code, exit 2 included, and size it under the cap in +[Output caps](#output-caps-stated-by-the-reference). Pointer: for which exit codes still read JSON +output, see . As of: 2026-09-27. Recheck +trigger: that section changes whether JSON output is read on exit 2. Composed via +`hook::emit_channels` / `hook::emit_skip_notice` +(`lib/hook-utils.sh`) alongside `additionalContext` in one JSON document: a hook with both +agent-channel content and a pending notice composes them there and never prints two objects. +Pointer: for how stdout is parsed as JSON, see . +As of: 2026-10-01. Recheck trigger: that section changes how a multi-object stdout is read. **Scope, required for exactly one situation:** a missing runtime prerequisite (binary, config file, `jq`) causes the hook to silently no-op instead of performing its check. Doctrine @@ -65,12 +68,14 @@ edits a file the user is working in, on the strength of an unrelated tool call, diff. The harness's own signal for it is a generic "PostToolUse hook modified `` after your edit (likely a formatter)" line that names no hook and shows no change, quoted from an observed session, not from a docs page, and used here only as an illustration of the shape such a -notice takes. What the docs settle is the negative this rule actually rests on, verified against - (fetched 2026-08-10): the three documented output channels -carry no file-change or diff surface, so a benign reflow and a wrong dictionary rewrite arrive -identically. Recheck trigger: a Claude Code release that adds a file-change or diff surface to the -hook output schema, whether a fourth output field or such a payload on one of -[the three](#the-three-surfaces), which would make this rule's disclosure requirement redundant. +notice takes. The rule rests on a negative: we found no file-change or diff surface among the hook +output channels, so a benign reflow and a wrong dictionary rewrite arrive identically. + +- **Pointer**: for the hook output fields, see . +- **As of**: 2026-08-10 +- **Recheck trigger**: a Claude Code release that adds a file-change or diff surface to the hook + output schema, whether a fourth output field or such a payload on one of + [the three](#the-three-surfaces), which would make this rule's disclosure requirement redundant. The person whose file was changed is the only one who can judge whether the change was correct, so **the hook must name what it changed on the user channel**, not only the agent one: what @@ -82,10 +87,11 @@ count, or the disclosure becomes the noise problem it was meant to prevent. **Not required** for two situations that are already visible or already correctly agent-scoped: -- **Exit-2 blocking paths.** A `PreToolUse` hook that blocks a tool call via exit code 2 is - already user-visible through Claude Code's own permission-denial UI, and Claude reads its stderr - as the denial reason. Repeating the block reason on `systemMessage` would be redundant, not more - observable. The field is not discarded on a block (see above), so a blocking hook may still carry +- **Exit-2 blocking paths.** We treat a `PreToolUse` block via exit code 2 as already visible to + both the user and Claude, with its stderr as the reason, so repeating the block reason on + `systemMessage` would be redundant, not more observable. Pointer: for the blocking message, see + . As of: 2026-10-01. Recheck trigger: that + section changes what a `PreToolUse` block shows or which text becomes its reason. The field is not discarded on a block (see above), so a blocking hook may still carry one, but only for content that meets the carve-out below, never for the reason itself. - **Legitimate advisory findings *the model can act on*.** A hook that surfaces a finding to Claude for it to act on (e.g. a lint result, a suggested fix) belongs on `additionalContext` only. That @@ -112,17 +118,22 @@ count, or the disclosure becomes the noise problem it was meant to prevent. 3. the emission is keyed to a **state transition**, not to every invocation. **Delivery may never be asserted.** The model channel may state that a choice belongs to the - operator; it may **never** state that the operator has seen it. No documented behavior tells a hook - whether an operator is present. `systemMessage` is documented only as a message shown to the user, - and nothing upstream describes its behavior in non-interactive runs, so a delivery claim is a fact - the hook cannot know in *any* mode, not only headless ones. Emitting to an unread operator channel - is harmless; telling the model a human holds the choice when none does is not. - - **Honest limit.** The docs state that `additionalContext` is inserted into the conversation and - saved to the transcript, and say no such thing about `systemMessage`; that the latter stays out of - model context is *inferred from the asymmetry*, not stated. If that inference is ever falsified, - this carve-out collapses, since content forbidden to the model would reach it either way, and the - correct response is to drop the payload, not to re-route it. + operator; it may **never** state that the operator has seen it. We found no documented way for a + hook to learn whether an operator is present, so a delivery claim is a fact the hook cannot know + in *any* mode, not only headless ones. Emitting to an unread operator channel is harmless; + telling the model a human holds the choice when none does is not. + + **Synchronous hooks only.** The carve-out holds only where `systemMessage` stays off the model + channel, and the reference decides that per hook kind and per event. We admit it from a + synchronous hook on an event whose own section leaves the field on the user channel, and never + from an `async` hook. Where the field would reach the model, drop the payload; never re-route it. + + - **Pointer**: for the field's delivery, see + and the event's section under ; for + background hooks, see . + - **As of**: 2026-10-01 + - **Recheck trigger**: either section changes where `systemMessage` is delivered, or an event's + section changes how it treats the field. **Repeat-notice discipline.** A missing-prerequisite notice behind a broad matcher (every `Write|Edit`, every `Bash` call) must not repeat on every invocation. Use `hook::require_jq` @@ -134,9 +145,11 @@ share the parent's context and would otherwise never see why the hook skipped. A latch renews with a one-line notice every `HOOK_NOTICE_RENEW_EVERY` skips (default 8). A plugin README states this as "once per session and agent, renewed every eighth skip", never "once per session". The exception is a missing external binary: `hook::notice_once prerequisite` latches on the session alone and each renewal keeps the full notice with its install route, so the README states "once per session, renewed with the install route every eighth skip". -**Important exit-code caveat, grounded in the fresh fetch:** on exit 0, **stderr is never shown to -the user or the agent**, and only stdout JSON is parsed. A bare `echo "..." >&2; exit 0` skip is -**not visible**, regardless of intent. `scripts/check-silent-skips.sh` **still treats a bare +**Important exit-code caveat:** on exit 0, **we treat stderr as visible to neither the user nor +the agent**; only stdout JSON carries a notice. A bare `echo "..." >&2; exit 0` skip is **not +visible**, regardless of intent. Pointer: for where exit-0 stderr goes, see +. As of: 2026-08-10. Recheck trigger: that +section starts showing exit-0 stderr in the transcript or to the model. `scripts/check-silent-skips.sh` **still treats a bare stderr write as a sanctioned visibility signal as of this doc's introduction.** That is incorrect for the exit-0 skip shapes the gate inspects, and the gate does not yet enforce the rule this doc states. **Gate correction is pending**, scoped into the same fleet-adoption follow-up PR (against @@ -160,49 +173,37 @@ candidate; give it a helper call. ### Output caps stated by the reference -Re-read 2026-09-05 against by the rung-1 route in -[upstream-drift](../upstream-drift/README.md#the-rungs): a `curl` of the raw-markdown channel, -317,632 bytes, first heading `# Hooks reference`, slug listed in `llms.txt`, SHA-256 -`c30a50b8192dadf4e6ba016e451685f57a6d1d2c360d268887a9a94022d29f3e`. The channel claims above still -match the page. Three cap facts the page states and this doc did not carry are recorded below, each -as a four-part record (claim, basis, as-of date, recheck trigger). Line numbers are positions in -that fetch, given so a re-check can find the span; the quoted text is the basis. - -1. **Output over the cap overflows to a file; it is not truncated.** Basis, line 913: "Hook output - strings, including `additionalContext`, `systemMessage`, and plain stdout, are capped at 10,000 - characters. Output that exceeds this limit is saved to a file and replaced with a preview and - file path, the same way a large valid Bash result is handled". As of 2026-09-05. Recheck - trigger: a read-time re-fetch of the page finds the 10,000 figure or the save-to-file behavior - under its JSON-output section changed or gone. What it means for a hook: an over-cap disclosure - is not lost, but it stops being the inline account the content-mutation rule above requires, and - in write mode the file has already been rewritten by then. A hook that must stay inline caps - itself under the figure with a truncation that keeps its counts and says it truncated, leaving - headroom for JSON escaping. The adopting reference is `plugins/typos-format/hooks/typos-format.sh` - (4,000 for `systemMessage`, 8,000 for `additionalContext`). - -2. **The `additionalContext` cap is per value; there is no pool shared across hooks.** Basis, line - 993: "When several hooks return `additionalContext` for the same event, Claude receives all of - the values. If a value exceeds 10,000 characters, Claude Code writes the full text to a file in - the session directory and passes Claude the file path with a short preview instead". As of - 2026-09-05. Recheck trigger: the same re-fetch finds the "all of the values" sentence changed or - a shared budget stated for the field. What it means for a hook: size the value against 10,000 - and never against what other hooks on the same event emit. +Three cap decisions, each with the section that owns the figure. -3. **`classifierContext` carries its own 2,000-character cap, shared across hooks, and governs - neither channel above.** The field is a `PostToolUse` `hookSpecificOutput` member addressed to - the auto-mode classifier ("requires Claude Code v2.1.236 or later", line 1999). Basis, lines - 2019 to 2021: "Claude Code caps the notes for one tool call at 2,000 characters and truncates the - rest. The cap is shared across every hook that responds to that call"; "Claude Code ignores the - field in the response of a hook that runs in the background"; "the classifier's transcript omits - read-only lookups such as file reads and searches. Claude Code discards a note attached to one - of those calls". As of 2026-09-05. Recheck trigger: the re-fetch finds the 2,000 figure, the - sharing rule, or the event list for the field changed. Why this doc records it: a 2026-09-04 - peer review read the 2,000-character shared cap as the `additionalContext` cap and filed the - typos-format 8,000-character self-cap as a bug; the report was withdrawn on this reading of the - page, and this is where the next reader should find the answer. The field is not a fourth - surface for this convention: it reaches the classifier, never the user or the model, so it - changes nothing about which channel a fleet hook writes a notice to. No fleet hook emits it - today, and a `PreToolUse` guard cannot: the page lists it under `PostToolUse` only. +1. **This convention sizes every user- or agent-channel string under 10,000 characters.** A + disclosure the content-mutation rule above requires must reach the reader inline and whole, so + a hook that must stay inline caps itself under the figure with a truncation that keeps its + counts and says it truncated, leaving headroom for JSON escaping. What happens to a value over + the cap is the pointer's to state. The adopting reference is `plugins/typos-format/hooks/typos-format.sh` (4,000 for + `systemMessage`, 8,000 for `additionalContext`). + - **Pointer**: for the output cap, see . + - **As of**: 2026-09-05 + - **Recheck trigger**: that section changes the 10,000 figure or what happens over it. + +2. **Each `additionalContext` value is sized on its own.** We size the value against the cap above + and never against what other hooks on the same event emit. + - **Pointer**: for how several hooks' values are delivered, see + . + - **As of**: 2026-09-05 + - **Recheck trigger**: that section states a budget shared across hooks. + +3. **`classifierContext` is not a fourth surface, and its cap governs neither channel above.** It + is a `PostToolUse` field addressed to the auto-mode classifier, never the user or the model, so + it changes nothing about which channel a fleet hook writes a notice to. No fleet hook emits it + today, and a `PreToolUse` guard cannot. Why this doc records it: a 2026-09-04 peer review read + the classifier note's shared cap as the `additionalContext` cap and filed the typos-format + 8,000-character self-cap as a bug; the report was withdrawn on this reading of the page, and + this is where the next reader should find the answer. + - **Pointer**: for the field and its limits, see + . + - **As of**: 2026-09-05 + - **Recheck trigger**: that section changes the field's cap, its sharing rule, or the events + that accept it. ### 3. OTel-style telemetry envelope @@ -215,21 +216,52 @@ skipped-for-cause). A pure inapplicability short-circuit before any check logic type, excluded path, missing prerequisite) does not need one; see the Conformance section below for the precise rule and why. -**Why a local file sink, not a real OTel exporter.** Claude Code strips every `OTEL_*` exporter -environment variable from hook subprocesses it spawns -(), so a hook process -cannot emit real OpenTelemetry even if it tried. The file-sink envelope is the only telemetry -surface available to a hook; this is a grounded constraint, not an oversight. - -**Deferred: `prompt_id` correlation.** Hook input JSON carries a `prompt_id` field (Claude Code -v2.1.196+) that matches the `prompt.id` attribute on real OpenTelemetry events, which would let -external tooling correlate a hook's local envelope with the same turn's real OTel stream. Adding -it is a `hook-telemetry` schema change (`schema_version` 1.0 → 1.1) touching every producer's -`data_json` construction, out of scope for this doc's three-surface convention. +**Why a local file sink, not a real OTel exporter.** A hook process does not receive the session's +`OTEL_*` exporter configuration, so it cannot emit real OpenTelemetry. The file-sink envelope is +the only telemetry surface we give a hook; this is a constraint, not an oversight. Pointer: for +which subprocesses get no `OTEL_*` variables, see +. As of: 2026-10-01. +Recheck trigger: that section starts passing `OTEL_*` variables to hook subprocesses. + +**Deferred: `prompt_id` correlation.** Carrying the hook input's `prompt_id` in the envelope would +let external tooling correlate a hook's local envelope with the same turn's real OTel stream. +Adding it is a `hook-telemetry` schema change (`schema_version` 1.0 → 1.1) touching every +producer's `data_json` construction, out of scope for this doc's three-surface convention. +Pointer: for the field, see . As of: +2026-10-01. Recheck trigger: that table drops the field or its OpenTelemetry match. melodic-software/claude-code-plugins#930 is closed: the per-session event log (`harness-ops`, melodic-software/claude-code-plugins#3750) records `prompt_id` per event, and the envelope-spine promotion is tracked at melodic-software/claude-code-plugins#3758. +## Text a hook adds for the model: frequency and phrasing + +A hook's `additionalContext` on a tool event lands beside the tool result, where untrusted tool +output also arrives. We treat agent-channel text that repeats often or tells the model what to do +as a prompt-injection risk, both for the hook's own text and for a genuine user message that +arrives in the same place. Two rules follow. + +- **Frequency.** Emit agent-channel text only when it changes what the model does next: a finding, + a state transition, or a missing prerequisite under the repeat-notice latch above. A check that + ran clean says nothing on the agent channel; its outcome goes to the telemetry envelope. A fleet + hook adds no per-call status line, no reminder repeated on every tool call, and no countdown or + budget line after tool results. +- **Phrasing.** Write facts with their source, never orders. Name the hook, what it observed, and + where (`typos-format: 2 misspellings in docs/a.md:12`), and state a remedy as a fact about the + project (`markdownlint: README.md:40 is 131 characters; this repo wraps markdown at 100`). Hook + text claims no authority it lacks + (system, administrator, user) and never presents itself as a message from the user. + +This section is documentation only: it measures no hook's emission rate and moves no hook to a +different event. + +- **Pointer**: for where `additionalContext` lands and how to phrase it, see + ; for the model behavior behind + the risk, see + . +- **As of**: 2026-10-01 +- **Recheck trigger**: either section changes where hook text lands or how the model treats text + that arrives beside tool results. + ## What this convention is not - **Not a diagnosis of host-level `PostToolUse` dispatch failure.** When every matching @@ -246,11 +278,11 @@ promotion is tracked at melodic-software/claude-code-plugins#3758. model can act on is itself a conformance defect (redundant user noise, or misrouting agent-actionable content to the user channel). - **Not a UI feature, but "no verbose surface exists" is the wrong reason.** Verbose surfaces do - exist and one of them carries hook output: "Async hook completion notifications are suppressed by - default. To see them, enable verbose mode with `Ctrl+O` or start Claude Code with `--verbose`" - (hooks reference, verified 2026-08-11). Alongside it are the `verbose` and `viewMode` settings, - the `--verbose` flag, `CLAUDE_CODE_DEBUG_LOG_LEVEL=verbose` for hook matcher counts, and - `--include-hook-events` for the stream-json event feed. + exist, and the hooks reference names one that carries background-hook output. Pointer: for that + surface, see . As of: 2026-08-11. + Alongside it are the `verbose` and `viewMode` settings, the `--verbose` flag, + `CLAUDE_CODE_DEBUG_LOG_LEVEL=verbose` for hook matcher counts, and `--include-hook-events` for + the stream-json event feed. The rule this doc needs does not depend on what those surfaces are, only on what an author may assume: **every one of them is off unless the consumer turned it on, and none of them changes @@ -286,10 +318,9 @@ promotion is tracked at melodic-software/claude-code-plugins#3758. > The counts above illustrate that error and support no rule: nothing in this doc's rules > depends on how many pages carry the word. They are deliberately floored ("at least 13") and > need no recheck, since a count that only ever grows cannot falsify the point it illustrates. The - > one claim here that the rules *do* rest on is the quoted `Ctrl+O` / `--verbose` sentence - > (basis: , rung-1 raw-markdown read, 2026-08-11), and it - > argues **for** the rule rather than against it, so its recheck trigger is the one on the - > paragraph above. + > one fact here that the rules *do* rest on is the background-hook surface the pointer above + > names, and it argues **for** the rule rather than against it, so its recheck trigger is the one + > on the paragraph above. ## Conformance @@ -319,9 +350,12 @@ Fleet audits check, per wired producer hook: (wrong tool type, excluded path, empty content, outside the project) that fires before any check logic runs carries no diagnostic information and does not need one. This matches how every current telemetry-emitting hook in the fleet is already shaped. +- Agent-channel text follows the frequency and phrasing rules in + [Text a hook adds for the model](#text-a-hook-adds-for-the-model-frequency-and-phrasing). Not + mechanically gated, but reviewed per hook. `scripts/check-silent-skips.sh` mechanically enforces the second point for the `command -v`-gated shapes it recognizes, **once its pending gate correction lands** (see the systemMessage section above). A bare stderr write does not actually satisfy the doctrine (exit-0 stderr is invisible -per the fresh fetch above), even though the gate does not yet reject it. After that correction, a +per the exit-code caveat above), even though the gate does not yet reject it. After that correction, a quiet skip needs a sanctioned helper call or an explicit `# silent-skip-ok:` annotation. diff --git a/docs/conventions/liveness-assertion/CHANGELOG.md b/docs/conventions/liveness-assertion/CHANGELOG.md index 64672fdf49..29a54c39a3 100644 --- a/docs/conventions/liveness-assertion/CHANGELOG.md +++ b/docs/conventions/liveness-assertion/CHANGELOG.md @@ -4,6 +4,11 @@ Notable changes to the liveness-assertion contract (SemVer). Changing the core c row's conformance bar, or an enforceability verdict is a major bump; additive guidance or new instance rows is a minor bump; docs-only clarification is a patch. +## [1.1.1] - 2026-10-01 + +Patch, docs-only. The Enforceability record names its routing-rule source as a pointer, per the +upstream-drift record shape. No core contract, taxonomy row or enforceability verdict changes. + ## [1.1.0] - 2026-08-28 Additive, minor. It adds a new instance row. The core contract, every taxonomy row's conformance diff --git a/docs/conventions/liveness-assertion/README.md b/docs/conventions/liveness-assertion/README.md index 627cb2ea9d..9077d36901 100644 --- a/docs/conventions/liveness-assertion/README.md +++ b/docs/conventions/liveness-assertion/README.md @@ -128,8 +128,8 @@ Classified per `melodic-software/standards` `conventions/engineering/enforceabil **Peel 1 defers all new mechanical enforcement.** Recorded with event triggers rather than dates: -- **Basis**: `enforceability-tiers.md` routing rule (worth-mechanizing defaults to "not yet" until - the contract exists); peel 1 publishes the contract only per +- **Pointer**: for the routing rule we apply (mechanize nothing until the contract exists), see + `enforceability-tiers.md`; peel 1 publishes the contract only per [#532](https://github.com/melodic-software/claude-code-plugins/issues/532) decision brief Option A. - **Recheck trigger (CI meta-check)**: peel 2+ lands a designed meta-check, **or** a second advisory-lane instance with annotations-only findings reaches `main` after this doc (the #510 diff --git a/docs/conventions/loop-lane/CHANGELOG.md b/docs/conventions/loop-lane/CHANGELOG.md index 3c812fb013..82ebf6185f 100644 --- a/docs/conventions/loop-lane/CHANGELOG.md +++ b/docs/conventions/loop-lane/CHANGELOG.md @@ -5,6 +5,22 @@ topology, the escalation contract, the capability-tier vocabulary, or any loop-l major bump, and additive guidance is a minor bump. A new model release re-audits the capability-tier table (§3); drift found by that audit is recorded here. +## [9.5.0] - 2026-10-01 + +Additive, minor. No topology, escalation-contract, tier-vocabulary, or §4 loop-layer invariant +changed. + +- **Alias binding (§3).** The tier table names each tier's Claude Code alias. The dated + alias-to-version table is removed: which model an alias resolves to is read live from Claude + Code's model page. +- **Known gaps (§3).** The classifier-fallback gap now covers every tier, the fast tier included, + and the usage-credit gap is kept. Both carry one pointer record. +- **Provider gap (§5).** The self-paced `/loop` shape is restated against the provider support the + scheduled-tasks page now documents. +- **Recheck trigger.** Any new model on Claude Code's model page re-reads the §3 alias binding. +- **Reviewer floor (§3).** A reviewer or verifier is never weaker than the implementer in effort + level as well as model tier. + ## [9.4.0] - 2026-09-29 Additive, minor. Section 6 records that the account-identity resolution is built on all three diff --git a/docs/conventions/loop-lane/README.md b/docs/conventions/loop-lane/README.md index f240fa11db..270ccefde8 100644 --- a/docs/conventions/loop-lane/README.md +++ b/docs/conventions/loop-lane/README.md @@ -282,35 +282,37 @@ record directory, with a `type: "http"` handler that POSTs the hook event's JSON } ``` -Every element is a documented first-party mechanism (verified against - and on -2026-07-27): - -- `type: "http"` handlers POST the hook's JSON input with `Content-Type: application/json` and are - supported in project `.claude/settings.json`, and in every other settings scope, on `PostToolUse`; - the one documented handler-type restriction that excludes them is on `SessionStart`. The seam is - therefore per-consuming-repo configuration; no plugin ships it. It is deterministic (the handler - fires on the matched lifecycle event, no model judgment) and carries no claude.ai subscription or - Remote Control dependency. -- The `if` field holds exactly one permission rule and is evaluated on `PostToolUse`. File rules - use the `Edit(...)` form, since Edit rules cover all file-editing tools, `Write` included, and a - `Write(path)` rule is never matched, and the single leading `/` anchors at the settings source - (`` for project settings). Each worktree checkout carries its own copy of the - tracked settings file, so by that settings-source rule the one tracked rule anchors at each - worktree's own root. That is an applied inference: the docs state worktree matching explicitly - only for local-settings rules. -- Header values interpolate environment variables only for names listed in `allowedEnvVars`. The - docs document interpolation for `headers` alone and say nothing about `url`, so treat the `url` - field as non-interpolating, an applied inference, and the reason the endpoint URL is tracked - config while the secret rides only in a header sourced from the operator's environment, never in - the repo. -- **Egress note.** The POST body is the full `PostToolUse` hook input, not just the record: - alongside `tool_input` (the record's path and content) it carries session metadata, for +Every element is a first-party mechanism. The bullets state what the seam relies on; each specific +is read live at the pointer. + +- **Pointer**: for the handler type, its fields and header interpolation, see + [HTTP hook fields](https://code.claude.com/docs/en/hooks#http-hook-fields); for the `if` field, + see [Common fields](https://code.claude.com/docs/en/hooks#common-fields) and + [Read and Edit](https://code.claude.com/docs/en/permissions#read-and-edit); for the POST body, + see [Common input fields](https://code.claude.com/docs/en/hooks#common-input-fields); for + failures, see [HTTP response handling](https://code.claude.com/docs/en/hooks#http-response-handling). +- **As of**: 2026-07-27 +- **Recheck trigger**: a Claude Code release note or docs change touching HTTP hook handlers, the + `if` field, header interpolation, or how edit rules anchor their paths. + +- The seam is an `http` handler on `PostToolUse` in the consuming repo's project + `.claude/settings.json`: per-consuming-repo configuration that no plugin ships. We rely on it as + deterministic (it fires on the matched lifecycle event, no model judgment) and as carrying no + claude.ai subscription or Remote Control dependency. +- The `if` field carries exactly one rule, in the `Edit(...)` form (never `Write(path)`), with a + single leading `/` so it anchors at the settings source. Each worktree checkout carries its own + copy of the tracked settings file, so we rely on the one tracked rule anchoring at each + worktree's own root. That is an applied inference, not a documented guarantee. +- The secret rides only in a header whose variable is listed in `allowedEnvVars`, sourced from the + operator's environment and never from the repo. We treat the `url` field as non-interpolating, + an applied inference, which is why the endpoint URL is tracked config. +- **Egress note.** We treat the POST body as the full `PostToolUse` hook input, not just the + record: alongside `tool_input` (the record's path and content) it carries session metadata, for example `session_id`, `cwd`, and `transcript_path`, which are absolute local paths and project identity. Configuring the hook is the consuming repo's deliberate opt-in to that egress; point the URL only at an endpoint trusted with it. -- A non-2xx response or a connection failure is a non-blocking error: a dead endpoint never blocks - a lane. +- We rely on a failed POST (a non-2xx response or a connection failure) being non-blocking: a dead + endpoint never blocks a lane. **Destination is the consumer's choice.** The URL is any HTTP endpoint the consuming repo controls: a generic webhook receiver, an internal alerting service, or a relay that reshapes the @@ -318,14 +320,22 @@ payload for a chat service (a Slack incoming webhook expects its own JSON shape raw hook payload, so Slack reach goes through a relay). Two non-deterministic layers may ride alongside, never instead: the built-in `PushNotification` tool, and model-driven outbound send via a chat plugin (UNVERIFIED here: confirm the plugin and its send capability against its own docs -before relying on it). `PushNotification` "sends a desktop notification, and a phone push when -Remote Control is connected"; it prompts for no permission, but the model decides when to call it. -Its phone leg therefore inherits every condition the Remote Control page enumerates under -Requirements, plus its mobile-push setup steps. One condition matters here in particular: -`DISABLE_TELEMETRY`, `DO_NOT_TRACK`, `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC`, and -`DISABLE_GROWTHBOOK` each disable the feature-flag evaluation Remote Control depends on (verified -2026-08-04: , -). Only the http hook is the deterministic leg. +before relying on it). The model decides when to call `PushNotification`, and its phone leg +depends on Remote Control, so we treat that leg as absent whenever any Remote Control requirement +or mobile-push setup step is unmet, including on a machine that turns feature-flag fetching off. +Only the http hook is the deterministic leg. + +- **Pointer**: for the tool, see the `PushNotification` row of the + [Tools reference](https://code.claude.com/docs/en/tools-reference) (the tool table sits under the + page title, with no section of its own); + for the phone leg's conditions, see + [Remote Control requirements](https://code.claude.com/docs/en/remote-control#requirements) and + [Mobile push notifications](https://code.claude.com/docs/en/remote-control#mobile-push-notifications); + for the variables that turn fetching off, see + [Features that need feature-flag fetching](https://code.claude.com/docs/en/env-vars#features-that-need-feature-flag-fetching). +- **As of**: 2026-08-04 +- **Recheck trigger**: a Claude Code release note changes `PushNotification` or the Remote Control + requirements. **The seam binds to the session's project, never to the repository a lane targets.** The record path is relative to the session's checkout, and the hook that fires is the one in that session's @@ -351,11 +361,11 @@ closed laptop or a dead process emits no hook event at all; the record write cov running but unattended, and lane-down detection stays with the stop gate and telemetry freshness (§4). -**A configured hook can also fail silently.** An env-var name absent from `allowedEnvVars` -interpolates as an empty string (documented: "references to unlisted variables are replaced with -empty strings"); a listed name unset in the operator's environment has no value to supply and -plausibly interpolates the same way, an applied inference not stated in the docs. Either way, a -non-2xx response or connection failure is a non-blocking error, so a misconfigured hook can 401 on +**A configured hook can also fail silently.** We treat a header variable missing from +`allowedEnvVars` as interpolating to an empty string (for the rule, see +[HTTP hook fields](https://code.claude.com/docs/en/hooks#http-hook-fields), as of 2026-07-27), and +a listed variable unset in the operator's environment as doing the same, an applied inference. +Either way a failed POST is non-blocking, so a misconfigured hook can 401 on every escalation while the lane runs on with nothing surfaced outside debug logs. Verify the leg when wiring it, by writing a throwaway record file with the Write tool and confirming the endpoint received the POST, and treat webhook silence across cycles that filed escalations as a @@ -365,56 +375,74 @@ check-the-hook signal, never as proof of health. Model selection is expressed as **capability tiers defined by order, never by family name**. Capability does not track family across generations (a current mid-tier model can equal a prior -top-tier one), so a tier named for a family silently rots. Three ordered tiers: +top-tier one), so a tier named for a family silently rots. Three ordered tiers, each bound to a +Claude Code alias (see "Alias binding" below): -| Tier | Role | -|---|---| -| frontier | Complex-stamped items; every security-surface work class, always | -| strong | Default implementer / worker | -| fast | Orchestrator and mechanical items; never weaker than the implementer it reviews | +| Tier | Role | Alias | +|---|---|---| +| frontier | Complex-stamped items; every security-surface work class, always | `best` | +| strong | Default implementer / worker | `opus` | +| fast | Orchestrator and mechanical items; never weaker than the implementer it reviews | `sonnet` | Fixed rules: an advisor or reviewer is **at least as capable** as the main model it checks (equal pairings are valid, and a fast orchestrator paired with an advisor at or above the main tier is the -recommended shape); a reviewer or verifier is never weaker than the implementer; a security-surface +recommended shape); a reviewer or verifier is never weaker than the implementer, in model tier or +in effort level; a security-surface work class routes to the frontier tier unconditionally. -### Current alias binding (re-audited 2026-08-12) - -The dated resolution of the ordered tiers to live aliases, the artifact the "new model release" -recheck trigger re-derives. Sourced from live fetches of - and - on 2026-08-12 (#1293); the -resolutions re-verified 2026-09-23 against both pages after the Opus 5.5 and Fable 5.1 releases, -and the fast row re-verified 2026-10-01 against both after Sonnet 5.5 became the `sonnet` alias: - -| Tier | Alias | Resolves to today | -|---|---|---| -| frontier | `best` | Fable 5.1 where the organization has access, else the latest Opus | -| strong | `opus` | Opus 5.5 | -| fast | `sonnet` | Sonnet 5.5 on the Anthropic API; the model-config provider table lists older Sonnet versions elsewhere | - -- **frontier binds `best`, not `fable`.** `best` is the docs' live handle for exactly the frontier - tier's meaning, "the model the `fable` alias resolves to where Fable is available to you, - otherwise the same model as `opus`", so a frontier dispatch self-heals where Fable is unavailable (it requires organization access - and Claude Code v2.1.170+, and can bill to usage credits) instead of failing or silently running - a stale pin. Two Fable caveats ride along as **known gaps**: its safety classifiers can trigger - automatic model fallback "most often in cybersecurity and biology domains", and frontier is the - tier every security-surface work class routes to, and no lane detects that fallback today (Opus - 5.5 and Sonnet 5.5 carry the same classifiers, so the strong and fast tiers share this gap); and in - non-interactive mode a Fable request that would bill usage credits bills them without a consent - prompt, which is the shape every unattended lane runs in. -- **strong binds `opus`.** The docs' own starting recommendation, "start with Claude Opus 5.5 for - most workloads". Opus 5.5 and Fable 5.1 both have reliable knowledge through June 2026, so - freshness does not separate them, and raw capability order (Fable above Opus) does not decide - the binding alone. -- **fast binds `sonnet`.** "Best combination of speed and intelligence", native 1M context, Jun - 2026 reliable cutoff: enough headroom to orchestrate and to review mechanical items without - breaching the reviewer floor. Its effort default and its early-stop and skipped-check tendencies - at lower effort are in the playbooks Sonnet 5.5 chapter, not restated here. -- **`haiku` is admissible nowhere in these lanes today.** Its 200k context sits against 1M - everywhere else, and its Feb 2025 reliable cutoff predates the harness surfaces these lanes - operate on; since the fast tier also covers reviewers and the implementer is always - sonnet-or-above, binding `haiku` anywhere would breach the reviewer-never-weaker floor. +### Alias binding + +The tier table above binds each tier to an alias, never to a model version, so a release that +moves an alias needs no edit here. Which model an alias resolves to, on each provider, is read live +from the model page whenever it matters, and never restated in this convention. + +- **Pointer:** for what each alias resolves to on each provider, see + [Model aliases](https://code.claude.com/docs/en/model-config#model-aliases); for each model's + position and capabilities, see + [models overview: latest models comparison](https://platform.claude.com/docs/en/about-claude/models/overview#latest-models-comparison). +- **As of:** 2026-10-01. +- **Recheck trigger:** any new model on Claude Code's model page. + +The reasons behind each binding: + +- **frontier binds `best`, not `fable`.** We bind the frontier tier to `best` because that alias + already carries the tier's meaning, Fable where the organization has it and Opus otherwise, so a + frontier dispatch self-heals where Fable is unavailable instead of failing or silently running a + stale pin. For Fable's access, version and billing requirements, see + [Work with Fable](https://code.claude.com/docs/en/model-config#work-with-fable). +- **strong binds `opus`.** We bind the strong tier to `opus` because the models overview names the + model it resolves to as the general starting point; raw capability order (Fable above Opus) does + not decide the binding alone. +- **fast binds `sonnet`.** We bind the fast tier to `sonnet` for its speed relative to the tiers + above and its context headroom: enough to orchestrate and to review mechanical items without + breaching the reviewer floor. The tier has nothing to do with Claude Code's fast mode, a separate + speed setting for Opus (for fast mode, see + [Speed up responses with fast mode](https://code.claude.com/docs/en/fast-mode)). +- **`haiku` is admissible nowhere in these lanes today.** The fast tier also covers reviewers and + the implementer is always `sonnet` or above, so binding `haiku` anywhere would breach the + reviewer-never-weaker floor. We also read the model it resolves to as having a smaller context + window and an older knowledge cutoff than these lanes need (for both, see the models overview). + +**Known gaps carried with the binding.** No lane detects either of these today, so each is recorded +here rather than left as an unstated assumption: + +- **Classifier fallback, on every tier.** The models the `best`, `opus` and `sonnet` aliases + resolve to run safety classifiers. A flagged request can re-run on a different model, after + which the session stays there, or end in a refusal for a category with nowhere to fall back + to. Security work trips them most often, and frontier is the tier every security-surface work + class routes to, but the strong and fast tiers carry the same gap: a lane's tier can drop + mid-run, or a cycle can stop on a refusal, with nothing in the lane noticing. +- **Usage-credit consent.** In non-interactive mode, the shape every unattended lane runs in, a + Fable request that would bill usage credits bills them without a consent prompt. +- **Pointer:** for the fallback targets, the categories without one, and the provider setup, see + [Automatic model fallback](https://code.claude.com/docs/en/model-config#automatic-model-fallback) + and + [Security research and biology workloads](https://code.claude.com/docs/en/model-config#security-research-and-biology-workloads); + for usage-credit consent, see + [Fable and usage credits](https://code.claude.com/docs/en/model-config#fable-and-usage-credits). +- **As of:** 2026-10-01. +- **Recheck trigger:** a model gains or loses a fallback target, a new model on Claude Code's model + page runs safety classifiers, or non-interactive mode starts asking for usage-credit consent. **Independence, where a dispatch stands in for human ratification.** The one dispatch that resolves a blocker in place of a human decision, the explicit-`autopilot` merge-authority exception (above), @@ -439,24 +467,32 @@ is the boundary's stated justification, so the boundary is revisited when that p path whose outcome stops being gate-decidable acquires the independence requirement, recorded as a versioned entry in [`CHANGELOG.md`](CHANGELOG.md) rather than silently. -**Runtime resolution is by model alias only.** The bare family-word aliases -(`fable` / `opus` / `sonnet` / `haiku`) are the live-updating handles that resolve to the current -recommended model for the provider and update over time; a dated model name is a pinned snapshot and -is never written into a lane body. Aliases are the only handle guaranteed under subscription OAuth, -so they are the runtime path; the Models API list endpoint is the **build/audit-time** verification -path, since it may require an API key a loop session lacks. No lane hard-codes a model ID. (Alias -semantics verified against on 2026-08-04.) +**Runtime resolution is by model alias only.** A lane names an alias (`best` / `fable` / `opus` / +`sonnet` / `haiku`), never a dated model name, because we treat the alias as the handle that +follows the provider's recommendation while a model ID is a pinned snapshot. We treat aliases as +the only handle that works under subscription OAuth, so they are the runtime path; the Models API +list endpoint is the **build/audit-time** verification path, since it may require an API key a loop +session lacks. No lane hard-codes a model ID (Pointer: for alias semantics, see +[Model aliases](https://code.claude.com/docs/en/model-config#model-aliases). As of: 2026-10-01. +Recheck trigger: the model page stops describing aliases as moving with the provider's +recommendation). -Tier tables are built from a live official-docs fetch at authoring time, never from recall. Any new -model release re-audits the tier table, and the trigger is recorded in this convention's -[`CHANGELOG.md`](CHANGELOG.md). +The binding is built from a live official-docs read at authoring time, never from recall. Any new +model on Claude Code's model page re-reads the reasons under "Alias binding"; a binding that +changes is recorded in this convention's [`CHANGELOG.md`](CHANGELOG.md). ### Rate-limit windows -Subscription (Pro/Max) usage is bounded by a rolling five-hour window and a weekly cap. The weekly -cap's exact model scoping and numeric limits are volatile and are **not** restated here. See the -official [Anthropic support article](https://support.claude.com/en/articles/11049741-what-is-the-max-plan) -(verified 2026-07-23). The operable pause floor lives in the rate-limit guard binding (§6). +The lanes model subscription (Pro/Max) usage as two limits, one over the last five hours and one +over the week, and +act only on the operable pause floor in the rate-limit guard binding (§6); the cap's scoping and +limits are not restated here. + +- **Pointer**: for the subscription usage windows and caps, see the + [Anthropic support article](https://support.claude.com/en/articles/11049741-what-is-the-max-plan). +- **As of**: 2026-07-23 +- **Recheck trigger**: the support article changes the window structure, or a Claude Code release + note changes the rate-limit fields the guard reads. ## 4. Loop-layer invariants @@ -468,10 +504,10 @@ state**: when every remaining open item is human-gated or escalated and no PR is reports and stops cleanly rather than idling forever. Without it, an overnight drain deadlocks on the first unanswered escalation. -A standing lane is additionally bounded by the `/loop` launch surface's **seven-day expiry**: a -`/loop` ends automatically seven days after it starts, on either launch shape (§5) and idle backoff -notwithstanding (, verified -2026-07-27, broadened from the 2026-07-23 stamp's self-paced-only wording). A standing lane +A standing lane is additionally bounded by the `/loop` launch surface's **seven-day expiry**, +which we treat as binding both launch shapes (§5), idle backoff notwithstanding. Pointer: +[Seven-day expiry](https://code.claude.com/docs/en/scheduled-tasks#seven-day-expiry). As of: +2026-07-27. Recheck trigger: a Claude Code release note changes `/loop` expiry. A standing lane therefore requires a relaunch owner, today always the operator, for whom `harness-ops` `lanes` `restart` is a one-command path (operator-initiated by contract; see the cycle-budget paragraph below). The lane records its loop-started timestamp in the lane's #502 telemetry block so the @@ -479,11 +515,12 @@ approaching expiry is visible ahead of time, and an expiry hit is handled exactl cycle-budget hit below: a restart-request into the #502 block, then a clean stop. **Self-pacing.** A lane paces itself through `/loop` with the interval omitted; Claude schedules the -next iteration with `ScheduleWakeup`, whose delay is clamped between one minute and one hour. -`ScheduleWakeup` is called at the end of each iteration and is not operator-callable (verified -against and - on 2026-07-27, no drift from the prior -2026-07-23 stamp). Idle raises the delay toward the ceiling. The self-pacing section the +next iteration with `ScheduleWakeup`, and no lane expects an operator to call it. Idle raises the +delay toward the ceiling. Pointer: for the delay bounds and who calls the tool, see +[Let Claude choose the interval](https://code.claude.com/docs/en/scheduled-tasks#let-claude-choose-the-interval) +and the `ScheduleWakeup` row of the [Tools reference](https://code.claude.com/docs/en/tools-reference). +As of: 2026-07-27. Recheck trigger: a Claude Code release note changes `ScheduleWakeup` or +self-paced `/loop`. The self-pacing section the `source-control:babysit-prs` skill owns is the worked precedent. **The prompt runs fresh; the session does not.** Each cycle re-sends the lane's prompt verbatim into @@ -655,12 +692,13 @@ session renders a status line, so an unattended lane samples nothing, and an emp unobserved rather than zero; the figures are **account-scope**, so the three-lane topology means concurrent lanes move the same windows and a per-cycle rise is one lane's own consumption only when that lane is the sole active session; and they are a **percentage of a subscription window, not a -token count**, absent entirely for non-subscription auth. No lane claims a token count, because none -is *readable* at a cycle boundary: the machine-readable token fields a session exposes are -current-context occupancy, not session totals. A machine-readable cumulative *cost* field does -exist, and is session-scoped, so it would attribute to a lane, but the guard's tee does not -forward it; widening the tee is a guard-side change this invariant deliberately does not make -(, verified 2026-07-28). +token count**, absent entirely for non-subscription auth. No lane claims a token count, because we +found no session-total token field readable at a cycle boundary. The status line's cumulative +*cost* field would attribute to a lane, but the guard's tee does not forward it; widening the tee +is a guard-side change this invariant deliberately does not make. Pointer: for the fields a +session exposes, see [Available data](https://code.claude.com/docs/en/statusline#available-data). +As of: 2026-07-28. Recheck trigger: a Claude Code release adds or changes a token or cost field in +the status line input. **Headless-config floor.** A headless lane launch never blocks on an interview: it takes explicit or persisted config, or tier defaults, and logs the assumption. The interactive path may run a @@ -719,33 +757,35 @@ All three adopters have shipped. This owner doc landed ahead of them, per the co rule; the table above is a live consumer list, not a forward reference. **Launch surfaces.** A lane launches interactively via `/loop`, the primary surface and a bundled -skill needing no install (, verified -2026-08-02), or headless via the `harness-ops` `lanes` launcher, which stores the one-line lane +skill needing no install (Pointer: [Bundled skills](https://code.claude.com/docs/en/skills#bundled-skills). +As of: 2026-08-02. Recheck trigger: `/loop` leaves the bundled set), or headless via the `harness-ops` `lanes` launcher, which stores the one-line lane prompt through its `prompt_dir` interface (#480). `lanes` is a **supporting, strictly one-directional** launcher: it launches the lane; no lane body ever requires, imports, or degrades without `harness-ops`. Every mention of `lanes` in a lane body is presence-gated with the `/loop` fallback documented at the site, per the [seam-phrasing convention](../seam-phrasing/README.md). **Two launch shapes, selected per invocation, and neither deprecates the other.** Supplying an interval -(`/loop 15m …`) converts it to a cron expression and fires on that fixed schedule, subject to -jitter; omitting it hands the delay to Claude, which picks one per iteration within the §4 bounds -and is not jittered. `ScheduleWakeup` reschedules a *self-paced* loop only, so it is not the pacing -mechanism once an interval is supplied -(, verified -2026-07-27). The §4 seven-day expiry binds both shapes. Both are current; this note reconciles which -applies where and changes neither. - -Jitter is the scheduler's deterministic offset on a *cron* task: up to 30 minutes after the -scheduled time, or up to half the interval for a task running more often than hourly. +(`/loop 15m …`) gives a fixed, jittered cron schedule; omitting it gives the self-paced shape, where +Claude picks each delay within the §4 bounds. We treat `ScheduleWakeup` as the pacing mechanism of +the self-paced shape only. The §4 seven-day expiry binds both shapes. Both are current; this note +reconciles which applies where and changes neither. Pointer: for both shapes, see +[Run on a fixed interval](https://code.claude.com/docs/en/scheduled-tasks#run-on-a-fixed-interval) +and +[Let Claude choose the interval](https://code.claude.com/docs/en/scheduled-tasks#let-claude-choose-the-interval); +for the size of the offset a cron task gets, see +[Jitter](https://code.claude.com/docs/en/scheduled-tasks#jitter). As of: 2026-07-27. Recheck +trigger: a Claude Code release note changes either `/loop` shape or the jitter rule. - **A lane always omits the interval.** Two §4 invariants need the self-paced shape and neither survives a cron schedule. *Idle backoff*, the standing shape's "idle backs off toward longer wakeups", derives the next delay from what the cycle just observed, which a fixed cadence cannot - consume. And a self-paced loop can **end itself**, because Claude calls `ScheduleWakeup` with - `stop: true`, which is how the drain shape's terminal state stops a lane cleanly; a fixed-interval - loop keeps running until stopped by hand or until the seven-day expiry, so a drain lane launched - that way cannot honor its own stop condition - (, verified 2026-07-27). Self-paced is + consume. And a self-paced loop can **end itself** (Claude calls `ScheduleWakeup` with + `stop: true`), which is how the drain shape's terminal state stops a lane cleanly; we treat a + fixed-interval loop as running until stopped by hand or until the seven-day expiry, so a drain + lane launched that way cannot honor its own stop condition (Pointer: + [Stop a loop](https://code.claude.com/docs/en/scheduled-tasks#stop-a-loop). As of: 2026-07-27. + Recheck trigger: that section changes how a fixed-interval or self-paced loop ends). + Self-paced is the lane shape by construction, not by preference. Two of the lane's other per-cycle signals, the adaptive-cap streak and seam exit 8 counted as dirty, govern *how much work a cycle takes on*, not when the next one fires, and are unaffected by either shape. The drain-exit snapshot is @@ -762,16 +802,21 @@ while its cadence mapping (the self-pacing cadence contract owned by the `source-control:babysit-prs` skill) is the self-paced contract the `babysit-loop` lane consumes. Reading either as the other's default is the confusion this note exists to prevent. -**Known gap: the self-paced shape is provider-conditional.** On Amazon Bedrock, Claude Platform on -AWS, Google Cloud's Agent Platform, and Microsoft Foundry, an omitted interval does **not** hand the -delay to Claude: the prompt runs on a fixed ten-minute schedule and `ScheduleWakeup` is unavailable -(, -, verified 2026-07-27). A lane launched there keeps +**Known gap: the self-paced shape is version-conditional off the first-party API.** We record a lane +launched on Microsoft Foundry, Amazon Bedrock, Google Cloud's Agent Platform, or Claude Platform on +AWS, or with feature-flag fetching turned off, as having the self-paced shape only on a Claude Code +version at or above the floor the scheduled-tasks page names for those providers; below that floor +it runs without it. Pointer: for the provider and version conditions on a dynamic `/loop`, see +[Let Claude choose the interval](https://code.claude.com/docs/en/scheduled-tasks#let-claude-choose-the-interval) +and the `ScheduleWakeup` row of the [Tools reference](https://code.claude.com/docs/en/tools-reference). +As of: 2026-10-01. Recheck trigger: a Claude Code release note changes `/loop` on a non-first-party +provider or the version floor. A lane launched below the floor keeps the loop but loses both properties the bullet above depends on: idle backoff cannot lengthen the wake, and the lane cannot end itself, so a **drain** lane there deadlocks on the first unanswered escalation exactly as §4's terminal state exists to prevent, and runs until stopped by hand or until -the seven-day expiry. No lane detects the provider today, so this is recorded as a known gap rather -than left as an unstated assumption, on the model §6 uses for the single-account assumption. +the seven-day expiry. No lane detects the provider or the version today, so this is recorded as a +known gap rather than left as an unstated assumption, on the model §6 uses for the single-account +assumption. ## 6. Rate-limit guard binding @@ -953,16 +998,15 @@ This contract is versioned in [`CHANGELOG.md`](CHANGELOG.md). A change to the to escalation contract, the tier vocabulary, or any loop-layer invariant is a major bump; additive guidance is a minor bump. -**Recheck triggers** ([upstream-drift](../upstream-drift/README.md) owns the stamp-and-trigger -discipline). Two. A firing that finds drift lands its outcome as a changelog entry; a no-drift -firing refreshes the claim's verification date in place, with no entry and no bump: +**Recheck triggers** ([upstream-drift](../upstream-drift/README.md) owns the record shape and the +trigger discipline). Two. A firing that finds drift lands its outcome as a changelog entry; a +no-drift firing refreshes the record's as-of date in place, with no entry and no bump: -- Any new model release re-audits the capability-tier table (§3). -- Any change to this convention, or to a consuming lane, that RELIES on an upstream-sourced claim - re-verifies that claim against its cited page first and refreshes the claim's verification date - with the outcome. +- Any new model on Claude Code's model page re-reads the §3 alias binding and its known gaps. +- Any change to this convention, or to a consuming lane, that RELIES on an upstream-pointed record + re-reads that record's pointer first and refreshes its as-of date with the outcome. -The upstream surfaces these claims rest on, the `/loop` seven-day expiry, the `ScheduleWakeup` +The upstream surfaces these records point at, the `/loop` seven-day expiry, the `ScheduleWakeup` bounds, model-alias semantics, and the rate-limit windows, move on a research-preview cadence. Where re-verification finds drift, the changed value lands here as a recorded entry rather than silently inside a lane body. diff --git a/docs/conventions/native-references/CHANGELOG.md b/docs/conventions/native-references/CHANGELOG.md index d4d92bdb4b..6352a98ea2 100644 --- a/docs/conventions/native-references/CHANGELOG.md +++ b/docs/conventions/native-references/CHANGELOG.md @@ -6,6 +6,22 @@ major change; additive guidance is minor; clarification is a patch. The doc ship unnumbered, which this file reads as **1.0**; the entry below is the first recorded change and lands the changelog the README said would arrive with it. +## [3.4.0] - 2026-10-01 + +Minor: additive guidance. No required part of the description phrase, no canonical gate token and +no enforceability verdict changes. + +- **A Boundary reference file holds links-only records.** The reference file inside a skill that + backs a `## Boundary` section now holds our decision in our words, a pointer to the exact + upstream section, the as-of date and the recheck trigger, and no upstream text; a table form uses + the header `| Decision | Pointer | As of | Recheck when |`. Existing `native-*` and `bundled-*` + reference files keep the older four-part shape until + [#5684](https://github.com/melodic-software/claude-code-plugins/issues/5684) converts them. +- **The convention's own upstream specifics are restated as records.** The gating-axes table, the + description budget caveat and the suggest sentence's `` now point at the docs sections + with an as-of date and a recheck trigger in place of restated page text. +- **The cloud-sessions link follows the docs site's new heading id.** + ## [3.3.6] - 2026-09-30 Patch: clarification. diff --git a/docs/conventions/native-references/README.md b/docs/conventions/native-references/README.md index 15dcc4554c..4cb2dbf87e 100644 --- a/docs/conventions/native-references/README.md +++ b/docs/conventions/native-references/README.md @@ -21,9 +21,10 @@ This doc owns the phrasing of references **to native surfaces**. It does not own - **Whether a reference should exist at all.** That is a verdict, and verdicts live in the committed overlap store rendered into [`docs/native-surfaces.md`](../../native-surfaces.md). This doc governs the words once a verdict says a reference is warranted. -- **The stamp discipline on any upstream fact a reference restates.** - [`upstream-drift`](../upstream-drift/README.md) owns the four-part record (claim, basis, as-of - date, recheck trigger) and the observability bar its triggers must clear. +- **The record behind any upstream specific a reference depends on.** + [`upstream-drift`](../upstream-drift/README.md) owns its shape (our decision, a pointer to the + exact upstream section, the as-of date, the recheck trigger) and the observability bar its + triggers must clear. - **Instruction economy.** [`plugin-philosophy`](../../plugin-philosophy.md) owns the rule that every always-loaded description is a per-session tax. This doc keeps the phrase to one clause because of that rule; it does not restate it. @@ -33,21 +34,26 @@ This doc owns the phrasing of references **to native surfaces**. It does not own Native availability varies along at least four independent axes, so any static availability sentence is wrong somewhere by construction: -| Axis | Mechanism | +| Axis | Where it is set | |---|---| -| Settings / environment | `disableBundledSkills` and `CLAUDE_CODE_DISABLE_BUNDLED_SKILLS` remove bundled skills and workflows; `skillOverrides` maps a name to `on` / `name-only` / `user-invocable-only` / `off`; `DISABLE_DOCTOR_COMMAND` hides `/doctor` specifically | -| Plan | Some surfaces require a paid or specific plan tier | -| Platform / provider | Some surfaces are absent on some OSes, and several are unavailable on non-first-party model providers | -| Host surface | CLI, web/cloud, VS Code, and mobile expose different rosters; terminal-interface commands do not exist in a web session, and a cloud session carries session-provided skills a local CLI does not | - -Claim, basis, and trigger for that table, per [`upstream-drift`](../upstream-drift/README.md): -the four axes are documented on `https://code.claude.com/docs/en/settings-reference.md` -(`disableBundledSkills`, `skillOverrides`), `https://code.claude.com/docs/en/env-vars.md`, -`https://code.claude.com/docs/en/commands.md` ("Not every command appears for every user. -Availability depends on your platform, plan, and environment."), and -`https://code.claude.com/docs/en/cloud-environments.md`; verified 2026-08-23; **recheck trigger**: -a Claude Code release note or docs change adds, removes, or renames a gating axis, or a -`skillOverrides` state leaves the four-value set. +| Settings / environment | The `disableBundledSkills` setting and its environment twin, `skillOverrides`, and `DISABLE_DOCTOR_COMMAND` | +| Plan | The account's plan tier | +| Platform / provider | The operating system and the model provider | +| Host surface | CLI, web/cloud, VS Code, or mobile | + +We treat every axis as able to remove or hide a native surface, so no component states +availability. + +- **Pointer**: for the switches, see + [`disableBundledSkills`](https://code.claude.com/docs/en/settings-reference#disablebundledskills), + [skill visibility overrides](https://code.claude.com/docs/en/skills#override-skill-visibility-from-settings) + and [environment variables](https://code.claude.com/docs/en/env-vars); for plan and platform + gating, see [Commands](https://code.claude.com/docs/en/commands); for what a cloud session + carries, see + [What's available in cloud sessions](https://code.claude.com/docs/en/cloud-environments#what%E2%80%99s-available-in-cloud-sessions). +- **As of**: 2026-10-01 +- **Recheck trigger**: a Claude Code release note or docs change adds, removes, or renames a gating + axis, or a `skillOverrides` state leaves the four-value set. The consequence is the rule: **a component never states that a native surface is present, absent, enabled, or unavailable.** It states what to do *if the surface resolves in the session*, and the @@ -92,19 +98,18 @@ write "otherwise this skill", which is noise the shared budget pays for. ### Budget caveat -Descriptions are subject to two limits, and a baked phrase is the best available routing surface, -not a guaranteed one: +Descriptions are subject to two limits, a per-entry length cap and a budget on the whole listing, +so we treat a baked phrase as the best available routing surface, not a guaranteed one: a phrase +may be cut short, or its whole description dropped from the listing. -- the combined `description` + `when_to_use` text is truncated at **1,536 characters** in the - listing by default (`skillListingMaxDescChars`); and -- the listing as a whole is capped at a **share of the context window** - (`skillListingBudgetFraction`, default 1%). On overflow the listing keeps every skill *name* and - drops whole descriptions, starting with the least-invoked skills. - -Basis: `https://code.claude.com/docs/en/skills.md` (Frontmatter reference; Troubleshooting → -"Skill descriptions are cut short") and `https://code.claude.com/docs/en/settings-reference.md`; -verified 2026-08-23. **Recheck trigger**: a release or docs change moves the 1,536 default, the -1% default, or the drop-order rule. +- **Pointer**: for the per-entry cap, the listing budget and what overflow drops, see + [Skill descriptions are cut short](https://code.claude.com/docs/en/skills#skill-descriptions-are-cut-short), + [`skillListingMaxDescChars`](https://code.claude.com/docs/en/settings-reference#skilllistingmaxdescchars) + and + [`skillListingBudgetFraction`](https://code.claude.com/docs/en/settings-reference#skilllistingbudgetfraction). +- **As of**: 2026-10-01 +- **Recheck trigger**: a release or docs change moves the per-entry cap default, the budget + default, or the drop-order rule. Two obligations follow. Keep the phrase to one clause, since it spends shared budget every session for every consumer. And where a fleet's listing plausibly overflows, the overlap store records a @@ -132,10 +137,13 @@ Boundary section spends no shared listing budget and changes no routing; the gat earns (below) has no reason to hold the section back. The section carries the conclusion: the surfaces by provenance class, the routing split, and the -mutation gate. The four-part records behind it (the basis each upstream specific rests on, its -as-of date, its recheck trigger, the extraction or docs evidence) live in a **reference file inside -the same skill**, linked from the section with a same-plugin relative path, so the body stays short -and the detail stays reachable. Modeled on the `review` plugin's organic pattern (`/review:quality-gate` +mutation gate. The records behind it live in a **reference file inside the same skill**, linked +from the section with a same-plugin relative path, so the body stays short and the detail stays +reachable. That file holds our decision in our words, a pointer to the exact upstream section, the +as-of date and the recheck trigger, and no upstream text; a table form uses the header +`| Decision | Pointer | As of | Recheck when |`. Existing `native-*` and `bundled-*` reference +files keep the older four-part shape until +[#5684](https://github.com/melodic-software/claude-code-plugins/issues/5684) converts them. Modeled on the `review` plugin's organic pattern (`/review:quality-gate` and `/review:fanout` each carry one): ```markdown @@ -171,8 +179,8 @@ Six properties the section keeps: phrase. Cross-plugin pointers are forbidden. 5. **Presence-gated language throughout**: the body inherits the description's gate; it never promotes a surface to available because the body is longer. -6. **Upstream specifics carry their basis and date**, per - [`upstream-drift`](../upstream-drift/README.md). +6. **Upstream specifics are pointed at, never restated**: each carries a pointer, an as-of date + and a recheck trigger, per [`upstream-drift`](../upstream-drift/README.md). ## Self-containment: shipped plugins never cite the registry @@ -209,10 +217,10 @@ Classified per `melodic-software/standards` `conventions/engineering/enforceabil |---|---| | `/harness-ops:audit-install-state` | `## Boundary` section for the bundled `doctor` skill (verdict `complementary`); no description phrase | | `/review:quality-gate`, `/review:fanout` | The organic Boundary pattern this doc generalizes; adopts the phrasing rules on next touch | -| `/harness-config:audit-instructions` | `## Boundary` section for the bundled `claude-api` skill's `prompt-audit` subcommand (verdict `complementary`, composite posture), four-part detail in the skill's own reference file; a description phrase for `claude-api`. A second `## Boundary` section for the bundled `doctor` skill's `prompt-audit` (integration `suggest`) | +| `/harness-config:audit-instructions` | `## Boundary` section for the bundled `claude-api` skill's `prompt-audit` subcommand (verdict `complementary`, composite posture), pointer records in the skill's own reference file; a description phrase for `claude-api`. A second `## Boundary` section for the bundled `doctor` skill's `prompt-audit` (integration `suggest`) | | `/evals:methodology` | `## Boundary` section for the bundled `claude-api` skill's `hillclimb` and `build-eval` subcommands (verdict `complementary`); a description phrase routing the model-and-effort sweep to `hillclimb`; detail in the skill's eval-design reference | | `/playbooks:fable-5` | `## Boundary` section for the bundled `claude-api` skill as the live-facts and cost-audit surface its chapters defer to (verdict `complementary`); a description phrase for its `route` row; detail in the pack's prompt-caching reference chapter | -| `/review:code-review`, `/review:security-review` | Description phrase + `## Boundary` section for the bundled `code-review` skill and the plugin-backed built-in `security-review` command (verdict `complementary`, CI lane versus session pass); four-part detail in each skill's `reference/` file | +| `/review:code-review`, `/review:security-review` | Description phrase + `## Boundary` section for the bundled `code-review` skill and the plugin-backed built-in `security-review` command (verdict `complementary`, CI lane versus session pass); pointer records in each skill's `reference/` file | | `/code-tidying:tidy`, `/code-tidying:batch-simplify` | `## Boundary` sections for the bundled `simplify` skill (verdict `complementary`, diff-anchored versus lane- and sweep-anchored); detail in each skill's reference or context file; no description phrase | | `/testing:run-e2e` | description phrase, `## Boundary` section and `## Native step` for the bundled `run` skill (verdict `complementary`, integration `wrap`; a look versus evidenced verification); detail in the skill's context file | | `/harness-ops:audit-performance`, `/harness-ops:audit-skill-visibility` | `## Boundary` sections for the bundled `doctor` skill (and `/skill-doctor` for the second), verdict `complementary`; the second also carries the description phrase for `/skill-doctor`; detail in each skill's `reference/` file | @@ -302,10 +310,10 @@ A `suggest` row addresses the person, not the model. The sentence shape is: If / is available in your session (), run it for . ``` -`` is a same-file four-part verification record per upstream-drift naming the surface's -own gate. For `/doctor` that gate is `DISABLE_DOCTOR_COMMAND` or a `skillOverrides` entry, -because `/doctor` survives `disableBundledSkills`. For every other bundled skill the gate -includes `disableBundledSkills` as well. Place the sentence at the start of the run when the +`` names the surface's own gate and links a same-file upstream-drift record for it (our +decision, a pointer to the section defining the gate, the as-of date, the recheck trigger). For +`/doctor` the basis names `DISABLE_DOCTOR_COMMAND` or a `skillOverrides` entry and not +`disableBundledSkills`; for every other bundled skill it names `disableBundledSkills` as well. Place the sentence at the start of the run when the surface covers everything the skill does, and at the end when coverage is partial. A model-disabled bundled skill is suggested as `/` exactly as a built-in command is. The wording is "reserved for the person to run", never "cannot be invoked" as an absolute. An @@ -337,4 +345,5 @@ change; the doc's README-only original state reads as 1.0. Upstream publishes no convention for deferring to its own surfaces (absence checked 2026-08-23 against the pages listed above and `https://code.claude.com/docs/llms.txt`), which is why this -repository owns one. +repository owns one. Recheck trigger: a Claude Code docs page publishes guidance on how a plugin +should refer to native surfaces. diff --git a/docs/conventions/recommendation-basis/CHANGELOG.md b/docs/conventions/recommendation-basis/CHANGELOG.md index 20bc473852..43d63714f3 100644 --- a/docs/conventions/recommendation-basis/CHANGELOG.md +++ b/docs/conventions/recommendation-basis/CHANGELOG.md @@ -4,6 +4,12 @@ Notable changes to the recommendation-basis contract (SemVer). Changing the grou label's values, or the re-emit shape is a major bump; additive guidance is a minor bump; docs-only clarification is a patch. +## [1.0.2] - 2026-10-01 + +Patch, docs-only. The Boundary bullet on durable records of upstream-derived facts names the record +the upstream-drift convention now requires: our decision, a pointer, an as-of date and a recheck +trigger. No grounding bar, label value or re-emit shape changes. + ## [1.0.1] - 2026-09-28 Adopters lists the skills that conform to 1.0.0 across `planning`, `source-control`, `github`, diff --git a/docs/conventions/recommendation-basis/README.md b/docs/conventions/recommendation-basis/README.md index 3f28820d98..067c8fc910 100644 --- a/docs/conventions/recommendation-basis/README.md +++ b/docs/conventions/recommendation-basis/README.md @@ -34,9 +34,10 @@ re-statement. It does not own: which in turn defers to the contract `/discovery:research` states. This doc uses those terms and does not redefine them. - **Durable records of upstream-derived facts.** A recommendation written into a committed file as - a standing decision carries the four-part record the - [upstream-drift convention](../upstream-drift/README.md) defines. The `Basis:` label covers what - is said to the user in session; the stamp covers what is stored. + a standing decision carries the record the + [upstream-drift convention](../upstream-drift/README.md#required-parts) requires: our decision, a + pointer, an as-of date and a recheck trigger. The `Basis:` label covers what is said to the user + in session; the record covers what is stored. ## What counts as a recommendation diff --git a/docs/conventions/upstream-drift/CHANGELOG.md b/docs/conventions/upstream-drift/CHANGELOG.md index 382d8b3d13..ab29506e7a 100644 --- a/docs/conventions/upstream-drift/CHANGELOG.md +++ b/docs/conventions/upstream-drift/CHANGELOG.md @@ -4,6 +4,21 @@ Notable changes to the upstream-drift contract (SemVer). Changing a required par name, or an enforceability verdict is a major bump; additive guidance is a minor bump; docs-only clarification is a patch. +## [2.0.0] - 2026-10-01 + +Major under this contract's own rule: the required parts change. + +A conforming record now stores no upstream text, quoted or paraphrased. Its parts are our decision +in our own words, a pointer to the exact upstream section, an as-of date and a recheck trigger. The +restated claim and its basis are gone. A probed behavior points at the probe and may state what it +observed; a source conflict is recorded only as "pages X and Y disagree on topic T"; and one named +exception, "Old-patterns mapping tables", admits an old-to-current name table inside a skill's +"Old patterns" section where the skill-authoring guidance recommends one. The fetch +route's "No verbatim quote, no claim" rule becomes "No read, no verdict": the matched span stays in +the run's working data, never in the record. The worked instances no longer quote upstream pages. +The Adopters rows are restated for the new shape, and records still in the 1.x shape are tracked in +[#5684](https://github.com/melodic-software/claude-code-plugins/issues/5684). + ## [1.7.0] - 2026-09-30 Additive guidance; minor under this contract's own rule. No required part, canonical name, or diff --git a/docs/conventions/upstream-drift/README.md b/docs/conventions/upstream-drift/README.md index 06861f9a10..e2e19f0771 100644 --- a/docs/conventions/upstream-drift/README.md +++ b/docs/conventions/upstream-drift/README.md @@ -17,8 +17,9 @@ Owner doc for **how this repository records a fact or decision derived from a source it does not own**, whether an official doc page, an upstream issue thread, or a probed platform behavior, so the record -stays honest as the upstream moves. One name and one shape: a dated **verification stamp** paired -with a **recheck trigger**, the stated observable event that obliges re-deriving the record. +stays honest as the upstream moves. One name and one shape: our decision in our own words, a +**pointer** to the exact upstream section, an **as-of date**, and a **recheck trigger**, the stated +observable event that obliges re-deriving the record. The upstream text itself is never stored. The fleet previously practiced this in five-plus places under four names: "recheck triggers" ([hook-config-delivery](../hook-config-delivery/README.md)), "revisit triggers" @@ -58,35 +59,68 @@ the upstream form list is the org standard's to change. ## A date is never authority -A dated verification stamp is an **as-of record**: it tells the reader when the claim last matched -its source, and nothing more. It never confers standing authority. A stale stamp reads identically +A dated verification stamp is an **as-of record**: it tells the reader when the decision was last +derived from its source, and nothing more. It never confers standing authority. A stale stamp reads identically to a fresh one, and upstream surfaces move without notice: Claude Code changes its own conventions between releases, sometimes with no version signal on the surface in question, and experimental surfaces churn outright. The part of the record that matters is therefore the **trigger**, not the -date: anything restating a volatile upstream specific carries a stated re-derivation event, or it -is drift waiting to happen. Before acting on any stamped claim, re-fetch the cited basis. The -stamp is the ceiling on how current the claim can be, never a guarantee. +date: anything depending on a volatile upstream specific carries a stated re-derivation event, or +it is drift waiting to happen. Before acting on any record, re-read the section its pointer names. +The as-of date is the ceiling on how current the decision can be, never a guarantee. The discipline covers two record kinds, one shape: -- a **verified-fact stamp**: a restated upstream specific ("verified 2026-07-17 against \"); +- a **dependent decision**: something this repository does because of an upstream specific, with + the pointer saying where that specific lives; - a **recorded decision**: a deferral or rejection derived from upstream facts as they stood on a date, whose premises can rot the same way the facts can. ## Required parts -A conforming record carries four parts: - -1. **The claim or decision**: what exactly was verified, or what was decided and on what premise. -2. **The basis**, the specific source it was derived against: the official page URL (with anchor - where one exists), the upstream issue, or the probe/method for an empirical finding. "Verified" - with no stated basis is not re-checkable. -3. **The as-of date**: when the derivation happened. +A conforming record stores no upstream text, quoted or paraphrased, not even one line (the one +named exception is [old-patterns mapping tables](#old-patterns-mapping-tables)). It carries: + +```markdown + +- **Pointer**: for , see . +- **As of**: YYYY-MM-DD +- **Recheck trigger**: +``` + +1. **The decision**: what this repository does or decided, in our own words. It may name the topic + the upstream page covers; it never states what the page says about it. +2. **The pointer**: the official page URL with the anchor of the exact section, or the upstream + issue. A reader who needs the specific reads it there, live. A blog post is never the pointer + where a main docs section covers the topic: it appears only as a "correlate with \" + note beside that pointer. Where no docs page covers it yet, the record links the post as that + note, says no docs page covers the topic as of the date, and its recheck trigger is a docs page + starting to cover it, at which point the pointer moves there. +3. **The as-of date**: when the decision was last derived from the page. 4. **The recheck trigger**: the observable event that obliges re-derivation. -Prefer the pointer: where a surface can defer to the live source at read time, cite it and restate -nothing. Then no stamp is needed at all. The four-part record is the fallback for surfaces that -must restate a volatile specific to function. +Two cases have their own form: + +- **A probed behavior** has no upstream page. The pointer names the probe (the script, pull request + or issue holding its evidence), and the decision may state what the probe observed, in our words: + the observation is ours, not the upstream's. +- **A source conflict** is recorded only as "pages X and Y disagree on topic T", with both links, + the as-of date and a trigger. Neither page's position is restated. + +When a surface needs the specific at run time, it fetches it from the pointer +([the fetch route](#reading-the-basis-the-fetch-route)). A catalog row that must fire without a +live fetch keeps its firing rule in our own words; the rule is our decision, not the page's text. + +### Old-patterns mapping tables + +One named exception admits upstream names into a file. A skill may carry a table mapping old API +or interface names to their current ones, but only inside an "Old patterns" section placed where +the skill-authoring guidance recommends one. The table holds names only, never descriptions of +behavior, and the section carries the record parts below it. Everywhere else the rule above holds. + +- **Pointer**: for where the guidance recommends an "Old patterns" section, see + . +- **As of**: 2026-10-01 +- **Recheck trigger**: that section stops recommending an "Old patterns" section, or moves. ## The observability bar @@ -100,24 +134,25 @@ qualify: a trigger whose firing cannot be checked is a date with extra words. A firing is a record-maintenance event, and the procedure follows what the trigger guards: -- **A four-part record.** Re-fetch the cited basis and re-derive the claim or decision from what is - actually there, never patching the record from memory. Refresh the as-of date **with the outcome**, - drift or no drift. On a versioned surface a drift outcome lands as a changelog entry; refreshing - a date with no verdict change is no entry and no version bump. +- **A pointer record.** Re-read the section the pointer names and re-derive the decision from what + is actually there, never patching the record from memory. A section that moved gets a new + pointer. Refresh the as-of date **with the outcome**, drift or no drift. On a versioned surface + a drift outcome lands as a changelog entry; refreshing a date with no verdict change is no entry + and no version bump. - **A named trigger guarding an in-repo decision** ([Adopters](#adopters) says which rows these - are). There is no cited basis to re-fetch and no as-of date to refresh: re-derive the decision + are). There is no pointer to re-read and no as-of date to refresh: re-derive the decision from the state the trigger names. The decision guarded is in-repo; the firing event can live anywhere, upstream included. Record the outcome durably where the decision lives: the record - itself or the owning surface's changelog. A re-derivation that ends up restating an upstream - specific adopts the four required parts in the refreshed record. The durable outcome is the part - this kind shares with the stamped kind. + itself or the owning surface's changelog. A re-derivation that ends up depending on an upstream + specific adopts the required parts in the refreshed record. The durable outcome is the part + this kind shares with the pointer kind. Whichever the kind, where re-derivation finds drift the changed value lands in the owning record, never silently in a consuming surface. ### Read-time validation is not a firing -The standing rule to re-fetch a cited basis before acting on a stamped claim +The standing rule to re-read a record's pointer before acting on it ([a date is never authority](#a-date-is-never-authority)) is per-use validation: it protects the act, not the record, and a lookup that finds no drift obliges no edit anywhere. A record kept current this way states divergence as its trigger, and "a read-time re-fetch finds the source no @@ -126,7 +161,7 @@ lookup, is what fires, and only a firing invokes the maintenance procedure above ## Reading the basis: the fetch route -Re-fetching a cited basis is the first step of every firing above, so **how** the page is read is +Re-reading the section a pointer names is the first step of every firing above, so **how** the page is read is part of the contract. A summarizing fetch of a long docs page is not a read of that page: it truncates, and a summarizer asked what the page contains then answers from the truncated span. That answer is indistinguishable from a genuine absence, so a truncated fetch does not merely fail. It @@ -136,8 +171,9 @@ rows missing ([#2182](https://github.com/melodic-software/claude-code-plugins/pu Three rules bind every read, whichever rung it comes from: -- **No verbatim quote, no claim.** A record's basis is the text, not a paraphrase of it. A verdict - of "current" states the quoted span it matched. +- **No read, no verdict.** A verdict rests on the text read, not on a paraphrase or a summary of + it. A verdict of "current" names the span it matched in the run's working data; the span never + enters the record. - **A truncated read supports no absence claim, ever.** If the fetch stops short, say so and mark the item unverified. "Not in the response" is never "not on the page". The reader cannot tell those apart, which is the entire failure this rung ladder exists to prevent. @@ -164,11 +200,10 @@ types. The Anthropic profile is the default. A page whose channel does not resol declares is recorded unread with a reason, so the reader drops a rung and says so; the script never falls back to another channel itself. -Rung 1 is the default. It was verified against `env-vars` on 2026-08-10: `curl` returned -`text/markdown`, 361,797 bytes over 458 lines carrying 315 variable rows including the full -`CLAUDE_CODE_MAX_*` range, and two fetches seconds apart hashed identically -(SHA-256 `43a805b4cfffd9aae5e36cec42f3a271dc92ddead26db76cd401d61ff4048584`). That same fetch -re-confirmed the header finding below: `Last-Modified` came back equal to `Date`. +Rung 1 is the default. A probe against `env-vars` on 2026-08-10 showed the raw channel returning +the whole page, including the rows the summarizing fetches had dropped, and two fetches seconds +apart hashing identically. That same fetch re-confirmed the header finding below: `Last-Modified` +came back equal to `Date`. The route is not new here; it is **hoisted from two surfaces that each derived it independently**. `/harness-ops:changelog`'s read-actions context carried it page-scoped ("`curl` the @@ -189,12 +224,11 @@ page it is reading and drops a rung when it does not resolve. A rung-1 fetch can return `200`, `text/markdown`, and a complete untruncated body that is **someone else's page**. A retired slug is silently aliased to its successor: no redirect, no -`Location` header, no notice in the body. Verified 2026-08-11: -`https://code.claude.com/docs/en/slash-commands.md` returns `200` with 82,668 bytes whose first -heading is `# Extend Claude with skills`, **byte-identical to `skills.md`** (both SHA-256 -`a833dd5c96b9b111de0daec5fc6436e210c8cdc009e51306d32438746db0b5a5`), while the rendered URL reports -`0` redirects. This is not a catch-all: an invented slug (`nonexistent-page-xyz.md`) returns a clean -`404`, so the alias is specific to slugs that once existed. +`Location` header, no notice in the body. A probe on 2026-08-11 found +`https://code.claude.com/docs/en/slash-commands.md` returning `200` with a body **byte-identical to +`skills.md`**, while the rendered URL reported `0` redirects. This is not a catch-all: an invented +slug (`nonexistent-page-xyz.md`) returned a clean `404`, so the alias is specific to slugs that once +existed. The failure this produces is worse than truncation, because truncation at least yields text you can see is short. Here a search for a term the *requested* page owns comes back empty against a full, @@ -209,10 +243,9 @@ Two checks, both cheap, and a run does them before it trusts a body: alias. Verified across ten slugs on 2026-08-11: the nine live ones each appear as `docs/en/.md`; `slash-commands` appears in no such entry (only an unrelated `agent-sdk/slash-commands`), which is exactly the one that aliased. -- **Read the body's own first heading before quoting it.** `skills.md` and a live `.md` both - say what they are on line 5. A heading that does not match the page you asked for ends the read; - a title that merely differs in wording from the slug does not (`sub-agents.md` is titled "Create - custom subagents", `costs.md` "Manage costs effectively", both correct). +- **Read the body's own first heading before trusting it.** A page says what it is in its first + heading. A heading for a different page than the one you asked for ends the read; a title that + merely differs in wording from the slug does not. A slug missing from `llms.txt` is not automatically a dead end: it may have been renamed, and the index is the place to find the successor. Fetch the successor and cite **that** slug, rather than @@ -235,13 +268,11 @@ it, and both produce a claim that reads as researched: `hooks`", or, if the sweep really covered the index, "not documented on any page listed in `llms.txt` as of ``", which is a much larger and much more expensive claim. - **Searching the phrase instead of the capability.** A literal string can be absent while the - thing it names is documented in other words on the same page. Worked instance, verified - 2026-08-11 on `hooks.md`: the phrase "verbose hooks" appears **zero** times, yet the page itself - documents "Async hook completion notifications are suppressed by default. To see them, enable - verbose mode with `Ctrl+O` or start Claude Code with `--verbose`", and separately - "set `CLAUDE_CODE_DEBUG_LOG_LEVEL=verbose` to see additional log lines such as hook matcher - counts and query matching". A phrase search would have returned nothing and licensed "no verbose - hooks toggle exists". That is false, from a complete, untruncated read of the right page. + thing it names is documented in other words on the same page. Worked instance, 2026-08-11 on + [`hooks`](https://code.claude.com/docs/en/hooks): the phrase "verbose hooks" appeared **zero** + times, yet the page documented two separate ways to see more hook output, in other words. A + phrase search would have returned nothing and licensed "no verbose hooks toggle exists". That + was false, from a complete, untruncated read of the right page. So an absence claim states the corpus and the terms tried, and a claim that a *capability* is missing searches the capability's plausible vocabulary, not one phrasing of it. This bit the fleet @@ -258,9 +289,9 @@ convenient. A mirror read is admissible only when it is **verbatim** and its currency is **corroborated against the page's own content**, never against the mirror's self-reported sync time alone, which is a claim by the party whose freshness is in question. The corroboration names a fact that only a sync -later than some known upstream change could carry, and the record states it. The worked instance: -`ericbuess/claude-code-docs` `docs/env-vars.md` was accepted because it carried the v2.1.224 -removal of the 200-subagent-per-session cap, which no pre-v2.1.224 sync can contain. +later than some known upstream change could carry, and the record names that release. The worked +instance: `ericbuess/claude-code-docs` `docs/env-vars.md` was accepted because it carried a change +from Claude Code v2.1.224, which no earlier sync can contain. A record resting on a mirror **says on its face that it is one rung below a primary read**, and states retirement of that basis as part of its trigger: a later primary read of the same range @@ -324,8 +355,10 @@ every correct citation and every in-repo mention alike, and a gate whose false-p routine suppression trains authors to bypass it. That is worse than no gate, because it converts a real signal into noise with an approved silencer. -- **Basis**: `melodic-software/standards` `conventions/engineering/enforceability-tiers.md` (the - reasoning-only tier and the worth-mechanizing routing rule), plus the worked instance above. +- **Pointer**: for the reasoning-only tier and the worth-mechanizing routing rule, see + `melodic-software/standards` `conventions/engineering/enforceability-tiers.md`; the worked + instance is above. +- **As of**: 2026-08-12 - **Recheck trigger**: a third unstamped upstream-fact carrier reaches `main` after this decision (two are already on the record: `plugin-quality`, corrected in its 0.4.0, and `architecture`, whose false claim was removed in its 0.5.1), **or** a detector is demonstrated that separates an @@ -346,32 +379,38 @@ carriers are recorded that way in [#2297](https://github.com/melodic-software/claude-code-plugins/issues/2297). The rows are not all the same thing, and the table says which is which. A **conforming record** -carries the four required parts for an upstream-derived claim or decision. A **named trigger** -shares the canonical name, the observability bar, and +carries the [required parts](#required-parts) for a decision that depends on something +upstream-owned. A **named trigger** shares the canonical name, the observability bar, and [its own firing procedure](#when-a-trigger-fires), but guards an in-repo decision: in scope for the -name, outside the four-part requirement, which binds only records that restate something -upstream-owned. This narrows what a row advertises; it does not widen the -contract to fit its exceptions. +name, outside the pointer requirement, which binds only records that depend on something +upstream-owned. This narrows what a row advertises; it does not widen the contract to fit its +exceptions. + +At 2.0.0 the required parts changed from a restated claim plus its basis to our decision plus a +pointer. The rows below were converted with that release. Records elsewhere still in the 1.x +four-part shape are tracked in +[#5684](https://github.com/melodic-software/claude-code-plugins/issues/5684) and adopt the new shape +on touch. | Surface | Was | What a reader can rely on | |---|---|---| -| [hook-config-delivery](../hook-config-delivery/README.md) §Recheck triggers | already the canonical name | Conforming records: version-pinned facts table with per-fact basis, a per-row verified version and date, and fact-scoped event triggers. | -| [loop-lane](../loop-lane/README.md) §Versioning | "Re-derivation triggers" | Conforming records: dated upstream-claim stamps; drift outcomes recorded in its changelog. | -| [plugin-philosophy](../../plugin-philosophy.md) component-stances staleness disclaimer | unlabeled discipline | Conforming records: per-row claim, linked page, and verified date; the re-fetch-before-acting rule is [read-time validation](#read-time-validation-is-not-a-firing), and every row's stated trigger is a fetch diverging from the row. | -| [plugin-philosophy](../../plugin-philosophy.md#recorded-gate-runs) recorded gate runs | new with this table | Conforming records of the second kind: **recorded decisions**, one per platform surface the Native-first adoption gate has been run against, carrying an adopt/defer/decline verdict, the quoted upstream basis it rests on, and a trigger written per row rather than the generic divergence-at-fetch. A verdict is re-derived when its own trigger fires, not on any fetch that differs. | -| [official-docs](../../official-docs.md) staleness warning and per-row verified dates | unlabeled discipline | Conforming records: same shape as the component-stances table: link + date, divergence-at-fetch as the stated trigger. | -| [migration-playbook](../../migration-playbook.md) decision records | "Revisit trigger", and "Re-trigger" on the plugin-acceptance review record | Mixed: the dated component-decision records cite upstream bases and conform; the org-internal records (e.g. the ratification and plugin-acceptance review records) are named triggers; the skill-quality retrofit record is a third kind, terminal exclusions that state "no recheck trigger" by design, decided out, so nothing fires. | -| [ecosystem-commands](../ecosystem-commands/README.md) task-runner deferral | "Revisit triggers" | Named triggers only: an undated in-repo deferral; not a four-part record. | -| `/ai-slop:audit`, the tell catalog it loads, §Upstream-drift record | new with 1.5.0 | Conforming record: revision-pinned four-part record over the Wikipedia source page (claim, `oldid` basis, as-of date, recurring recheck trigger: each `ai-slop` release and each fleet audit, chosen over per-revision after measuring the page at 50+ edits/week), plus a recorded fetch-gap note for two source sections the same trigger covers. | -| `/docs-hygiene:write-for-humans`, the source records it loads | new with docs-hygiene 0.18.0 | Conforming records: one four-part record per external writing standard the skill falls back to (Diátaxis, Google developer documentation style, ASD-STE100, Global English), each carrying claim, basis, as-of date, and an observable recheck trigger. Three are publication events (an STE issue, a Global English edition, a Diátaxis revision); the Google record's is a page-content divergence, because that guide is a continuously-edited site with no edition to pin. The contract admits either shape, and the record names which one it is. The STE record additionally states a fidelity ceiling: the layer is a principles subset, not the specification, so a document written to it is not thereby STE-conformant. | +| [hook-config-delivery](../hook-config-delivery/README.md) §Recheck triggers | already the canonical name | Conforming records: per-fact pointer, a per-row verified version and date, and fact-scoped event triggers. | +| [loop-lane](../loop-lane/README.md) §Versioning | "Re-derivation triggers" | Conforming records: dated pointers; drift outcomes recorded in its changelog. | +| [plugin-philosophy](../../plugin-philosophy.md) component-stances staleness disclaimer | unlabeled discipline | Conforming records: per-row decision, linked section, and as-of date; the re-read-before-acting rule is [read-time validation](#read-time-validation-is-not-a-firing), and every row's stated trigger is a read diverging from the row. | +| [plugin-philosophy](../../plugin-philosophy.md#recorded-gate-runs) recorded gate runs | new with this table | Conforming records of the second kind: **recorded decisions**, one per platform surface the Native-first adoption gate has been run against, carrying an adopt/defer/decline verdict, a pointer to the upstream section it rests on, and a trigger written per row rather than the generic divergence-at-read. A verdict is re-derived when its own trigger fires, not on any read that differs. | +| [official-docs](../../official-docs.md) staleness warning and per-row verified dates | unlabeled discipline | Conforming records: same shape as the component-stances table: link + date, divergence-at-read as the stated trigger. | +| [migration-playbook](../../migration-playbook.md) decision records | "Revisit trigger", and "Re-trigger" on the plugin-acceptance review record | Named triggers only: the org-internal records (e.g. the ratification and plugin-acceptance review records) are named triggers; the skill-quality retrofit record states "no recheck trigger" by design, decided out, so nothing fires. The dated component-decision records keep the 1.x shape and are tracked in #5684. | +| [ecosystem-commands](../ecosystem-commands/README.md) task-runner deferral | "Revisit triggers" | Named triggers only: an undated in-repo deferral with no upstream pointer. | +| `/ai-slop:audit`, the tell catalog it loads, §Upstream-drift record | new with 1.5.0 | Conforming record: a revision-pinned pointer to the Wikipedia source page (`oldid`), as-of date, and a recurring recheck trigger: each `ai-slop` release and each fleet audit, chosen over per-revision after measuring the page at 50+ edits/week. | +| `/docs-hygiene:write-for-humans`, the source records it loads | new with docs-hygiene 0.18.0 | Conforming records: one pointer record per external writing standard the skill falls back to, each carrying the pointer, as-of date, and an observable recheck trigger: a publication event for an edition-pinned standard, a page-content divergence for a continuously edited one. The record names which. | Elsewhere the name binds on touch: living surfaces still saying "revisit trigger", "re-trigger", "re-derivation trigger", or "what would reopen it" (several plugin reference docs already use the canonical `## Recheck triggers` heading) adopt the canonical name, the observability bar, and their -kind's firing procedure the next time they change; a surface restating an upstream-owned specific -additionally adopts the required parts. **History is never rewritten**: `CHANGELOG.md` entries, -dated audit records, and ADR sections keep the wording they shipped with; a new ADR uses the -canonical name going forward. +kind's firing procedure the next time they change; a surface depending on an upstream-owned +specific additionally adopts the required parts and drops any restated upstream text. **History +is never rewritten**: `CHANGELOG.md` entries, dated audit records, and ADR sections keep the +wording they shipped with; a new ADR uses the canonical name going forward. ## Why this name diff --git a/docs/finding-your-unknowns.md b/docs/finding-your-unknowns.md index 3b5d8d7914..390510bf46 100644 --- a/docs/finding-your-unknowns.md +++ b/docs/finding-your-unknowns.md @@ -10,11 +10,10 @@ pattern catalog and the boundaries (when HTML, when not; what deliberately stays un-codified). Sibling docs: `plugin-philosophy.md` (governance), `glossary.md` (vocabulary), `migration-playbook.md` (delivery). -**Sources and permission basis.** The material derives from public posts by their named -author (see [Sources](#sources-and-citation-shape)). This doc quotes short attributed -verbatim excerpts under fair-quotation practice; no license is claimed and bulk -reproduction is avoided. Quotes are reproduced exactly as published, punctuation -included, and are never edited to fit this repo's style rules. +**Sources.** The material derives from public posts by their named author (see +[Sources](#sources-and-citation-shape)). This doc stores no text from them, quoted or +paraphrased: it records what this repository adopted, in our words, and points at the source +section each decision rests on, which a reader opens to read the author's own words. ## Contents @@ -33,29 +32,27 @@ included, and are never edited to fit this repo's style rules. ## Why this exists -The methodology's economic argument, in the author's words: "Every explainer, brainstorm, -interview, prototype, and reference is a cheap way to find out what you didn't know before -it gets expensive to fix." (Field guide, [Sources](#sources-and-citation-shape) S1.) Each -pass below trades a few minutes of artifact review for a class of rework. +We adopted the methodology for its economics: each artifact is a cheap way to learn something +before it becomes expensive to fix, and each pass below trades a few minutes of artifact review +for a class of rework. For the author's own argument, see S1. -Caution on the framing: the author's stronger thesis, that output quality is now -bottlenecked by the human's ability to clarify the model's unknowns, is a single -practitioner's vendor-published claim and is treated here as direction, not doctrine. +Caution on the framing: the author's stronger thesis about where output quality is now +bottlenecked (S1) is a single practitioner's vendor-published claim and is treated here as +direction, not doctrine. ## The unknowns taxonomy -Four quadrants, asked as "what are your unknowns?" before prompting: +We ask "what are your unknowns?" before prompting, across four quadrants: -- **Known knowns**: what the prompt already states. -- **Known unknowns**: questions you know to ask but haven't answered yet. -- **Unknown knowns**: things you assume without realizing you're assuming them; the - agent can't see them until you disclose them. -- **Unknown unknowns**: the pothole you didn't know the road could have; only an - artifact that shows you the terrain surfaces these. +- **Known knowns**: what the prompt already says. +- **Known unknowns**: open questions you are aware of. +- **Unknown knowns**: assumptions you hold without noticing; the agent cannot see them until + you state them. +- **Unknown unknowns**: risks you have no reason yet to look for; only an artifact that shows + the terrain brings these out. -The draft article's quadrant taglines ("questions you know to ask", "the pothole you -didn't know the road could have") appear only in the X draft (S4), which is the citable -source for draft-only content. +The author's own quadrant taglines appear only in the X draft (S4), which is the citable source +for draft-only content. Findings that surface during an unknowns pass fall into four types (adopted as `discovery:blindspot`'s output taxonomy): **Landmine** (a change that will break @@ -65,17 +62,15 @@ longer shows), **Convention** (an unwritten team rule the work must follow), and Two diagnostics ride the taxonomy: -- Over-specifying and under-specifying are the same failure seen from two sides: both - mean the split between what you locked and what you left open didn't match your actual - unknowns. -- When a long-horizon task comes back wrong, check the unknowns and the plan's +- We treat over-specifying and under-specifying as one failure: in both, the split between + what you locked and what you left open did not match your actual unknowns. +- When a task that ran for hours returns a wrong result, we check the unknowns and the plan's adaptability before blaming the model: the usual root cause is an unknown that was never surfaced, not a capability gap. -The lifecycle is a loop: what an artifact teaches you becomes the starting map for the -next round. The author frames this as matching the map to the territory (S1, "Matching -map and territory"), cited here as his metaphor, not adopted as house vocabulary (see -`glossary.md` rejected terms). +We run the method as a loop: what an artifact teaches becomes the starting map for the next +round. The author's own framing of this loop (S1, section "Matching map and territory") is his +metaphor, not adopted as house vocabulary (see `glossary.md` rejected terms). ## The five-pass pre-implementation workflow @@ -101,7 +96,7 @@ independently corroborates it; see `session-flow` plugin). ## Prompt-pattern catalog -Patterns the corpus demonstrated that have no owning skill; each entry is one canonical +Patterns the corpus demonstrated that have no owning skill; each entry is one house prompt-line to adapt. Patterns with an owning skill are listed in the [workflow](#the-five-pass-pre-implementation-workflow) above. Invoke the skill instead. @@ -120,9 +115,9 @@ prompt-line to adapt. Patterns with an owning skill are listed in the - **Quiz me before I merge**: served by `/education:quiz-me`; the merge gate itself stays with `/verification:confirm` (one mechanism per concern). -Reconciliation note: the corpus's "tweakable plan" ordering (high-tweak decisions first, -mechanical work collapsed) is already `planning:plan`'s documented presentation default; -it needed no new mode here. +Reconciliation note: the corpus's tweakable-plan ordering is already `planning:plan`'s +documented presentation default (high-tweak decisions first, mechanical work collapsed); it +needed no new mode here. ## Reply-affordance convention @@ -143,13 +138,10 @@ validation answer set). Fleet audits check those surfaces against this section. ## Export-button rule -**The rule.** An interactive HTML artifact always ends with an export affordance that -turns UI state back into something the user can paste or commit. In the author's words: -"The trick is always to end with an export: a "copy as JSON" or "copy as prompt" button -that turns whatever I did in the UI back into something I can paste into Claude Code." -(S2, "Custom editing interfaces".) The doctrine recurs three times independently in the -corpus; it is what keeps a throwaway editor inside the agent loop instead of becoming a -dead end. +**The rule.** Every interactive HTML artifact our skills emit ends with a control that +copies the state the user built out as text they can paste into the session or commit. For the author's version of this rule, see S2, section "Custom editing +interfaces". The doctrine recurs three times independently in the corpus; it is what keeps a +throwaway editor inside the agent loop instead of becoming a dead end. **Who is bound.** Skills that emit interactive HTML artifacts cite this section. @@ -173,37 +165,34 @@ registry row per `plugin-philosophy.md` "Convention registry". ## When HTML, and when not -The corpus's examples index (S3) organizes twenty demos into nine categories: -exploration and planning, code review and understanding, design, prototyping, -illustrations and diagrams, decks, research and learning, reports, and custom editing -interfaces. Those categories double as the "when is HTML worth it" taxonomy: reach for a -rendered page when the information is spatial (diffs, call graphs), comparative -(side-by-side directions), interactive (motion you can only feel), or recurring (reports -that benefit from structure and color). +The corpus's examples index (S3) groups its demos by category. Our test for when HTML is worth +it: reach for a rendered page when the information is spatial (diffs, call graphs), +comparative (side-by-side directions), interactive (motion you can only feel), or recurring +(reports that benefit from structure and color). - **Density rubric**: HTML earns its cost through tables, CSS, SVG, interaction, and spatial layout. Markdown pushed past its density limit produces the degraded workarounds (ASCII diagrams, unicode color) that signal you wanted a page. -- **Reading ceiling**: the author's ~100-line markdown ceiling is a practitioner - anecdote, recorded as such, not a measured threshold. -- **Sharing**: the publish-and-share argument is satisfied in this environment by the - Artifact tool; nothing extra to build. +- **Reading ceiling**: the author's markdown length ceiling (S2) is a practitioner + anecdote, recorded as such, not a measured threshold; we set none. +- **Sharing**: the publish-and-share need is met in this environment by the Artifact tool + (see [Share session output as artifacts](https://code.claude.com/docs/en/artifacts)); + nothing extra to build. - **Scoping rule**: HTML artifacts are for ephemeral and published outputs. They never - replace version-controlled instruction surfaces. HTML diffs are noisy (the author's - own admission) and generation costs 2-4x the markdown equivalent, so plans, skills, - and docs stay markdown in git. + replace version-controlled instruction surfaces. HTML diffs are noisy and generating HTML + costs more than the markdown equivalent (for the author's own estimate, see S2), so plans, + skills, and docs stay markdown in git. ## The buy-in pattern -For work that needs stakeholder agreement, the corpus's buy-in document has five -sections: demo first; the pitch; pre-answered objections; spec at a glance; risk and -rollback with named per-person asks and a deadline. The pre-answered-objections element -is the industry-standard core: Amazon's PR/FAQ carries an internal FAQ anticipating hard -leadership questions (Bezos 2017 shareholder letter; Bryar & Carr's Working Backwards), -and every surveyed RFC process requires drawbacks/alternatives-considered sections: Rust -RFCs, Oxide RFDs, Google design docs, Uber-style RFCs. In all of those orgs the -persuasion artifact and the decision record are one document with a lifecycle, which is -why this repo extends existing planning artifacts rather than minting a parallel one. +For work that needs stakeholder agreement, we use a buy-in document whose core is pre-answered +objections (for the corpus's version, see S2): demo first; the pitch; pre-answered objections; +spec at a glance; risk and rollback with named per-person asks and a deadline. Pre-answered +objections are the industry-standard core: Amazon's PR/FAQ and every surveyed RFC process (Rust +RFCs, Oxide RFDs, Google design docs, Uber-style RFCs) carry the same element; see the buy-in +grounding in [Sources](#sources-and-citation-shape). In all of those orgs the persuasion +artifact and the decision record are one document with a lifecycle, which is why this repo +extends existing planning artifacts rather than minting a parallel one. **Objection-evidence checklist** (reusable in PR descriptions): for each objection you expect, write the question, the factual answer, and the evidence citation, before @@ -212,30 +201,21 @@ through the [workflow](#the-five-pass-pre-implementation-workflow). ## Cautions from the source author -The corpus carries its own warning against exactly the move a plugin marketplace is -tempted to make, and this repo treats it as binding (it is why the deltas that landed are -judgment-preserving contract lines and doc entries, never generator skills): - -> I’m a little bit afraid that people will read this article and turn it into a /html -> skill or something. While there might be some value in that, I want to emphasize that -> you don’t need to do much to get Claude to do this. You can just ask it to “make a HTML -> file” or “make a HTML artifact”. -> -> The trick is knowing what you want the artifact to do and how you might use it. You may -> over time make a skill, but for now I’d suggest just prompting from scratch to get a -> hang of how to use it in different cases. (S2, "How to Get Started".) +The author warns against exactly the move a plugin marketplace is tempted to make: turning the +method into a dedicated generator skill instead of prompting for the artifact directly (S2, +section "How to Get Started"). This repo treats that caution as binding, which is why the deltas +that landed are judgment-preserving contract lines and doc entries, never generator skills. Two companions to the warning: -- **Stay in the loop** is the evaluation lens for any artifact tooling: "All of the above - is to say that I think the real reason I use HTML is that I feel much more in the loop - with Claude." (S2, "Stay in the Loop".) Tooling that produces artifacts the user never +- **Stay in the loop** is our evaluation lens for any artifact tooling (for the author's + framing, see S2, section "Stay in the Loop"). Tooling that produces artifacts the user never forms judgment about fails this criterion even when it satisfies density, sharing, and ease. -- **Throwaway-editor doctrine**: a custom editing interface is "not a product, or a - reusable tool". It is built for the exact thing being worked on and discarded. The - marketplace instinct to generalize a good throwaway into a shipped generator is the - failure mode the warning names. +- **Throwaway-editor doctrine**: we build a custom editing interface for the exact thing being + worked on and discard it; it is never a product or a reusable tool. The marketplace instinct + to generalize a good throwaway into a shipped generator is the failure mode the warning + names. ## Heuristics awaiting evidence @@ -259,11 +239,13 @@ they graduate into a skill body only on observed, repeated stumble evidence: Citations in this doc use: URL, ISO retrieval date, and `sha256:` over the raw snapshot bytes captured at retrieval. Content drift produces a new citation, never an -in-place hash edit. +in-place hash edit. Recheck trigger: a re-retrieval whose hash differs from the recorded one. +No Claude docs page covers this methodology as of 2026-10-01, so the author's posts stay the +sources, each read at its link. - **S1**: "A field guide to Claude Fable 5: Finding your unknowns", Thariq Shihipar, - Anthropic blog, published 2026-07-06. - `https://claude.com/blog/a-field-guide-to-claude-fable-finding-your-unknowns` + Anthropic blog, published 2026-07-06. No docs page covers it: + (correlate with `https://claude.com/blog/a-field-guide-to-claude-fable-finding-your-unknowns`) (retrieved 2026-09-01, `sha256:ac8229699555d38eb0dfe6c80dd2e85353f30471a7abff0894d342b5107aad26`) - **S2**: "Using Claude Code: The Unreasonable Effectiveness of HTML", X article by the diff --git a/docs/official-docs.md b/docs/official-docs.md index 3b501ec2f7..bfb09357e7 100644 --- a/docs/official-docs.md +++ b/docs/official-docs.md @@ -11,11 +11,11 @@ training-data recall. > prior fetch. The authoritative, self-updating master list is > [`https://code.claude.com/docs/llms.txt`](https://code.claude.com/docs/llms.txt); if a page listed > here is missing from it, or a page you need isn't listed here, treat `llms.txt` as the source of -> truth and update this file. Every row below was verified against a live fetch on the date shown, and +> truth and update this file. Each row's **As of** date is when the page was last read live, and > that date is the ceiling on how current the row still is, not a guarantee. A fetch that no longer > matches a row is that row's recheck trigger: update the row, refreshing its date with the > outcome. The [upstream-drift convention](conventions/upstream-drift/README.md) owns this -> stamp-and-trigger discipline, and its +> record shape and trigger discipline, and its > [fetch route](conventions/upstream-drift/README.md#reading-the-basis-the-fetch-route) owns how to > read the page you re-fetch: several of these pages are long enough that a summarizing fetch > truncates them and then reports what it never reached as absent. Read the `.md` channel verbatim @@ -24,16 +24,15 @@ training-data recall. ## Plugin components → doc page One row per plugin component type, per the current [Plugins reference](https://code.claude.com/docs/en/plugins-reference). -`Commands` is the legacy flat-markdown form of a skill, and the [Skills](https://code.claude.com/docs/en/skills) -page is authoritative for both. Statusline is not its own plugin component: it is one of the two -settings keys (`subagentStatusLine`) a plugin's `settings.json` may set. Channels are declared via a -`channels` manifest field bound to an MCP server, not a separate file location. Workflows have no -per-component section in the Plugins reference. That page carries the slot in its standard-layout -and file-locations tables, and the [Workflows](https://code.claude.com/docs/en/workflows) page is -authoritative for the component. The manifest (`.claude-plugin/plugin.json`) is the container these -components are declared in, not a component, so it has no row. - -| Component | Official doc page | Verified date | +We cite the [Skills](https://code.claude.com/docs/en/skills) page for both skills and legacy +`commands/`. Statusline gets no row: we treat it as a settings key a plugin's `settings.json` may +set, not a component. Channels get a row for the `channels` manifest field, not a file location. +For workflows we cite the [Workflows](https://code.claude.com/docs/en/workflows) page. The manifest +(`.claude-plugin/plugin.json`) is the container these components are declared in, not a component, +so it has no row. For which slots and settings keys a plugin carries, see +[Standard layout](https://code.claude.com/docs/en/plugins-reference#standard-layout). + +| Component | Official doc page | As of | |---|---|---| | Skills (`skills/`) | | 2026-08-06 | | Commands: legacy flat-file skills (`commands/`) | | 2026-08-06 | @@ -41,18 +40,18 @@ components are declared in, not a component, so it has no row. | Workflows (`workflows/`) | | 2026-08-06 | | Hooks (`hooks/hooks.json`) | | 2026-08-06 | | MCP servers (`.mcp.json`) | | 2026-08-06 | -| LSP servers (`.lsp.json`) | | 2026-08-06 | +| LSP servers (`.lsp.json`) | | 2026-10-01 | | Output styles (`output-styles/`) | | 2026-08-06 | -| Themes (`themes/`) | | 2026-08-06 | +| Themes (`themes/`) | | 2026-10-01 | | Monitors (`monitors/monitors.json`) | | 2026-08-06 | | Channels (`channels` manifest field) | | 2026-08-06 | -| Executables (`bin/`) | | 2026-08-06 | +| Executables (`bin/`) | | 2026-10-01 | | Settings (`settings.json` defaults) | | 2026-08-12 | | Dependencies (`dependencies` manifest field) | | 2026-08-06 | ## Authoring -| Page | Official doc page | Verified date | +| Page | Official doc page | As of | |---|---|---| | Create plugins | | 2026-08-06 | | Plugins reference (schemas, variables, CLI) | | 2026-08-06 | @@ -75,14 +74,14 @@ components are declared in, not a component, so it has no row. | Run parallel sessions with worktrees | | 2026-08-06 | | Tools reference (includes the Monitor tool) | | 2026-08-06 | | Run agents in parallel: compares subagents, agent view, agent teams, dynamic workflows | | 2026-08-10 | -| Orchestrate agent teams: experimental, `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS` | | 2026-08-10 | -| Cross-session messaging: `ListAgents`/`SendMessage`, `crossSessionInbound`; v2.1.224+ (native Windows v2.1.234+) | | 2026-08-24 | +| Orchestrate agent teams: status and the `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS` switch | | 2026-08-10 | +| Cross-session messaging: `ListAgents`/`SendMessage`, `crossSessionInbound`, and the version floors | | 2026-08-24 | | Manage sessions: resume, branch, transcript storage | | 2026-08-10 | | Checkpointing: what `/rewind` does and does not restore | | 2026-08-10 | | Feature availability: per-feature matrix by model provider and subscription plan (not by host surface, see Platforms) | | 2026-08-10 | | Platforms and integrations: the host-surface index (CLI, Desktop, IDEs, web, mobile) | | 2026-08-10 | -| Ultrareview: human-confirmed, metered cloud review; no programmatic entry point | | 2026-08-10 | -| Chrome: browser integration delivered as the built-in `claude-in-chrome` skill | | 2026-08-10 | +| Ultrareview: the cloud review, how it is started and billed | | 2026-08-10 | +| Chrome: the browser integration and how it is delivered | | 2026-08-10 | The two `best-practices` rows share only their slug: the platform page is the cross-product Agent Skills guide for skill bodies, the Claude Code page is the harness guide for CLAUDE.md, permissions, @@ -90,10 +89,10 @@ and sessions, and they are distinct documents, so cite the one you mean by its f ## Distribution / marketplace -| Page | Official doc page | Verified date | +| Page | Official doc page | As of | |---|---|---| | Create & distribute a marketplace | | 2026-08-06 | -| GitHub Enterprise Server: marketplaces on a self-hosted instance; `owner/repo` always resolves to github.com | | 2026-08-10 | +| GitHub Enterprise Server: marketplaces on a self-hosted instance, and how `owner/repo` resolves | | 2026-08-10 | | Discover & install plugins | | 2026-08-06 | | Plugin dependencies (version constraints) | | 2026-08-06 | | Recommend plugins for your org (plugin relevance) | | 2026-08-06 | @@ -108,7 +107,7 @@ SDK-based host. ## Configuration / settings -| Page | Official doc page | Verified date | +| Page | Official doc page | As of | |---|---|---| | Settings | | 2026-08-12 | | Server-managed settings | | 2026-08-06 | @@ -128,9 +127,10 @@ and embedded sample prompts, is authored against these pages. They live on `plat master list is [`https://platform.claude.com/docs/llms.txt`](https://platform.claude.com/docs/llms.txt). -| Page | Official doc page | Verified date | +| Page | Official doc page | As of | |---|---|---| | Prompting best practices (all current models) | | 2026-08-08 | +| Prompting Claude Fable 5.1 | | 2026-10-01 | | Prompting Claude Fable 5 | | 2026-08-08 | | Prompting Claude Opus 5.5 | | 2026-09-23 | | Prompting Claude Sonnet 5.5 | | 2026-10-01 | @@ -145,14 +145,14 @@ This marketplace authors model-graded eval fixtures for its skills (see the migr eval warrant policy) and ships the `evals` plugin distilling this guidance, so the platform-side evaluation pages are plugin-relevant here alongside the prompting-doctrine rows above. -| Page | Official doc page | Verified date | +| Page | Official doc page | As of | |---|---|---| | Define success criteria and build evaluations | | 2026-08-08 | | Evals cookbook (source: `anthropics/claude-cookbooks` `misc/building_evals.ipynb`) | | 2026-08-08 | ## Reference / schemas -| Page | Official doc page | Verified date | +| Page | Official doc page | As of | |---|---|---| | Docs index (discover any other page) | | 2026-08-06 | | CLI reference | | 2026-08-06 | @@ -169,8 +169,8 @@ page has produced inconsistent readings of the same entries. Second, a changelog what changed **in a version**, so always pin the version, and pair it with the topic page rather than replacing it, since the topic page stays authoritative for mechanism and semantics. -Machine-readable JSON Schemas (editor validation only; Claude Code ignores the `$schema` field at -load time, already cited in this repo's `CLAUDE.md`): `marketplace.json` → +Machine-readable JSON Schemas, which we use for editor validation only and never as a load-time +contract: `marketplace.json` → [`https://json.schemastore.org/claude-code-marketplace.json`](https://json.schemastore.org/claude-code-marketplace.json), `plugin.json` → [`https://json.schemastore.org/claude-code-plugin-manifest.json`](https://json.schemastore.org/claude-code-plugin-manifest.json) diff --git a/docs/plugin-philosophy.md b/docs/plugin-philosophy.md index 3c7f67a436..7b07c697fd 100644 --- a/docs/plugin-philosophy.md +++ b/docs/plugin-philosophy.md @@ -54,15 +54,15 @@ this rule entirely, being neither skill, agent, nor schema content. Identifying the manifest is for.) A git config **vendor section that is not a publisher name** is the git-native place for a -convention any `git config --get` consumer must be able to read. git-config(1) Variables states -that "Other git-related tools may and do use their own variables. When inventing new variables for -use in your own tool, make sure their names do not conflict with those that are used by Git itself -and other popular tools, and describe them in your documentation" (verified 2026-09-08 against -[git-config(1) Variables](https://git-scm.com/docs/git-config#_variables); recheck trigger: that -paragraph being rewritten). Git already owns `worktree.*` (`worktree.guessRemote`, -`worktree.useRelativePaths`; [git-worktree(1) Configuration](https://git-scm.com/docs/git-worktree#_configuration), -verified 2026-09-08; recheck trigger: git-config(1) adding a new `worktree.*` key), so a placement -key cannot live there. Popular tools typically name the section after the *tool* (`ghq.root`, +convention any `git config --get` consumer must be able to read. We name such a section so it +collides with neither Git's own variables nor other popular tools', and document it. Pointer: for +third-party tool variables, see +[git-config(1) Variables](https://git-scm.com/docs/git-config#_variables). As of: 2026-09-08. +Recheck trigger: that paragraph being rewritten. Git already owns `worktree.*` +(`worktree.guessRemote`, `worktree.useRelativePaths`; Pointer: +[git-worktree(1) Configuration](https://git-scm.com/docs/git-worktree#_configuration). As of: +2026-09-08. Recheck trigger: git-config(1) adding a new `worktree.*` key), so a placement key +cannot live there. Popular tools typically name the section after the *tool* (`ghq.root`, `git-town.*`, `lfs.*`, `wt.basedir`, `delta.*`, `hub.protocol`); a publisher-named key (`melodic.*`) is still an org-agnosticism defect, and a plugin-named key (`source-control.*`) still couples consumers to this marketplace's plugin identity. The worktree-placement key is therefore a @@ -109,13 +109,12 @@ Keep plugins horizontally decoupled: Claude Code installs automatically) or guarded behind an "if installed" check with the documented fallback. A bare unguarded cross-plugin reference is a defect. -This follows Claude Code's own distinction between standalone configuration and plugins. Standalone -configuration is for "personal workflows, project-specific customizations, quick experiments". -Plugins are for "sharing with teammates, distributing to community, versioned releases, reusable -across projects" -([create plugins](https://code.claude.com/docs/en/plugins#when-to-use-plugins-vs-standalone-configuration), -verified 2026-08-10). Namespaced skill invocations are part of that isolation, not an -implementation detail. +This follows Claude Code's own distinction between standalone configuration and plugins: this +marketplace ships plugins because it distributes versioned capability to other people and projects. +Pointer: for when to use a plugin rather than standalone configuration, see +. As of: +2026-08-10. Recheck trigger: that section moves or drops the distinction. Namespaced skill +invocations are part of that isolation, not an implementation detail. ### Hardcoded consumer specifics @@ -136,17 +135,18 @@ A skill name is an imperative verb phrase; the plugin namespace supplies the obj (`/machine-health:audit`, `/source-control:commit`). Names compose into instruction sentences, such as "/discovery:explore the module, then /planning:interview me", and one grammar keeps every name in the marketplace predictable. This is a deliberate, documented deviation from the official authoring -guidance's gerund preference, and the guidance sanctions it: gerunds are what it says to "consider -using", action-oriented names (`process-pdfs`, `analyze-spreadsheets`) are listed under "Acceptable -alternatives", and what it puts under Avoid is "inconsistent patterns within your skill collection", -which is exactly the consistency this section supplies -([skill authoring best practices, "Naming conventions"](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#naming-conventions), -verified 2026-09-10). Neither source treats one form as required: the agentskills.io -specification's own example names are noun phrases (`pdf-processing`, `data-analysis`, -`code-review`; [specification](https://agentskills.io/specification), verified 2026-09-10), and -Claude Code validates no naming form (`claude plugin validate` 2.1.263 passed a non-conforming -name, verified 2026-09-10). Recheck this paragraph when the page's Avoid list changes, when the -specification's example names change, or when `claude plugin validate` starts rejecting a form. +guidance's gerund preference. We read that guidance as allowing action-oriented names and as +asking above all for one consistent pattern within a skill collection, which is what this section +supplies. Neither the guidance nor the Agent Skills specification requires one form, and Claude +Code validates no naming form (our probe: `claude plugin validate` 2.1.263 passed a non-conforming +name on 2026-09-10). + +- **Pointer**: for skill naming, see + + and . +- **As of**: 2026-09-10 +- **Recheck trigger**: the guidance's list of names to avoid changes, the specification's example + names change, or `claude plugin validate` starts rejecting a form. Verb meanings are fixed: @@ -192,15 +192,17 @@ that no other skill here performs. Every exception is an entry on this list, decided per name. A name class is never blanket-sanctioned. -A plugin skill declares no frontmatter `name`. The field is optional and defaults to the directory -name ([frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference), fetched -2026-08-10), and the directory here is already the name the skill is documented and invoked by, so -declaring it restates the path in the character set the Agent Skills specification allows. The one -effect a declaration still buys is the bare alias: in a plugin skill a declared `name` also registers -the bare `/` alongside the namespaced command, unless another command already owns that token -([how a skill gets its command name](https://code.claude.com/docs/en/skills#how-a-skill-gets-its-command-name)). -Declare it only to take that alias deliberately, and only with the value the directory already -carries. A `name` that *differs* from its directory is out of bounds here even though the harness +A plugin skill declares no frontmatter `name`. We rely on the directory supplying the name, and +the directory here is already the name the skill is documented and invoked by, so declaring it +restates the path in the character set the Agent Skills specification allows. The one effect we +would declare it for is the bare `/` alias beside the namespaced command. Declare it only to +take that alias deliberately, and only with the value the directory already carries. + +- **Pointer**: for the `name` field, see + ; for the bare alias, see + . +- **As of**: 2026-08-10 +- **Recheck trigger**: `name` stops being optional, or a declared `name` stops adding a bare alias. A `name` that *differs* from its directory is out of bounds here even though the harness honors it: it would relocate the last command segment away from the directory that `scripts/check-skill-leaf-names.sh` derives every leaf from, desynchronizing the cross-plugin collision registry from the commands that actually resolve. That is why `skill-quality`'s check 1 @@ -209,17 +211,18 @@ Never degrade a name to dodge a built-in command: plugin skills are namespaced a with other levels. When a name matches a built-in, the bare token still belongs to the built-in; the namespaced form is the plugin skill's only command. -Resolution settled, **display** follows from it. Before v2.1.216 a declared `name` replaced the -whole command, so the menu showed the bare form and the namespaced one did not autocomplete; that -is gone. The skills page no longer pins that history; the basis is now the v2.1.216 -[changelog](https://code.claude.com/docs/en/changelog) entry, "Fixed plugin skills with a `name` -frontmatter field losing their plugin prefix in slash-command autocomplete" (verified 2026-08-31; -recheck trigger: a fetch of the changelog or the skills page no longer matching this record). What -the skills page pins instead is a successor quirk: a `name` that itself carries the plugin's own -prefix was doubled from v2.1.216 through v2.1.245 and is not re-prefixed on v2.1.246 or later -([how a skill gets its command name](https://code.claude.com/docs/en/skills#how-a-skill-gets-its-command-name), -fetched 2026-08-31), moot under this doctrine because the only sanctioned value is the bare -directory name. The rest is observed in the client rather than documented +Resolution settled, **display** follows from it. We rely on a plugin skill keeping its plugin +prefix in autocomplete whether or not it declares `name`. A version-dependent display quirk for a +`name` that carries the plugin's own prefix is moot under this doctrine, because the only +sanctioned value is the bare directory name. + +- **Pointer**: for the autocomplete fix, see the + [changelog entry 2.1.216](https://code.claude.com/docs/en/changelog#2-1-216); for the prefixed-`name` + quirk, see . +- **As of**: 2026-08-31 +- **Recheck trigger**: a fetch of the changelog or the skills page no longer matching this record. + +The rest is observed in the client rather than documented (2.1.225): the picker labels a row with the command it resolves, `/planning:plan` prefix and all, and appends a bare alias in parentheses only when what you typed prefix-matches that alias, so a skill declaring no `name` never renders the stuttering `/plugin:skill (skill)`. Re-observe before @@ -275,57 +278,59 @@ native one matures into fitness. ### Recorded gate runs Platform surfaces the gate has been run against, recorded in the -[upstream-drift](conventions/upstream-drift/README.md) four-part shape: claim, basis, as-of date, -trigger. Defer and decline are results, not omissions; the trigger, never the date, is what obliges -re-deriving a row. +[upstream-drift](conventions/upstream-drift/README.md#required-parts) record shape: each row holds +our verdict and reason in our words, the surface link (to the section the reason rests on) as the +pointer, the as-of date, and the recheck trigger. No row restates the page. Defer and decline are +results, not omissions; the trigger, never the date, is what obliges re-deriving a row. -| Surface | Verdict | Basis and reason | Recheck trigger | Verified | +| Surface (pointer) | Verdict | Decision and reason | Recheck trigger | As of | |---|---|---|---|---| -| [Run agents in parallel](https://code.claude.com/docs/en/agents) | Adopt, as a citation | The upstream comparison of every way Claude Code runs multiple agents: subagents, agent view, agent teams, dynamic workflows. Adopted as the [dispatch ladder](#dispatch-ladder)'s canonical index and cited there, never restated, so the menu an author chooses from cannot go stale inside this file. | The page adds or drops a parallelism surface. | 2026-08-10 | -| [Feature availability](https://code.claude.com/docs/en/feature-availability) | Adopt, as a citation | Per-feature availability by model provider and subscription plan, the canonical input to the [cross-platform contract](#cross-platform-contract), cited there. Its "platform" sense is the *provider* platform (Anthropic Console, Amazon Bedrock, Claude Platform on AWS, Google Cloud's Agent Platform, Microsoft Foundry), never the host surface a consumer runs in; that axis is [Platforms and integrations](https://code.claude.com/docs/en/platforms), a separate row below. Copying it is barred by [evidence and validation](#evidence-and-validation): a provider matrix is exactly the volatile table that rule names. | A plugin proposes narrowing its platform support. Re-fetch the matrix then, never trust a restatement. | 2026-08-10 | -| [Agent teams](https://code.claude.com/docs/en/agent-teams) | Defer | Fails gate 2 and stops there: "Agent teams are experimental and disabled by default", gated behind `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`, with limitations the page states outright: "No nested teams: teammates cannot spawn their own teammates", and `/resume` and `/rewind` do not restore in-process teammates. Defer rather than decline: the gap question stays open while the surface is opt-in and churning, and no plugin may depend on a team meanwhile. | The page drops the experimental warning or the `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS` requirement. | 2026-08-10 | -| [Cross-session messaging](https://code.claude.com/docs/en/cross-session-messaging) | Decline | Fails gate 1: the channel is for "independent sessions that you start and steer yourself", not for a skill dispatching a worker. It could not be a portable rung either: "not available on Amazon Bedrock, Claude Platform on AWS, Google Cloud's Agent Platform, or Microsoft Foundry". Re-derived 2026-08-24 after the prior trigger's Windows leg fired: native Windows is now supported ("v2.1.234 or later on native Windows"), which removes one portability leg but moves neither surviving premise, so the verdict stands. | Either surviving premise moves: the page stops scoping the channel to sessions you steer yourself, or its Availability section stops excluding any of those four providers. | 2026-08-24 | -| [Sessions](https://code.claude.com/docs/en/sessions) | Decline | Fails gate 1. Resume restores "the full history, including tool calls and results", which is the authoring story the [inline-template conventions](#inline-template-conventions) exist to withhold, so it is the opposite of a fresh-eyes rung rather than a missing one. A human session-management surface with no plugin-authoring interface. | `sessions` grows a plugin-facing interface: a manifest field, a tool, or skill frontmatter. | 2026-08-10 | -| [Platforms and integrations](https://code.claude.com/docs/en/platforms) | Adopt, as a citation | The upstream index of every host Claude Code runs in, covering CLI, Desktop, VS Code, JetBrains, web, and mobile, plus the integrations beside them. Adopted as the [cross-platform contract](#cross-platform-contract)'s canonical input for the host axis, which the already-adopted [feature availability](https://code.claude.com/docs/en/feature-availability) does not carry: that page's axes are provider and plan, and it scopes itself to what runs locally: "The Claude Code CLI and everything that runs locally work on every provider." The host axis matters because a host can withhold the plugin system outright rather than one capability: a Desktop session in WSL 2 lists "connectors and plugins" among features that "aren't available in WSL sessions yet"; on mobile, "commands that only run in the terminal interface, such as `/plugin` and `/resume`, don't work from the app"; Desktop's Cowork tab sources its plugins from claude.ai configuration, "not from the CLI's `~/.claude` directory"; and the VS Code extension carries only a "Subset" of the CLI's "Commands and skills", so a skill this fleet ships may simply not be reachable there. None of those four is restated in the contract. Only the rule they establish is. | `platforms` adds or drops a host, or `feature-availability` grows a host-surface axis, which would make this citation redundant. | 2026-08-10 | -| [GitHub Enterprise Server](https://code.claude.com/docs/en/github-enterprise-server) | Decline | Does not fail gate 1 by subject: it is a real plugin-distribution surface, "Plugin marketplaces \| ✅ Supported", and the only page in this run that names one. It fails on need. Nothing in this repo documents a GHES-hosted mirror or fork of this marketplace, and no README anywhere ships a full-git-URL install path, the form GHES requires. Census of the 65 plugin READMEs: 54 carry the literal `/plugin marketplace add melodic-software/claude-code-plugins`; 9 carry no install block; `dometrain` points at another github.com marketplace; and `github`, being marketplace-agnostic, uses the placeholder `/`. All of those are the same `owner/repo` shorthand, which the page says "always resolves to github.com", correct for this marketplace, and the one place the finding could bite: a consumer redistributing the `github` plugin from a GHES-hosted marketplace would follow that README and silently resolve to github.com instead of their own host. Otherwise the GHES-specific obligations land on a consumer running their own instance, not on this marketplace: full git URL, `extraKnownMarketplaces` pre-registration, `hostPattern` allowlisting. | This repo documents a GHES-hosted mirror or fork, or any README gains an install path that is not `owner/repo` shorthand, a full git URL being the form that means a non-github.com host is in play. Also fires if `plugins/github/README.md` starts naming a concrete GHES-hosted marketplace. | 2026-08-10 | -| [Ultrareview](https://code.claude.com/docs/en/ultrareview) | Decline | Fails gate 1: no interface a plugin can reach. Each run is human-gated and metered: "Claude Code shows a confirmation dialog with the review scope, your remaining free runs, and the estimated cost", then "typically \$5 to \$25 in usage credits". So it can never be a rung in an automated dispatch ladder. Nor is `review:fanout` a custom rebuild of it that Native-first would retire: fanout normalizes many in-session finding producers into one ranked report, where this is one confirmed cloud run. | The page documents a non-interactive or programmatic entry point. | 2026-08-10 | -| [Chrome](https://code.claude.com/docs/en/chrome) | Decline | Fails gate 1: a consumer-installed browser integration delivered as a built-in skill, since Claude Code "asks for permission to use the `claude-in-chrome` skill", so a plugin has nothing to declare here and must not rebuild automation the platform already ships. Recorded rather than dismissed because the page only *looked* cited: the repo's sole reference is a `docs/en/browser` URL that now returns 404, inside `plugins/playbooks/skills/boris/vendor/SKILL.md`, a verbatim upstream baseline kept for drift detection, which is why it is deliberately not hand-edited here. | A plugin proposes shipping browser automation, or `/playbooks:update` refreshes the boris baseline and the stale slug persists. | 2026-08-10 | -| [Mods (hooks modules)](https://github.com/anthropics/claude-code/tree/main/mods) | Defer | Fails gate 2 and stops there. A mod is a plugin whose behavior lives in one `register(on, options)` hooks module running in-process. Anthropic's own `mods/README.md` states the interface "may change between releases without notice"; a mod you write is off by default behind the rollout gate `tengu_plugin_hooks_modules`, whose default is `false` and which `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS` only overrides per process; and the feature has zero mentions in the official docs (all 197 pages via `llms-full.txt`) or in `CHANGELOG.md`, checked at Claude Code 2.1.278. Defer rather than decline: the surface is real and shipping, so the gap question stays open, and no plugin may depend on it meanwhile. Recorded in [ADR 0035](adr/0035-defer-claude-code-mods-with-five-go-criteria.md). | All five go criteria hold: a test mod loads with the enable flag unset, the official docs mention the feature, [#92533](https://github.com/anthropics/claude-code/issues/92533) is closed, the official docs state the throw and timeout semantics and the engine default on an uncaught throw is settled upstream (the generated `.d.ts` JSDoc already states the mechanism, so a JSDoc hit does not meet this), and the early-access warning is gone from `mods/README.md`. Commands and expected outputs: [go-no-go.md](upstream/claude-code-mods/go-no-go.md). Any one failing is no-go. | 2026-09-19 | -| [Checkpointing](https://code.claude.com/docs/en/checkpointing) | Decline | Nothing to adopt, and the reason is the outcome: `/rewind` cannot be a mutating skill's undo story, because "Checkpointing does not track files modified by bash commands" and, for any subagent other than a foreground forked skill, "rewinding doesn't restore the edits. Use git to revert them." The restored carve-out is narrow, covering only a `context: fork` skill running in the foreground, so a skill that mutates through a shell script or a background worker states a git-based rollback and never leans on `/rewind`. | The limitations section drops either the bash-command or the subagent exclusion. | 2026-08-10 | +| [Run agents in parallel](https://code.claude.com/docs/en/agents#choose-an-approach) | Adopt, as a citation | The upstream comparison of every way Claude Code runs multiple agents: subagents, agent view, agent teams, dynamic workflows. Adopted as the [dispatch ladder](#dispatch-ladder)'s canonical index and cited there, never restated, so the menu an author chooses from cannot go stale inside this file. | The page adds or drops a parallelism surface. | 2026-08-10 | +| [Feature availability](https://code.claude.com/docs/en/feature-availability) | Adopt, as a citation | Per-feature availability by model provider and subscription plan, the canonical input to the [cross-platform contract](#cross-platform-contract), cited there. We read its "platform" sense as the *provider* platform, never the host surface a consumer runs in; that axis is [Platforms and integrations](https://code.claude.com/docs/en/platforms), a separate row below. Copying it is barred by [evidence and validation](#evidence-and-validation): a provider matrix is exactly the volatile table that rule names. | A plugin proposes narrowing its platform support. Re-fetch the matrix then, never trust a restatement. | 2026-08-10 | +| [Agent teams](https://code.claude.com/docs/en/agent-teams#limitations) | Defer | Fails gate 2 and stops there: the feature is marked experimental and is off unless `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` is set, and its stated limitations cover nesting and resume. Defer rather than decline: the gap question stays open while the surface is opt-in and churning, and no plugin may depend on a team meanwhile. | The page drops the experimental warning or the `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS` requirement. | 2026-08-10 | +| [Cross-session messaging](https://code.claude.com/docs/en/cross-session-messaging#when-to-use-cross-session-messaging), [availability](https://code.claude.com/docs/en/cross-session-messaging#availability) | Decline | Fails gate 1: we read the channel as one between sessions a person starts and steers, not a worker a skill dispatches. Re-derived 2026-10-01 after the provider leg of the prior trigger fired: same-machine messaging now reaches every provider, so portability no longer bars a same-machine rung, but gate 1 holds the verdict on its own. | The page stops scoping the channel to sessions you steer yourself, or a plugin surface (a manifest field, skill frontmatter, or a tool a skill may call) gains a way to start and address a session. | 2026-10-01 | +| [Sessions](https://code.claude.com/docs/en/sessions#what-a-resumed-session-restores) | Decline | Fails gate 1. Resume restores the prior conversation in full, which is the authoring story the [inline-template conventions](#inline-template-conventions) exist to withhold, so it is the opposite of a fresh-eyes rung rather than a missing one. A human session-management surface with no plugin-authoring interface. | `sessions` grows a plugin-facing interface: a manifest field, a tool, or skill frontmatter. | 2026-08-10 | +| [Platforms and integrations](https://code.claude.com/docs/en/platforms) | Adopt, as a citation | The upstream index of every host Claude Code runs in, covering CLI, Desktop, VS Code, JetBrains, web, and mobile, plus the integrations beside them. Adopted as the [cross-platform contract](#cross-platform-contract)'s canonical input for the host axis, which the already-adopted [feature availability](https://code.claude.com/docs/en/feature-availability) does not carry: that page's axes are provider and plan, scoped to what runs locally. The host axis matters because a host can withhold the plugin system outright rather than one capability: Desktop sessions in WSL 2, the mobile app, Desktop's Cowork tab, and the VS Code extension each limit plugins, terminal-only commands, or skills relative to the CLI, so a skill this fleet ships may simply not be reachable there (read each host's page from the index for the specifics). None of those four is restated in the contract. Only the rule they establish is. | `platforms` adds or drops a host, or `feature-availability` grows a host-surface axis, which would make this citation redundant. | 2026-08-10 | +| [GitHub Enterprise Server](https://code.claude.com/docs/en/github-enterprise-server#plugin-marketplaces-on-ghes) | Decline | Does not fail gate 1 by subject: it is a real plugin-distribution surface, and the only page in this run that names one. It fails on need. Nothing in this repo documents a GHES-hosted mirror or fork of this marketplace, and no README anywhere ships a full-git-URL install path, the form GHES requires. Census of the 65 plugin READMEs: 54 carry the literal `/plugin marketplace add melodic-software/claude-code-plugins`; 9 carry no install block; `dometrain` points at another github.com marketplace; and `github`, being marketplace-agnostic, uses the placeholder `/`. All of those are the same `owner/repo` shorthand, which we treat as resolving to github.com, correct for this marketplace, and the one place the finding could bite: a consumer redistributing the `github` plugin from a GHES-hosted marketplace would follow that README and silently resolve to github.com instead of their own host. Otherwise the GHES-specific obligations land on a consumer running their own instance, not on this marketplace: full git URL, `extraKnownMarketplaces` pre-registration, `hostPattern` allowlisting. | This repo documents a GHES-hosted mirror or fork, or any README gains an install path that is not `owner/repo` shorthand, a full git URL being the form that means a non-github.com host is in play. Also fires if `plugins/github/README.md` starts naming a concrete GHES-hosted marketplace. | 2026-08-10 | +| [Ultrareview](https://code.claude.com/docs/en/ultrareview#run-ultrareview-non-interactively), [pricing](https://code.claude.com/docs/en/ultrareview#pricing-and-free-runs) | Decline | Fails gate 1 for automated dispatch. Re-derived 2026-10-01 after the prior trigger fired (a non-interactive entry point now exists): each run is metered, and we read the page as treating the person who starts a run as the one consenting to its billing, so no skill may launch one on its own, and a person stays free to run it. Nor is `review:fanout` a custom rebuild of it that Native-first would retire: fanout normalizes many in-session finding producers into one ranked report, where this is one consented cloud run. | The page lets a run that Claude starts count as consent to its billing, or runs stop being metered. | 2026-10-01 | +| [Chrome](https://code.claude.com/docs/en/chrome) | Decline | Fails gate 1: a consumer-installed browser integration the platform ships itself, so a plugin has nothing to declare here and must not rebuild automation the platform already ships. Recorded rather than dismissed because the page only *looked* cited: the repo's sole reference is a `docs/en/browser` URL that now returns 404, inside `plugins/playbooks/skills/boris/vendor/SKILL.md`, a verbatim upstream baseline kept for drift detection, which is why it is deliberately not hand-edited here. | A plugin proposes shipping browser automation, or `/playbooks:update` refreshes the boris baseline and the stale slug persists. | 2026-08-10 | +| [Mods (hooks modules)](https://github.com/anthropics/claude-code/tree/main/mods) | Defer | Fails gate 2 and stops there. A mod is a plugin whose behavior lives in one `register(on, options)` hooks module running in-process. Anthropic's own `mods/README.md` marks the interface as unstable between releases; a mod you write is off by default behind the rollout gate `tengu_plugin_hooks_modules`, whose default is `false` and which `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS` only overrides per process; and the feature has zero mentions in the official docs (all 197 pages via `llms-full.txt`) or in `CHANGELOG.md`, checked at Claude Code 2.1.278. Defer rather than decline: the surface is real and shipping, so the gap question stays open, and no plugin may depend on it meanwhile. Recorded in [ADR 0035](adr/0035-defer-claude-code-mods-with-five-go-criteria.md). | All five go criteria hold: a test mod loads with the enable flag unset, the official docs mention the feature, [#92533](https://github.com/anthropics/claude-code/issues/92533) is closed, the official docs state the throw and timeout semantics and the engine default on an uncaught throw is settled upstream (the generated `.d.ts` JSDoc already states the mechanism, so a JSDoc hit does not meet this), and the early-access warning is gone from `mods/README.md`. Commands and expected outputs: [go-no-go.md](upstream/claude-code-mods/go-no-go.md). Any one failing is no-go. | 2026-09-19 | +| [Checkpointing: bash changes](https://code.claude.com/docs/en/checkpointing#bash-command-changes-not-tracked), [subagent edits](https://code.claude.com/docs/en/checkpointing#subagent-edits-not-restored) | Decline | Nothing to adopt, and the reason is the outcome: `/rewind` cannot be a mutating skill's undo story, because checkpoints do not cover bash-command changes or the edits of most subagents. The restored carve-out is narrow, covering only a `context: fork` skill running in the foreground, so a skill that mutates through a shell script or a background worker states a git-based rollback and never leans on `/rewind`. | The limitations section drops either the bash-command or the subagent exclusion. | 2026-08-10 | ## Component stances -> **Staleness disclaimer.** The platform changes constantly. Every row carries the date its facts -> were verified against the linked official page. Always re-fetch the current page before acting on -> a row; never trust this table alone. A fetch that diverges from a row is that row's recheck -> trigger: update the row, refreshing its verified date with the outcome. The -> [upstream-drift convention](conventions/upstream-drift/README.md) owns this stamp discipline. +> **Staleness disclaimer.** The platform changes constantly. Each row states our stance in our +> words; its component link is the pointer, and its as-of date is when the stance was last derived +> from that page. Always re-fetch the current page before acting on a row; never trust this table +> alone. A fetch that diverges from a row is that row's recheck trigger: re-derive the row and +> refresh its as-of date with the outcome. The +> [upstream-drift convention](conventions/upstream-drift/README.md) owns this record discipline. -| Component | Stance | Rationale and constraints | Verified | +| Component (pointer) | Stance | Rationale and constraints | As of | |---|---|---|---| -| [Skills](https://code.claude.com/docs/en/skills) | Primary surface | The default unit of capability. Newer frontmatter is adopted case-by-case through the adoption gate: `paths`, `context: fork` (+ `agent`), `arguments`, skill-scoped `hooks` with `once`, and `model` (the override lasts for the current turn and is not saved; in auto mode a model auto mode does not support is not used and the session keeps its model; with `context: fork` the value sets the forked subagent's model). `model` verified 2026-09-29 against the [frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference). Recheck when that row changes what `model` accepts or when auto mode stops keeping the session model. | 2026-09-29 | -| [`commands/`](https://code.claude.com/docs/en/plugins-reference) | Prohibited | Officially merged into skills; docs direct "use `skills/` for new plugins". Existing flat commands migrate to skill directories. | 2026-07-17 | -| [Agents](https://code.claude.com/docs/en/sub-agents) | Adopt on need | Plugin agents do not support `hooks`, `mcpServers`, or `permissionMode` (security restriction). Design within that limit rather than working around it. | 2026-07-17 | -| [Workflows](https://code.claude.com/docs/en/workflows) | Adopt on need | Native and not experimental: a script in `workflows/`, or wherever the `workflows` manifest field points (that field replaces the default scan), runs as a plugin-namespaced `/plugin:name` command. Availability, not maturity, is the constraint: workflows are paid-plan-gated, a consumer can switch them off (`disableWorkflows`, `CLAUDE_CODE_DISABLE_WORKFLOWS`), and an org can disable them fleet-wide in managed settings; so, as with `bin/`, never make a workflow the only path to a capability. Not "Wait": the [deferred workflow engines](adr/0020-defer-three-medley-surfaces-with-explicit-recheck-triggers.md) are a named candidate carrying a live trigger, so the gap is identified rather than hypothetical. None ship in this fleet today. | 2026-07-27 | -| [Hooks](https://code.claude.com/docs/en/hooks) | Adopt on need | Exec form (`args`) is mandatory wherever `${user_config.*}` appears, because shell form errors since v2.1.207; otherwise read the `CLAUDE_PLUGIN_OPTION_` mirror. Windows exec form spawns a real executable such as a `.exe` with the `args` array and no shell, so a shebang script or a `.cmd`/`.bat` shim is not a `command`, and neither is a bare `bash`, `sh`, `python`, or `python3` (a failed launch is non-blocking, so a guard then enforces nothing). Shell form with `"shell": "bash"` stays legal where no `${user_config.*}` appears; every plugin hook row uses exec form, `"command": "node"` with the script path in `args`, except the guardrails and disk-hygiene SessionStart node notice rows and the harness-ops hook-failure-audit Stop row, which run in shell form with `"shell": "bash"` because they must work when `node` is missing. `node` must be on `PATH`, and Claude Code does not guarantee it: exec form resolves `command` on `PATH` ([Exec form and shell form](https://code.claude.com/docs/en/hooks#exec-form-and-shell-form)), and the installed `claude` binary does not itself invoke Node ([Install with npm](https://code.claude.com/docs/en/setup#install-with-npm)), both fetched 2026-09-29. A hook that cannot start is a non-blocking error, so a guard whose `node` is missing enforces nothing and the transcript notice is the only signal ([Other exit codes](https://code.claude.com/docs/en/hooks#other-exit-codes)). `scripts/check-hook-exec-form.sh` rejects a bare name other than `node`. `scripts/check-exec-form-windows-probe.sh` rejects a script path used as `command`; its non-Windows skip does not authorize converting `.sh` rows. The four-part record is [Windows exec-form probe](#windows-exec-form-probe). Hooks modules ("mods"), the in-process TypeScript hook form, are deferred: see the mods row under [Recorded gate runs](#recorded-gate-runs) and [ADR 0035](adr/0035-defer-claude-code-mods-with-five-go-criteria.md). | 2026-09-29 | -| [MCP servers](https://code.claude.com/docs/en/mcp) | Adopt on need | Clears the plugin-acceptance security review for egress and trust delegation. Also the only component type that can cost a consumer their prompt cache: every other kind only appends to the request, while enabling or disabling a plugin that provides an MCP server forces a full re-read whenever the server's tools load into the prefix instead of being deferred by tool search ([actions that invalidate the cache](https://code.claude.com/docs/en/prompt-caching#actions-that-invalidate-the-cache), verified 2026-08-10). | 2026-08-10 | -| [LSP servers](https://code.claude.com/docs/en/plugins-reference) | Adopt on need | Consumer must have the language-server binary; declare the prerequisite per the failure-behavior rules. | 2026-07-17 | -| [Output styles](https://code.claude.com/docs/en/plugins-reference) | Adopt on need | No additional constraints. | 2026-07-17 | -| [`bin/`](https://code.claude.com/docs/en/plugins) | Adopt on need | Executables join the Bash tool's `PATH` while the plugin is enabled; names must be collision-safe (plugin-prefixed), because the platform does not namespace them. That `PATH` delivery is per-session and can silently fail ([anthropics/claude-code#68066](https://github.com/anthropics/claude-code/issues/68066)), so never depend on bare-name invocation: invoke via `${CLAUDE_PLUGIN_ROOT}/bin/`, and note that a `bash "…/bin/x"` invocation does not match a `Bash(x:*)` allow rule. | 2026-07-17 | -| [Plugin `settings.json`](https://code.claude.com/docs/en/plugins) | `agent` prohibited by default | Supports only `agent` and `subagentStatusLine`. `agent` takes over the main thread, a consumer-hostile default for a marketplace plugin; any exception requires documented justification in the plugin README. | 2026-07-17 | -| [Monitors](https://code.claude.com/docs/en/plugins-reference) | Wait | Experimental (`experimental.monitors`); interactive-CLI-only, unsandboxed at hook trust level, no `${user_config.*}` and no `CLAUDE_PLUGIN_OPTION_*` in monitor processes; keep running after mid-session disable. Re-verify before each audit. | 2026-07-17 | -| [Themes](https://code.claude.com/docs/en/plugins-reference) | Wait | Experimental (`experimental.themes`); schema may change between releases. Re-verify before each audit. | 2026-07-17 | -| [Channels](https://code.claude.com/docs/en/plugins-reference) | Wait | No longer carries an official experimental label, but fails the adoption gate today: no fleet gap it fills. Re-verify before each audit. | 2026-07-17 | -| [Dependencies](https://code.claude.com/docs/en/plugin-dependencies) | Adopt on need (hard requires only) | See the design boundary: hard requires only, semver-constrained, released via `{name}--v{version}` tags. None exist in this fleet today. | 2026-07-17 | +| [Skills](https://code.claude.com/docs/en/skills) | Primary surface | The default unit of capability. Newer frontmatter is adopted case-by-case through the adoption gate: `paths`, `context: fork` (+ `agent`), `arguments`, skill-scoped `hooks` with `once`, and `model`, which we use only as a per-turn override, including for a forked subagent. Pointer for `model`: the [frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference). Recheck trigger: that row changes what `model` accepts, or auto mode stops keeping the session model. | 2026-09-29 | +| [`commands/`](https://code.claude.com/docs/en/plugins/components#commands) | Prohibited | Superseded by skills upstream; every new capability goes in `skills/`. Existing flat commands migrate to skill directories. | 2026-07-17 | +| [Agents](https://code.claude.com/docs/en/plugins/components#frontmatter-fields-in-plugin-agents) | Adopt on need | Plugin agents do not support `hooks`, `mcpServers`, or `permissionMode` (security restriction). Design within that limit rather than working around it. | 2026-07-17 | +| [Workflows](https://code.claude.com/docs/en/workflows#distribute-a-workflow-in-a-plugin) | Adopt on need | Native and not experimental: a script in `workflows/`, or wherever the `workflows` manifest field points (that field replaces the default scan), runs as a plugin-namespaced `/plugin:name` command. Availability, not maturity, is the constraint: workflows are paid-plan-gated, a consumer can switch them off (`disableWorkflows`, `CLAUDE_CODE_DISABLE_WORKFLOWS`), and an org can disable them fleet-wide in managed settings; so, as with `bin/`, never make a workflow the only path to a capability. Not "Wait": the [deferred workflow engines](adr/0020-defer-three-medley-surfaces-with-explicit-recheck-triggers.md) are a named candidate carrying a live trigger, so the gap is identified rather than hypothetical. None ship in this fleet today. | 2026-07-27 | +| [Hooks](https://code.claude.com/docs/en/hooks) | Adopt on need | Exec form (`args`) is mandatory wherever `${user_config.*}` appears, because shell form errors since v2.1.207; otherwise read the `CLAUDE_PLUGIN_OPTION_` mirror. On Windows, exec form launches an executable file (a `.exe`, for example) directly with the `args` array and no shell, so a shebang script or a `.cmd`/`.bat` shim is not a `command`, and neither is a bare `bash`, `sh`, `python`, or `python3` (a failed launch is non-blocking, so a guard then enforces nothing). Shell form with `"shell": "bash"` stays legal where no `${user_config.*}` appears; every plugin hook row uses exec form, `"command": "node"` with the script path in `args`, except the guardrails and disk-hygiene SessionStart node notice rows and the harness-ops hook-failure-audit Stop row, which run in shell form with `"shell": "bash"` because they must work when `node` is missing. `node` must be on `PATH`, and we do not assume a Claude Code install brings it (pointers: [Exec form and shell form](https://code.claude.com/docs/en/hooks#exec-form-and-shell-form), [Install with npm](https://code.claude.com/docs/en/setup#install-with-npm)). We treat a hook that cannot start as a guard that enforced nothing, with the transcript notice as the only signal (pointer: [Other exit codes](https://code.claude.com/docs/en/hooks#other-exit-codes)). `scripts/check-hook-exec-form.sh` rejects a bare name other than `node`. `scripts/check-exec-form-windows-probe.sh` rejects a script path used as `command`; its non-Windows skip does not authorize converting `.sh` rows. The record is [Windows exec-form probe](#windows-exec-form-probe). Hooks modules ("mods"), the in-process TypeScript hook form, are deferred: see the mods row under [Recorded gate runs](#recorded-gate-runs) and [ADR 0035](adr/0035-defer-claude-code-mods-with-five-go-criteria.md). | 2026-09-29 | +| [MCP servers](https://code.claude.com/docs/en/mcp) | Adopt on need | Clears the plugin-acceptance security review for egress and trust delegation. Also the only component type that can cost a consumer their prompt cache: every other kind only appends to the request, while enabling or disabling a plugin that provides an MCP server forces a full re-read whenever the server's tools load into the prefix instead of being deferred by tool search (pointer: [actions that invalidate the cache](https://code.claude.com/docs/en/prompt-caching#actions-that-invalidate-the-cache)). | 2026-08-10 | +| [LSP servers](https://code.claude.com/docs/en/plugins/components#lsp-servers) | Adopt on need | Consumer must have the language-server binary; declare the prerequisite per the failure-behavior rules. | 2026-07-17 | +| [Output styles](https://code.claude.com/docs/en/plugins/components#themes-and-output-styles) | Adopt on need | No additional constraints. | 2026-07-17 | +| [`bin/`](https://code.claude.com/docs/en/plugins/components#executables) | Adopt on need | A plugin's executables reach the Bash tool's `PATH` for as long as the plugin stays on; names must be collision-safe (plugin-prefixed), because the platform does not namespace them. That `PATH` delivery is per-session and can silently fail ([anthropics/claude-code#68066](https://github.com/anthropics/claude-code/issues/68066)), so never depend on bare-name invocation: invoke via `${CLAUDE_PLUGIN_ROOT}/bin/`, and note that a `bash "…/bin/x"` invocation does not match a `Bash(x:*)` allow rule. | 2026-07-17 | +| [Plugin `settings.json`](https://code.claude.com/docs/en/plugins/components#default-settings) | `agent` prohibited by default | Supports only `agent` and `subagentStatusLine`. `agent` takes over the main thread, a consumer-hostile default for a marketplace plugin; any exception requires documented justification in the plugin README. | 2026-07-17 | +| [Monitors](https://code.claude.com/docs/en/plugins/components#monitors) | Wait | Experimental (`experimental.monitors`); interactive-CLI-only, unsandboxed at hook trust level, no `${user_config.*}` and no `CLAUDE_PLUGIN_OPTION_*` in monitor processes; keep running after mid-session disable. Re-verify before each audit. | 2026-07-17 | +| [Themes](https://code.claude.com/docs/en/plugins/components#themes-and-output-styles) | Wait | Experimental (`experimental.themes`); schema may change between releases. Re-verify before each audit. | 2026-07-17 | +| [Channels](https://code.claude.com/docs/en/plugins/components#channels) | Wait | No longer carries an official experimental label, but fails the adoption gate today: no fleet gap it fills. Re-verify before each audit. | 2026-07-17 | +| [Dependencies](https://code.claude.com/docs/en/plugins/dependencies) | Adopt on need (hard requires only) | See the design boundary: hard requires only, semver-constrained, released via `{name}--v{version}` tags. None exist in this fleet today. | 2026-07-17 | ### Windows exec-form probe `scripts/check-exec-form-windows-probe.sh` rejects an exec-form `command` that is not a real Windows executable ([#3686](https://github.com/melodic-software/claude-code-plugins/issues/3686)). It does not rewrite rows. A `.sh` path, a `.cmd`/`.bat` shim, or bare `bash` as `command` stays illegal. `scripts/check-hook-exec-form.sh` keeps rejecting bare `bash` with the script in `args`. Every shipped hook row is exec form, except the three shell-form rows named in the Hooks row above: `"command": "node"` with `hooks/exec-bash.mjs` (canonical `lib/exec-bash.mjs`, copied by `scripts/sync-exec-bash.sh`) and then the script. The launcher finds Git Bash and never `System32\bash.exe`. A default-off option is `--require-true NAME` (exit 0 unless `CLAUDE_PLUGIN_OPTION_NAME` is `true`). A default-on option is `--run-if-unset-or-true NAME` (exit 0 only when that variable is set to something other than `true`). Skill-frontmatter `args` is a YAML sequence, one element per argument. -- **Claim:** On Windows, exec form (`args` present) resolves `command` as an executable and spawns it directly with `args` as the argument vector. There is no shell, so a shebang is not honored, and `command` must be a real executable such as a `.exe`. `.cmd` and `.bat` shims cannot be spawned. If a Windows spawn of that shape drops `args` or the process image is `bash.exe`, the fleet sweep stops. -- **Basis:** [Hooks reference](https://code.claude.com/docs/en/hooks), section "Exec form and shell form". Verbatim, from a full raw-markdown read of `https://code.claude.com/docs/en/hooks.md` (330,813 bytes, SHA-256 `57e3b47d55acfbae3dcdc112866c8c0f75528d8b5c4fca9bfcdaa904d4728218`; the slug is listed in `https://code.claude.com/docs/llms.txt`): "On Windows, exec form requires `command` to resolve to a real executable such as a `.exe`." The same section states that exec form has no shell and that `shell` is "Ignored when `args` is set". Args-drop is [anthropics/claude-code#90495](https://github.com/anthropics/claude-code/issues/90495), open as of this date. +- **Decision:** every exec-form `command` is a real Windows executable (`node`), never a shebang script, a `.cmd`/`.bat` shim, or bare `bash`, and no row relies on a shell or on the `shell` field while `args` is set. If a Windows spawn of that shape drops `args` or the process image is `bash.exe`, the fleet sweep stops. +- **Pointer:** for exec form on Windows, see (read from the raw `.md`, 330,813 bytes, SHA-256 `57e3b47d55acfbae3dcdc112866c8c0f75528d8b5c4fca9bfcdaa904d4728218`); for the args drop, see [anthropics/claude-code#90495](https://github.com/anthropics/claude-code/issues/90495), open as of this date. - **As of:** 2026-09-28. -- **Recheck:** the hooks page changes that Windows sentence, stops ignoring `shell` when `args` is set, or #90495 closes. +- **Recheck trigger:** the hooks page changes what exec form requires of `command` on Windows, stops ignoring `shell` when `args` is set, or #90495 closes. On a non-Windows host the spawn half prints a `SKIP` line and exits 0 when the row spellings are clean. That skip is fail-soft: it does not show that #90495 is absent, and it does not authorize converting `.sh` rows. On Windows the script spawns `node.exe` (a PE image, not a `.cmd`) with an args array and a stdin payload. A missing sentinel or a `bash.exe` image exits 1 with `ARGS-DROP` and the sweep stays stopped. `EXEC_FORM_WINDOWS_PROBE_LIVE=1` adds an opt-in `claude` hook run; without `claude` on `PATH` that half also skips fail-soft. @@ -471,8 +476,9 @@ a consumer sets a value through; this rule settles where the plugin's own values - Docs cite the owner and do not restate the number. A copy generated from the owner is not a second owner. - Measurement records keep their numbers: a recorded result is data about one run, not a restatement. -- A volatile external specific restated in a skill body carries the four-part record of the - [upstream-drift convention](conventions/upstream-drift/README.md). +- A skill body never restates a volatile external specific. It states our decision and carries the + pointer, as-of date, and recheck trigger of the + [upstream-drift convention](conventions/upstream-drift/README.md#required-parts). - A value read from two languages lives in a JSON file both read. Worked example: `animation`'s brush defaults live in `skills/rotoscope/scripts/brush.json`, read by @@ -523,10 +529,10 @@ hook and a `git`-dependent tier set whose absence the native prompt cannot see, back to its default and which has no external prerequisite, correctly ships none. A setup skill was written for it and deliberately dropped rather than kept for symmetry. -The uniform contract: the skill is named `setup`, sets `disable-model-invocation: true`, matching -upstream's own rule for the flag, "for workflows with side effects that you want to trigger -manually" ([best practices](https://code.claude.com/docs/en/best-practices), verified 2026-08-10), -and offers `check` (read-only inspect and verify) and `apply` (idempotent configure) actions. This +The uniform contract: the skill is named `setup`, sets `disable-model-invocation: true`, because +setup has side effects a person should trigger by hand (Pointer: +. As of: 2026-08-10. Recheck trigger: +that section stops tying the flag to manually triggered side effects), and offers `check` (read-only inspect and verify) and `apply` (idempotent configure) actions. This contract is exception class (ii) of the fleet's invocation-mode rubric ([`docs/conventions/invocation-mode/`](conventions/invocation-mode/README.md)), which owns the default and the other reasons a skill may set the flag. The @@ -592,15 +598,13 @@ a separate `:check` skill with `disable-model-invocation: false`. It rea follows only its `check` section, so `setup` stays the one account of what is checked, and it installs and writes nothing. `setup` keeps `true`, and `apply` stays manual. A hook or probe names `/:check`, never `/:setup check`, which the flag hides from Claude. -`plugins/context7/skills/check/SKILL.md` is the shape to copy. Verification record. Claim: the flag -is set per skill in frontmatter, and the skills page documents no per-action invocation flag. -Basis: , field -`disable-model-invocation` ("Set to `true` to prevent Claude from automatically loading this -skill. Use for workflows you want to trigger manually with `/name`."), and - ("Two frontmatter fields let -you restrict this"), both read from the raw `.md` of that page on 2026-09-29. As of: 2026-09-29. -Recheck: a Claude Code release adds a per-action invocation flag, or that field's description -stops applying to the whole skill. +`plugins/context7/skills/check/SKILL.md` is the shape to copy. Record. Decision: we set the flag +per skill and split `check` into its own skill, because we found no per-action invocation flag. +Pointer: for `disable-model-invocation`, see + and +. As of: 2026-09-29. Recheck +trigger: a Claude Code release adds a per-action invocation flag, or that field stops applying to +the whole skill. The verb set is deliberately closed at `check` and `apply`: no standalone `remove`, `reset`, or `migrate` verb joins the mandatory contract (teardown, where genuinely needed, rides as a `remove` @@ -681,12 +685,14 @@ data: removing the plugin's own tracked setup config is teardown, whereas an app that mutates a managed inventory the plugin maintains (a status change over existing entries, say) is ordinary `apply` surface, not teardown, and does not trip that trigger. -Two native idioms are the sanctioned initialization surfaces (verified 2026-08-10 against the -[hooks reference](https://code.claude.com/docs/en/hooks) and -[plugins reference](https://code.claude.com/docs/en/plugins-reference)): the `Setup` hook event -(`--init-only`, or `--init`/`--maintenance` in `-p` mode) for headless and CI preparation, and a -`SessionStart` hook comparing a bundled manifest against its `${CLAUDE_PLUGIN_DATA}` copy for -runtime-dependency installation. +Two native idioms are the sanctioned initialization surfaces: the `Setup` hook event for headless +and CI preparation, and a `SessionStart` hook comparing a bundled manifest against its +`${CLAUDE_PLUGIN_DATA}` copy for runtime-dependency installation. Pointer: for the `Setup` event and +the flags that fire it, see ; for the data-directory +install idiom, see +. +As of: 2026-08-10. Recheck trigger: either section moves, or the `Setup` event or the data-directory +idiom is removed. These native idioms complement the `setup` skill; they do not compete with it, and native-first is honored either way. The skill is the interactive, discoverable consumer-configuration face @@ -737,10 +743,10 @@ also its subject. A machine-global or `@latest` install is not itself a reason t `install-typos`: cargo, Homebrew, Conda, pacman or a pre-built binary, reasons 1 and 2) are the current refusals. They stay; they are not defects against a missing subaction. -- **Claim:** install subactions fall into four shapes by what they install (consumer-repo +- **Decision:** install subactions fall into four shapes by what they install (consumer-repo dependency, machine-global CLI, plugin-owned dependencies, hook file); names are not converged; a refusal names the reasons that apply from the list and never excludes a sanctioned subaction. -- **Basis:** #3574. What each live subaction installs: `plugins/ruff-format/skills/setup/SKILL.md` +- **Pointer:** #3574. What each live subaction installs: `plugins/ruff-format/skills/setup/SKILL.md` (`install-ruff`), `plugins/biome-format/skills/setup/SKILL.md` (`install-biome`), `plugins/markdown-format/skills/setup/SKILL.md` (`install-lint`), `plugins/context7/skills/setup/SKILL.md` and `plugins/playwright/skills/setup/SKILL.md` @@ -754,7 +760,7 @@ current refusals. They stay; they are not defects against a missing subaction. `plugins/typos-format/skills/setup/SKILL.md`. Tokens such as `install-hint` and `install-browser` are not setup subactions. - **As of:** 2026-09-29. -- **Recheck:** a setup skill adds an install subaction that fits none of the four shapes, a live +- **Recheck trigger:** a setup skill adds an install subaction that fits none of the four shapes, a live subaction changes what or where it installs, or a maintainer converges the fleet onto one spelling. @@ -857,14 +863,14 @@ Optional platform integrations must degrade visibly and preserve the portable co [Feature availability](https://code.claude.com/docs/en/feature-availability) is this contract's canonical input: fetch it when a platform, provider, or plan question decides something, and restate -none of it here (verified 2026-08-10, [recorded gate runs](#recorded-gate-runs)). A capability the +none of it here (as of 2026-08-10; trigger in [recorded gate runs](#recorded-gate-runs)). A capability the platform itself does not ship on a supported OS is the platform's gap, never the "narrower, inherent platform boundary" a plugin may declare. The plugin still owes a portable path. That input carries two axes, model provider and subscription plan. The *host surface* a consumer runs in, whether CLI, Desktop, an IDE extension, web, or mobile, is a third, read separately from [Platforms and integrations](https://code.claude.com/docs/en/platforms) and the per-host pages it -indexes, cited and never restated (verified 2026-08-10, [recorded gate runs](#recorded-gate-runs)). +indexes, cited and never restated (as of 2026-08-10; trigger in [recorded gate runs](#recorded-gate-runs)). It is a distinct axis because a host can withhold the plugin system itself rather than one capability, and where no plugin loads there is no portable path for one to owe. Host-surface absence is therefore neither a plugin defect nor a boundary a plugin may declare: the OS rule above governs @@ -896,15 +902,16 @@ Every standing instruction this marketplace ships is a per-session tax on every whether or not the instruction ever fires: a CLAUDE.md line, a hook that corrects model behavior, a skill's always-loaded listing text. (Whether a skill's description enters that always-loaded listing at all is the invocation-mode choice, owned by the rubric at -[`docs/conventions/invocation-mode/`](conventions/invocation-mode/README.md).) Official doctrine is explicit: "CLAUDE.md is loaded -every session, so only include things that apply broadly… For each line, ask: 'Would removing this -cause Claude to make mistakes?' If not, cut it," and "If Claude already does something correctly -without the instruction, delete it or convert it to a hook" -([best-practices](https://code.claude.com/docs/en/best-practices), verified 2026-08-10). Anthropic -applied the same doctrine to Claude Code itself, removing over 80% of its system prompt for the -Opus 5 / Fable 5 generation with no measurable loss on its coding evaluations -([The new rules of context engineering for Claude 5 generation models](https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models), -verified 2026-08-08). Context load is not the only budget: the human maintainer's cognitive load +[`docs/conventions/invocation-mode/`](conventions/invocation-mode/README.md).) We keep a standing +line only if removing it would cause a mistake, and we delete or convert to a hook any instruction +the model already follows unaided. Pointer: for pruning standing instructions, see + and +. As of: 2026-08-10. +Recheck trigger: either section moves or stops recommending pruning. Anthropic reports applying +the same doctrine to Claude Code's own system prompt for the Claude 5 generation; no docs page +covers that as of 2026-08-08 +(correlate with ). +Context load is not the only budget: the human maintainer's cognitive load is its sibling constraint, and the write-time doctrine budgeting both lives in `docs-hygiene:write-for-agents`. Four rules follow: @@ -982,15 +989,19 @@ A context that produced work is structurally the weakest place to judge that wor made a mistake plausible is still active, so a self-check inherits the bias. A fresh-context (non-fork) subagent, generic or named, removes it: it starts in its own fresh context window, blind to the reasoning under review. A fork does not: it inherits the parent session's full conversation history, so -it carries the same bias forward -([subagents](https://code.claude.com/docs/en/sub-agents), verified 2026-08-10). -Upstream now states the doctrine, not only the mechanism: a fresh context "improves code review -since Claude won't be biased toward code it just wrote", and a verification subagent exists "so the -agent doing the work isn't the one grading it", and a reviewer in a fresh subagent context "sees only -the diff and the criteria you give it, not the reasoning that produced the change" -([best practices](https://code.claude.com/docs/en/best-practices), verified 2026-08-10). This -section is the authoring-time form of that guidance, applied where an invoker cannot be relied on -to remember it. +it carries the same bias forward. Upstream recommends fresh-context review as well as describing +the mechanism, and this section is the authoring-time form of that guidance, applied where an +invoker cannot be relied on to remember it. + +- **Pointer**: for what a fork inherits, see + ; for + fresh-context review, see + , + , and + . +- **As of**: 2026-08-10 +- **Recheck trigger**: a fork stops inheriting the parent conversation, or those sections stop + recommending a fresh-context reviewer. The rule: **a skill step whose output judges work produced in the same context delegates that judgment to a fresh-context (non-fork) subagent**, generic or named; what the rule requires is the fresh @@ -1036,18 +1047,20 @@ judgment and its target, never re-derives these rules. ### Dispatch ladder The default worker is a **generic fresh-context subagent carrying rich inline instructions**: the -task, the artifact, the criteria, and the output shape all travel in the dispatch prompt. A subagent -starts with a fresh, isolated context window and does not see the parent conversation -([subagents](https://code.claude.com/docs/en/sub-agents), verified 2026-08-10), which is exactly the -independence the checkpoint buys. A skill may prefer an installed **named agent** on the next rung, +task, the artifact, the criteria, and the output shape all travel in the dispatch prompt. A non-fork +subagent starts without the parent conversation (Pointer: +. As of: +2026-08-10. Recheck trigger: a non-fork subagent starts receiving the parent conversation), which +is exactly the independence the checkpoint buys. A skill may prefer an installed **named agent** on the next rung, but only when the named-agent bar below is met, and the site always states the generic fallback (presence-gate-plus-fallback, [seam phrasing](conventions/seam-phrasing/README.md)). The top rung, for high-stakes verdicts where correlated model blind spots are the risk, is a **cross-vendor advisor** when one is installed, on the same presence-gate shape with the same generic fallback. Those rungs are one choice among the platform's parallelism surfaces; -[run agents in parallel](https://code.claude.com/docs/en/agents) is the canonical upstream comparison -of all of them (verified 2026-08-10). Why the fleet takes the subagent rung today rather than agent +[run agents in parallel](https://code.claude.com/docs/en/agents#choose-an-approach) is the +canonical upstream comparison of all of them (as of 2026-08-10; trigger in +[recorded gate runs](#recorded-gate-runs)). Why the fleet takes the subagent rung today rather than agent teams or cross-session messaging is recorded once in the [gate runs](#recorded-gate-runs). Re-derive from that table's triggers instead of re-arguing it at a checkpoint site. @@ -1062,11 +1075,10 @@ A dispatch prompt at any rung: - **degrades when absent**: a preferred named agent or advisor that is not installed routes to the generic fresh-context subagent, never to a command that may not resolve; and - **bounds what counts as a finding**: correctness and the stated requirements, everything else - optional. Upstream names the failure this prevents: "A reviewer prompted to find gaps will - usually report some, even when the work is sound, because that is what it was asked to do", and - chasing all of them "leads to over-engineering" - ([best practices](https://code.claude.com/docs/en/best-practices), verified 2026-08-10). An - unbounded adversarial prompt buys noise at the same price as judgment. + optional. An unbounded adversarial prompt buys noise at the same price as judgment, and chasing + every reported gap over-engineers the work. Pointer: + . As of: + 2026-08-10. Recheck trigger: that section stops recommending a bounded review. ### Named-agent bar @@ -1075,12 +1087,13 @@ multiple sites (or repeats via description-triggered direct invocation) AND a mo pin, or an enforced tool restriction is required.** Otherwise the generic subagent with inline instructions is the simpler, equally independent form. On tool cages: an allowlist that includes Bash bars Edit/Write and recursive spawning but is **not read-only**, because Bash can write. State what the cage -actually enforces, never "read-only" ([plugin agents support `tools` frontmatter](https://code.claude.com/docs/en/plugins-reference), -verified 2026-08-10). +actually enforces, never "read-only" (Pointer: for plugin agent frontmatter, see +. As of: +2026-08-10. Recheck trigger: plugin agents stop supporting `tools`). **Exception: `planning:plan-reviewer`.** -- **Claim:** this agent departs from three defaults on purpose. Bar: it has one dispatch site, +- **Decision:** this agent departs from three defaults on purpose. Bar: it has one dispatch site, `/planning:plan` Step 3, and its description says not to invoke it directly, so the multiple-sites clause is unmet; the pin clause carries it, because a definition is the only way to bound this one review's effort (a generic Agent-tool dispatch has no per-invocation `effort`). For the same @@ -1090,108 +1103,120 @@ verified 2026-08-10). bound it further. Model: it pins `model: opus`; under the fleet's pinned default session, `opus` is the session tier, so it meets the [Model tiers](#model-tiers) rule that a consequential verdict runs at the session-model tier or above. -- **Basis:** the frontmatter of `plugins/planning/agents/plan-reviewer.md` (`model: opus`, +- **Pointer:** the frontmatter of `plugins/planning/agents/plan-reviewer.md` (`model: opus`, `effort: medium`, `maxTurns: 25`) and `plugins/planning/skills/plan/SKILL.md` Step 3; [#4256](https://github.com/melodic-software/claude-code-plugins/issues/4256), which measured a nested plan review at 31.6 minutes and 277k tokens at session effort; the closing comment on [#4849](https://github.com/melodic-software/claude-code-plugins/pull/4849), which kept `opus`. - **As of:** 2026-09-29. -- **Recheck:** a second dispatch site or direct use appears (the exception then ends), the Agent +- **Recheck trigger:** a second dispatch site or direct use appears (the exception then ends), the Agent tool gains a per-invocation `effort` parameter, or the agent's `model` or `effort` changes. ### Model tiers The ladder is relative to the session: **a consequential verdict runs at the session-model tier or above, never below; tedious or mechanical preparation may drop one tier.** The heavy default must be -explicit: an agent definition that omits `model` falls through to `CLAUDE_CODE_SUBAGENT_MODEL` and, -where that is unset, to the main conversation's model, the same model `inherit` selects -([subagents: model resolution](https://code.claude.com/docs/en/sub-agents#choose-a-model): "When -you omit it, Claude Code picks the model in the subagent model order", verified 2026-09-27; frontmatter accepts `sonnet`, `opus`, `haiku`, `fable`, a full model ID, or -`inherit`). Consumers hold one global fallback knob: `CLAUDE_CODE_SUBAGENT_MODEL`, set via the -settings `env` map. It ranks **third**, below the per-invocation `model` parameter and below -frontmatter, so it decides only where neither is set; setting it to `inherit` is the same as leaving -it unset. A structural frontmatter binding therefore holds against it, and the knob is a default for -unbound subagents rather than an override -([subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model), -verified 2026-09-11, recheck when a release note touches subagent model selection; -`env` applies to every session and spawned subprocess, -[settings](https://code.claude.com/docs/en/settings), verified 2026-08-10). - -**Decline `CLAUDE_CODE_SUBAGENT_MODEL_FORCE`.** Claim: do not set -`CLAUDE_CODE_SUBAGENT_MODEL_FORCE`. It forces one model onto every subagent, teammate, and -workflow agent and ignores per-spawn and definition `model` values, which erases the tier ladder -above. Basis: - (`CLAUDE_CODE_SUBAGENT_MODEL_FORCE`, Claude Code -v2.1.257 or later) and -. As of: 2026-09-28. -Recheck: that env-vars row stops ignoring definition and per-spawn `model` values, or the -subagents page stops describing the force switch. - -There is no per-plugin -model surface, because plugin `userConfig` declares only generic typed options with no model semantics -([plugins reference: user configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration), -verified 2026-08-10). Doctrine therefore travels by authoring-time conformance in each skill, not runtime -configuration. - -Tier-to-model mapping, dated 2026-10-01 (recheck trigger: a new Claude model family reaches GA, or -the session default model changes): - -| Tier | Model (2026-10-01) | +explicit: every agent definition in this repository pins `model`, because an agent that omits it +falls through the harness's resolution order and, on a machine with no consumer default, runs on +the main conversation's model. Consumers hold one global fallback knob, `CLAUDE_CODE_SUBAGENT_MODEL`, +set through the settings `env` map. We rely on it ranking below both the per-invocation `model` +parameter and frontmatter, so it decides only for a subagent neither binds: a structural +frontmatter binding holds against it, and the knob is a default for unbound subagents rather than +an override. + +- **Pointer:** for the resolution order and the values frontmatter `model` accepts, see + [subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model); for the + `env` key, see [settings reference: `env`](https://code.claude.com/docs/en/settings-reference#env). +- **As of:** 2026-10-01. +- **Recheck trigger:** a release note touches subagent model selection, or the variable stops + ranking below frontmatter. + +**Decline `CLAUDE_CODE_SUBAGENT_MODEL_FORCE`.** We do not set it, and a consumer who sets it gives +up the tier ladder above, because it puts one model on every subagent regardless of its pin or a +per-spawn `model`. + +- **Pointer:** for the force switch, see + [subagents: run every subagent on one model](https://code.claude.com/docs/en/sub-agents#run-every-subagent-on-one-model) + and its row on [environment variables](https://code.claude.com/docs/en/env-vars#variables). +- **As of:** 2026-10-01. +- **Recheck trigger:** the switch stops overriding definition and per-spawn `model` values, or the + subagents page stops describing it. + +There is no per-plugin model surface: we read plugin `userConfig` as typed options with no model +semantics, so doctrine travels by authoring-time conformance in each skill, not runtime +configuration (Pointer: +[plugins reference: user configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration). +As of: 2026-08-10. Recheck trigger: a `userConfig` option can select the model a plugin's subagent +runs on). + +The tier table names Claude Code aliases, never model versions, so a release that moves an alias +needs no edit here. Which model each alias resolves to is read live from the model page, and it +differs by provider: the same alias can name an older model on a cloud provider's platform than on +the Anthropic API. + +| Tier | Alias | |---|---| -| Consequential verdict (session tier or above) | The active session model; under the fleet's current `opus[1m]` pin that is Opus 5.5, with Fable 5.1 the rung above | -| Mechanical prep, one tier down | Sonnet 5.5 | -| Bulk mechanical sweeps | Haiku 4.5 | - -Row 1 is relative by construction: the invariant above makes the ladder relative to the active -session, so a session already running Fable 5.1 has no rung above and dispatches consequential -verdicts at its own tier. The named models are the resolution under the fleet's pinned session -default (`opus[1m]`, an alias): `opus` resolves to Opus 5.5 on the Anthropic API -([model-config](https://code.claude.com/docs/en/model-config), verified 2026-09-23), the model the -models overview says to "start with … for most workloads", while Fable 5.1 is among "the most -capable models in Claude Code", suited to tasks larger than a single sitting rather than to harder -verdicts at ordinary length. The `fable` alias resolves to Fable 5.1, except in a Claude apps -gateway session, where `fable` and `best` resolve to Fable 5; Fable 5 itself is selected by model -id -([model-config: work with Fable](https://code.claude.com/docs/en/model-config#work-with-fable), -verified 2026-09-28). Opus 5 and Opus 4.8 are legacy models. Row 2 is Sonnet 5.5: the `sonnet` -alias resolves to it on the Anthropic API, and Sonnet 5 is listed as a legacy model -([model-config](https://code.claude.com/docs/en/model-config) and -[models overview](https://platform.claude.com/docs/en/about-claude/models/overview), both -re-read 2026-10-01). Row 3 re-verifies unchanged: Haiku 4.5 remains the current Haiku. -The trigger itself re-tested negative: a further family, Claude Mythos 5, now appears upstream but -has not fired it: Mythos "is not generally available", offered invitation-only to approved -customers under Project Glasswing, so no lane may reach for it. The figures behind the cost ordering -below are upstream-owned -([pricing](https://platform.claude.com/docs/en/about-claude/pricing)) and are not restated here. -([model config](https://code.claude.com/docs/en/model-config), -[models overview](https://platform.claude.com/docs/en/about-claude/models/overview), both verified -2026-08-10.) +| Consequential verdict (session tier or above) | The session's own model, with no `model` passed; under the fleet's `opus[1m]` session pin that is `opus`, with `fable` the rung above | +| Mechanical prep, one tier down | `sonnet` | +| Bulk mechanical sweeps | `haiku` | + +- **Pointer:** for what each alias resolves to on each provider, see + [model config: model aliases](https://code.claude.com/docs/en/model-config#model-aliases). For + where the `sonnet` row's model fits against `opus`, see the + [models overview: latest models comparison](https://platform.claude.com/docs/en/about-claude/models/overview#latest-models-comparison) + and [model config: available models](https://code.claude.com/docs/en/model-config#available-models) + (correlate with ; + recheck when those docs pages cover what the post adds). +- **As of:** 2026-10-01. +- **Recheck trigger:** any new model on Claude Code's model page. + +Row 1 is relative by construction: a session already running the top model has no rung above and +dispatches consequential verdicts at its own tier. The cost ordering behind the rows is +upstream-owned and is not restated here (Pointer: +[pricing: model pricing](https://platform.claude.com/docs/en/about-claude/pricing#model-pricing). +As of: 2026-08-10. Recheck +trigger: any new model on Claude Code's model page). No row binds a model that is not generally +available to the fleet. + +**Override points.** The rows are repository defaults; a personal routing preference belongs in the +operator's own user-scope settings, never in this table. From narrowest to widest reach: a dispatch +site passes the per-invocation `model` (upward only for a verdict, per the rule above); an agent +definition's `model` frontmatter sets that agent's default; `CLAUDE_CODE_SUBAGENT_MODEL` sets the +default for subagents nothing else binds; `ANTHROPIC_DEFAULT_OPUS_MODEL`, +`ANTHROPIC_DEFAULT_SONNET_MODEL`, `ANTHROPIC_DEFAULT_HAIKU_MODEL` and +`ANTHROPIC_DEFAULT_FABLE_MODEL` pin which model an alias resolves to; and an `availableModels` +allowlist bounds every one of them (below). + +- **Pointer:** for the alias-pinning variables, see + [model config: environment variables](https://code.claude.com/docs/en/model-config#environment-variables); + for the allowlist, see + [model config: restrict model selection](https://code.claude.com/docs/en/model-config#restrict-model-selection). +- **As of:** 2026-10-01. +- **Recheck trigger:** the model page adds or removes an alias or an alias-pinning variable. That ladder is a cost ordering, and one capability does not travel down it: **interleaved thinking, -a thinking block between tool calls rather than only before the first and after the last.** Claude Code -models it per model, as the `interleaved_thinking` capability value -([model config: customize pinned model display and capabilities](https://code.claude.com/docs/en/model-config#customize-pinned-model-display-and-capabilities), -verified 2026-08-10; a pinned model's unlisted capabilities are disabled). The per-model roster is -upstream-owned. Resolve it at -[thinking: interleaved thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#interleaved-thinking), -which today states that interleaving is automatic on every model supporting adaptive thinking with -no beta header, and that Claude Haiku 4.5 does not support it (verified 2026-08-10, corroborated by -the model roster's adaptive-thinking column; recheck trigger: a new Haiku generation reaches GA, or -that page's per-model sentence changes). +a thinking block between tool calls rather than only before the first and after the last.** We treat +it as a per-model capability that the bottom tier's alias may lack, and resolve which models have it +from the live roster, never from this file. + +- **Pointer:** for which models interleave, see + [thinking: interleaved thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#interleaved-thinking); + for how Claude Code records the capability for a pinned model, see + [model config: customize pinned model display and capabilities](https://code.claude.com/docs/en/model-config#customize-pinned-model-display-and-capabilities). +- **As of:** 2026-10-01. +- **Recheck trigger:** any new model on Claude Code's model page, or that section's per-model + statement changes. The dispatch consequence, phrased as capability rather than family name so it survives an alias moving under it: **require interleaving only where extended reasoning between tool results decides the next call, meaning a mid-sweep judgment that has to change what gets called next. A task that chains -calls, or that reasons over its results at the end, does not need it.** The boundary is much -narrower than the capability's name suggests, and the same page draws it: "Consecutive tool calls do -not require interleaved thinking. Claude can chain tool calls with or without interleaved thinking; -interleaving changes where thinking blocks appear between tool calls, not whether tool calls can -chain." What the capability adds is a thinking block at that boundary, so what its absence removes is -deliberation *at that point*, not the tool result from context, and not the ability to act on it. -So the bottom tier row stands for bulk mechanical sweeps and for straightforward triage or research -passes that decide at the end; the case it does not cover is a fan-out whose worth is deliberating -partway through, where the next call must change because of what the last one returned. +calls, or that reasons over its results at the end, does not need it.** We read the capability as +adding deliberation at that point, not as a condition for chaining tool calls or for acting on a +tool result (same pointer). So the bottom tier row stands for bulk mechanical sweeps and for +straightforward triage or research passes that decide at the end; the case it does not cover is a +fan-out whose worth is deliberating partway through, where the next call must change because of +what the last one returned. The **dispatch-site** tier enforcement is structural at two binding sites: `plugins/implementation/agents/implementer.md` and @@ -1202,53 +1227,51 @@ recheck list: the trigger above re-audits **every** agent-frontmatter `model` va repository, which `git grep -n '^model:' -- 'plugins/*/agents/*.md'` enumerates rather than any list restated here. -That floor is the consumer's to lose. An enterprise `availableModels` allowlist applies "everywhere -a user can specify a model", frontmatter pins included, and where this document once recorded the -blocked-pin branch as unresolved upstream, upstream now resolves it, per surface and differently for -each. A blocked **subagent** override "falls back to the subagent's inherited model … rather than -failing the request", except that on the Anthropic API and Claude Platform on AWS a blocked *family -alias* instead follows the substitution rule and runs "on the newest permitted version of its -family", a v2.1.222 change the page dates, before which the alias fell back like any other blocked -value. A blocked **skill or command** override behaves differently again: "Claude Code ignores the -override, including a blocked family alias, and the skill or command runs on the session model." - -The earlier derivation's conclusion survives its replacement. A blocked subagent alias can still -land **below** the session, as when the session runs Opus 5.5, the lane is pinned `opus`, and the -allowlist permits only an older Opus. A blocked *cheap* pin lands on the inherited model, which is the session's and -therefore not cheap. So the tier invariant above is still not self-enforcing for a subagent lane: it -may depend on its pin in neither direction, and no error is raised either way. Only the skill and -command branch is now pinned down, and it degrades upward-bounded, to exactly the session model, -never below it. A design whose correctness needs a tier still needs a mechanism that is not a -frontmatter pin -([model config: restrict model selection](https://code.claude.com/docs/en/model-config#restrict-model-selection), -[sub-agents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model), both -verified 2026-08-10; recheck trigger: either page's blocked-override behavior for a subagent, -skill, or command changing). +That floor is the consumer's to lose. An enterprise `availableModels` allowlist reaches frontmatter +pins too, and Claude Code handles a blocked pin differently for a subagent than for a skill or +command. Our conclusion, with the per-surface rules left at the pointer: a subagent lane's tier is +not self-enforcing, because a blocked pin can land below the session (an `opus` pin under an +allowlist that permits only an older Opus) or on the inherited session model (a blocked cheap pin, +which is then not cheap), with no error either way. A skill or command lane with a blocked pin stays +on the session model, never below it. A design whose correctness needs a tier therefore needs a +mechanism that is not a frontmatter pin. + +- **Pointer:** for how a blocked override is handled on each surface, see + [model config: restrict model selection](https://code.claude.com/docs/en/model-config#restrict-model-selection) + and [subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model). +- **As of:** 2026-08-10. +- **Recheck trigger:** either page's blocked-override behavior for a subagent, skill, or command + changes. ### Effort tiers -Effort routes per lane the way model does. Skill and subagent frontmatter `effort` overrides the -session level while that lane is active, but never the `CLAUDE_CODE_EFFORT_LEVEL` environment -variable, and accepts all five level names including `max`; a level the active model does not -support falls back to the highest supported level at or below it -([skills: frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference), -[model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level), -verified 2026-08-10). The ladder itself is upstream-owned, covering level names, per-model -availability, and per-model defaults: resolve it from the model-config page at decision time, never -from this document. - -What a pin actually buys is bounded by how allocation works: thinking is adaptive, so the model -"evaluates each request and decides for itself whether to think and how much", and the caller sets -an intent and optionally the effort while the model "allocates reasoning where it judges reasoning -will help" ([steering thinking](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost), -verified 2026-08-03). A lane pin is therefore a posture, never a switch: a lane pinned `low` still -thinks where the model judges thinking earns its cost, and a turn carrying no thinking is that -mechanism working rather than a pin misfiring. Authoring conformance follows the posture: pin the -lane, then let allocation vary per request instead of writing prose that tries to force it uniform. - -Lane rules, dated 2026-07-29 (recheck trigger: a model change on any pinned lane, or the -model-config effort table changes, since the effort scale is calibrated per model, so the same level -name is not the same underlying value across models): +Effort routes per lane the way model does: a skill or agent pins `effort` in its frontmatter where +its work needs a level other than the session's, knowing the `CLAUDE_CODE_EFFORT_LEVEL` environment +variable and an effort cap still win over the pin. The ladder itself, meaning level names, which +models support effort, each model's default, and what happens to a level a model lacks, is +upstream-owned: resolve it from the model page at decision time, never from this document. + +- **Pointer:** for the levels each model supports and its default, see + [model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level); + for how a frontmatter pin ranks against the session, the variable and a cap, see + [model config: set the effort level](https://code.claude.com/docs/en/model-config#set-the-effort-level); + for the field itself, see + [skills: frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference). +- **As of:** 2026-10-01. +- **Recheck trigger:** any new model on Claude Code's model page, or the effort section changes how + a frontmatter pin ranks. + +We treat a lane pin as a posture, never a switch: thinking is adaptive, so a lane pinned `low` still +thinks where the model judges it worth the cost, and a turn with no thinking is not a pin +misfiring. Authoring conformance follows: pin the lane, then let allocation vary per request +instead of writing prose that tries to force it uniform (Pointer: for how the model decides when to +think, see +[steering thinking: how Claude decides when to think](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#how-claude-decides-when-to-think). +As of: 2026-08-03. Recheck trigger: that page stops describing thinking as decided per request). + +Lane rules (recheck trigger: any new model on Claude Code's model page, or a model change on any +pinned lane, because each model maps level names to its own depth, so one name does not mean the +same depth on two models): - **Consequential-output lanes with a frontmatter surface pin `high`**: verdicts, and research that feeds decisions, wherever the lane is a named agent or a skill doing that work in its own @@ -1256,165 +1279,285 @@ name is not the same underlying value across models): cost (the environment variable still wins, per above). The pin is not relative: on a model whose own default sits above `high`, it caps the lane below that model's default, and the recheck trigger above exists exactly for this. The reach is the mechanism's, not the rule's: a generic - Agent-tool dispatch carries no effort control, because the tool takes a per-invocation `model` - parameter with no effort counterpart - ([sub-agents](https://code.claude.com/docs/en/sub-agents), doc-silence corroborated by the live - tool schema, 2026-07-29), so it structurally inherits the session level and its floor is the + Agent-tool dispatch carries no effort control: we read the live Agent tool schema, which has a + per-invocation `model` parameter and no effort counterpart, and that probe has no stored + artifact (Pointer: for the per-call parameters the docs name, see + [subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model). + As of: 2026-07-29. Recheck trigger: the Agent tool gains an effort parameter), so it + structurally inherits the session level and its floor is the session baseline; promoting such a lane to a named agent is how it gains the pin (a required - effort pin satisfies the named-agent bar's pin clause). `planning:plan-reviewer` pins `medium` - by the [recorded exception](#named-agent-bar), and `implementation:phase-verifier`, - `review:ci-log-auditor`, and `review:doc-drift-detector` pin `medium` because each checks - against binary criteria ([pinned agents](#effort-tiers)). An orchestrator skill - whose consequential work executes in generic dispatches is likewise out of reach: a skill-level - pin governs the orchestrating conversation, and whether it propagates to subagents spawned - while the skill is active is undocumented, so treat propagation as unknown alongside the cache - caveat below. -- **Bulk mechanical sweeps may pin `low`.** Upstream pitches `low` for simpler tasks needing the - best speed and lowest cost, "such as subagents", and lower effort spends fewer tool calls - ([effort](https://platform.claude.com/docs/en/build-with-claude/effort)), but not at the model - ladder's own bottom rung, because the two ladders do not compose there. Effort is a per-model - capability and Haiku has none: "Models not listed here do not support effort", and no Haiku - appears in that table - ([model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level)), - which the model roster corroborates with adaptive thinking off for Claude Haiku 4.5 - ([models overview](https://platform.claude.com/docs/en/about-claude/models/overview), both - verified 2026-08-10). The documented unsupported-level fallback above does not reach this case: - it presupposes a supported level to fall back *to*, and here there is none. What the harness then - does with the pin, whether ignore it, warn, or fail, is **undocumented, and unverified here**; the - pages above establish the absent capability and nothing about the runtime handling, so no reading - of them settles it. The rule does not rest on that gap: a lane wanting the cheapest tier takes it - by model alone and omits the pin, because the dial it would be reaching for only exists one rung - up. -- **Every other lane omits the pin** and inherits the session level: effort is a general - preference, not a task-by-task decision - ([choosing a model and effort level](https://claude.com/blog/claude-model-and-effort-level-in-claude-code)). -- **No lane pins `max` without eval evidence.** Upstream warns it adds significant cost for - relatively small quality gains and can lead to overthinking. Deliberation helps only while - there is still evidence to find; past that point extra effort buys cost and latency and can - degrade the answer ([cut spend without losing quality](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cut-spend-without-losing-quality), - verified 2026-09-09). A pin above `high` (e.g. `xhigh`) - is a deliberate per-lane choice grounded in the target model's own recommended-levels guidance, - never a reflex. The miscalibration cuts the other way too: a lane set too low stops before it - has enough evidence, makes fewer tool calls, and skips the checks it would run unprompted, so - the answer looks finished while resting on partial information (same page, same verification). -- **Sweep model and effort together before raising either.** A stronger model at low effort can - beat a weaker or older model at high effort on both cost and quality, so a lane outgrowing its - level tests the newer model at lower effort before pinning the old one higher, on its own - evals. Cross-model economics and the flat-curve reading live in the fable-5 pack's - model-adaptation chapter for the newer model - (`plugins/playbooks/reference/model-adaptation/fable-5-1.md`, "Cross-model effort economics"); - current prices resolve through the `claude-api` skill at decision time. -- **Effort is the first lever in either direction; steering prose is the second.** Upstream states - the order plainly, to set the effort level matching the lane's workload, then "add prompt guidance - only if Claude's triggering still doesn't match your needs at that level", and gives the - rationale that lowering effort "is usually the better first lever, since it is a calibrated - control rather than a wording-sensitive instruction" - ([steering thinking: effort levels](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#effort-levels), - verified 2026-08-03). Both directions: shallow output from a pinned-`low` lane raises the lane's - effort rather than prompting around it, and a lane thinking more than the work needs lowers the - pin before any prose telling the model to think less. Upstream states that reduce direction - outright and warns it "may reduce quality on tasks that benefit from reasoning". A lane that must - hold its level for latency is the one case that reaches for steering prose first; it then owes - the measurement upstream asks for, a representative sample run with and without the guidance, - compared on trigger rate, output tokens, latency, and quality, because steering effectiveness is - wording-sensitive in a way a level is not. Authoring a lane's prose against its own pin, in - either direction, is the inversion this rule exists to catch. -- **Cache caveat**: changing effort between requests invalidates cached prompt prefixes, so a - skill pin firing mid-session is expected to cost the main conversation's cache (harness-side - request assembly unconfirmed), while a subagent pin is scoped to the subagent's own requests. So - treat skill-lane pins as cache-costly in cost-sensitive loops. State the outcome and not the - mechanism: the platform page and the harness page agree that an effort change forces a full - re-read but describe *why* differently, so an explanation that picks one is asserting more than - either source supports. Two corollaries follow. Setting a lane's effort explicitly to the model's - own default is a no-op that "does not break the cache", so a pin that merely documents the - default costs nothing. And **per-message steering is the cache-safe escape hatch**: guidance - appended to the newest user message "leaves earlier cache breakpoints intact, where a - configuration or effort change does not", which is what makes a skill's invocation-time - instructions cheaper than a mid-session pin. The convention that falls out, and the reason a lane - pin is a design-time choice rather than a per-task one: pick the level once and keep it, steer - per message when one turn needs more or less, and move the configuration only at natural breaks - between tasks ([steering thinking: prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#prompt-caching), - verified 2026-08-03). The harness page states the same convention in its own words, "Pick your - model and effort level at the top of a session, then save `/compact` for natural breaks between - tasks", and adds the interactive consequence a plugin author cannot see from the platform page - alone: once a conversation has started, Claude Code "shows a confirmation dialog before applying - an effort change that would invalidate the cache", so a mid-session change is a prompt the - consumer must clear rather than a silent cost. The same section independently corroborates the - no-op corollary above: a change resolving to the level already in effect "skips the dialog and - keeps the cache" ([prompt caching: changing effort level](https://code.claude.com/docs/en/prompt-caching#changing-effort-level), - verified 2026-08-10). Re-read 2026-09-29: on most models that dialog sentence still - holds, and each effort level has its own cache. On Opus 5.5, Sonnet 5.5, and Fable 5.1, with an - API key or a Claude subscription, changing effort keeps the cache and Claude Code applies the new level - without asking. The exception does not apply on Amazon Bedrock, Google Cloud's Agent Platform, - or a Claude apps gateway, when `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS` is set, or when the - organization has a HIPAA configuration. Before v2.1.260, Fable 5.1 invalidated the cache too - ([prompt caching: changing effort level](https://code.claude.com/docs/en/prompt-caching#changing-effort-level), - fetched 2026-09-29, 42,099 bytes; recheck trigger: that section drops the Opus 5.5 / Sonnet 5.5 / - Fable 5.1 exception or changes which providers it excludes). - -**Pinned `effort: high` agents.** - -- **Claim:** Eleven named agents pin `effort: high` so a session tuned down for cost does not - silently cheapen consequential workers, and four pin `effort: medium`: `plan-reviewer` by its - recorded exception, and `phase-verifier`, `ci-log-auditor`, and `doc-drift-detector` because - each checks against binary criteria. The lowering is owner-decided Option B, narrow, at - `medium` and not `low`, because a low-effort executor stops detecting that it is stuck. - `security-reviewer` and `architecture-guardian` stay `high`. There is no - per-invocation `effort` on Agent-tool dispatch, so a frontmatter pin is what holds a named - agent's lane. The `CLAUDE_CODE_EFFORT_LEVEL` environment variable overrides every pin at once - for the whole session (the environment variable still wins, per above), and a `maxEffortLevel` - or organization effort cap limits any pin above the cap. Both act on the whole session; neither - cited page documents a per-lane or per-plugin lever. -- **Basis:** The agent definitions on origin/main (2026-09-29). `effort: high`: `implementation` - `implementer`; `discovery` `explorer`, `researcher`, `intent-tracer`, and - `research-verifier`; `review` `code-reviewer`, `architecture-guardian`, - `ecosystem-specialist`, and `security-reviewer`; `plugin-quality` - `auditor`; `songwriting` `object-writer`. `effort: medium`: `implementation` `phase-verifier`; - `review` `ci-log-auditor` and `doc-drift-detector`; `planning` `plan-reviewer`. Issue - [#4253](https://github.com/melodic-software/claude-code-plugins/issues/4253) is the source of - the filed list of eleven, which omits `auditor`, `object-writer`, and `research-verifier`. The - Agent-tool gap is stated in this section ("a generic Agent-tool dispatch carries no effort - control"). Upstream, fetched 2026-09-29 from the raw `.md` channel: - [model config](https://code.claude.com/docs/en/model-config#set-the-effort-level) (109,848 - bytes), "Frontmatter effort applies when that skill or subagent is active, overriding the - session level but not the environment variable. A `maxEffortLevel` or organization effort cap - still limits the level the skill or subagent runs at"; - [sub-agents](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields) (107,466 - bytes), `effort`: "Effort level when this subagent is active. Overrides the session effort - level." Lowering effort beat an architecture change, and `low` is named for simpler subagent - tasks ([optimizing for cost and intelligence](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence), - [effort](https://platform.claude.com/docs/en/build-with-claude/effort), same fetch date). + effort pin satisfies the named-agent bar's pin clause). The named agents that pin `medium` + instead, and why, are listed under [pinned agents](#effort-tiers). An orchestrator skill + whose consequential work executes in generic dispatches reaches them through its own pin: a + skill-level pin governs the orchestrating conversation and, by our probe, the subagents it + dispatches (the skill-pin record under [override levers](#effort-tiers)); no docs page covers + that reach, so the cache caveat below still applies. +- **Read-only bulk mechanical sweeps may pin `low`.** We allow it where speed and cost matter more + than depth, subagent sweeps included, and never for a lane that changes code or verifies a change + (the [effort floor](#effort-floor)). Not at the model ladder's own bottom rung either, because the + two ladders do not compose there: we read the model the `haiku` alias resolves to as having no + effort support, so a pin there has no level to land on. What the harness does with such a pin, + whether ignore it, warn, or fail, is **unverified here**, and no page we read settles it. The rule + does not rest on that gap: a lane wanting the cheapest tier takes it by model alone and omits the + pin, because the dial it would be reaching for only exists one rung up (Pointer: for which models + support effort, see + [model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level); + for what a lower level trades away, see + [effort: how effort works](https://platform.claude.com/docs/en/build-with-claude/effort#how-effort-works). + As of: 2026-10-01. Recheck trigger: a Haiku model appears among the models that support effort). +- **Every other named agent pins the level its work's task row gives it**, never below `medium` + for code-changing or verifying work (the [effort floor](#effort-floor); the per-pin rows under + pinned agents name each row). Only a skill with no consequential output omits the pin and + inherits the session level. A pin is a design-time choice made once for that lane; the session + level belongs to the user, who overrides a pin through the [override levers](#effort-tiers). + - **Pointer:** for which level fits which kind of work, see + [model config: choose an effort level](https://code.claude.com/docs/en/model-config#choose-an-effort-level); + the two posts under the source conflict below are correlate notes only. + - **As of:** 2026-10-01. + - **Recheck trigger:** that section stops matching levels to kinds of work, or starts naming one + level for every task. + - **Source conflict:** the older effort post, + (correlate only), and + model config's choose-an-effort-level section, joined by the newer effort post, + (correlate only), disagree on whether effort is + a general preference or a per-task choice. As of: 2026-10-01. Recheck trigger: either post or + that section is revised on that point. +- **No lane pins `max` without eval evidence.** We treat it as the costliest level, whose gain a + lane must measure before using it. A pin above `high` (e.g. `xhigh`) is a deliberate per-lane + choice grounded in the target model's own recommended levels, never a reflex. We watch for + miscalibration in both directions: set too high, a lane keeps spending after the evidence runs + out; set too low, it stops early and skips checks, and its answer looks finished on partial + information (Pointer: for each level's trade-off, see + [effort: effort levels](https://platform.claude.com/docs/en/build-with-claude/effort#effort-levels); + for miscalibration, see + [cut spend without losing quality](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cut-spend-without-losing-quality). + As of: 2026-09-09. Recheck trigger: either section changes what it says a level above `high` + buys). +- **Sweep model and effort together before raising either.** A lane outgrowing its level tests the + newer model at lower effort before pinning the old one higher, on its own evals. Cross-model + economics and the flat-curve reading live in the fable-5 pack's model-adaptation chapter for the + newer model (`plugins/playbooks/reference/model-adaptation/fable-5-1.md`, "Cross-model effort + economics"); current prices resolve through the `claude-api` skill at decision time. +- **Effort is the first lever in either direction; steering prose is the second.** We set the + level to fit the lane's work and add prose only where that level still falls short (Pointer: + for why, see + [steering thinking: effort levels](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#effort-levels). + As of: 2026-08-03. Recheck trigger: that section stops ranking effort ahead of prompt + steering). Both directions: shallow output from a pinned-`low` lane raises the lane's + effort instead of adding prompt text, and a lane thinking more than the work needs lowers the + pin before any prose telling the model to think less. A lane that must hold its level for latency + is the one case that reaches for steering prose first, and it keeps that prose only after + comparing a sample of runs with and without it, because wording moves the result less + predictably than a level does. Authoring a lane's prose against its own pin, in either + direction, is the inversion this rule exists to catch. +- **Cache caveat.** For whether an effort change keeps the cache on a given model and route, read + [prompt caching: changing effort level](https://code.claude.com/docs/en/prompt-caching#changing-effort-level) + first. Where that section says the cache is lost, we pick a lane's level once and keep it, steer + per message when one turn needs more or less, and move configuration only between tasks. Where + it says the cache is kept, that rule does not apply, but we design no lane that depends on that + path. Two expectations hold either way: a skill pin firing mid-session counts as cache-costly + for the main conversation, since how the harness assembles that request is unconfirmed, while a + subagent pin touches only the subagent's own requests; and a pin equal to the model's own + default changes nothing, so a pin that only documents the default costs nothing. + - **Pointer:** for Claude Code's handling, see the section above; for the API side, see + [steering thinking: prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#prompt-caching). + - **As of:** 2026-10-01. + - **Recheck trigger:** either section changes what an effort change does to the cache, or which + models, providers or versions keep the cache across one. + +**Pinned agents.** Every named agent in this repository pins its effort, so a session tuned down for +cost does not silently cheapen a worker. Each pins the level that model config's task rows give its +kind of work, and never below `medium` for work that changes code or verifies a change (the +[effort floor](#effort-floor)). Eleven pin `effort: high`: `implementation` `implementer`; +`discovery` `researcher`, `intent-tracer`, and `research-verifier`; `review` `code-reviewer`, +`architecture-guardian`, `security-reviewer`, `ci-log-auditor`, and `doc-drift-detector`; +`plugin-quality` `auditor`; `songwriting` `object-writer`. Four pin `effort: medium`: `planning` +`plan-reviewer` by its [recorded exception](#named-agent-bar); `implementation` `phase-verifier`, +because it checks one phase against acceptance criteria fixed before it runs; and `review` +`ecosystem-specialist` and `discovery` `explorer`, because their work is clearly scoped tool use, +running a repository's declared commands and reading and indexing a scope. No pin goes below +`medium`, because a low-effort executor stops detecting that it is stuck. A frontmatter pin is what +holds a named agent's lane, since an Agent-tool dispatch passes no effort. + +- **Pointer:** the agent definitions themselves, listed by + `git grep -n '^effort:' -- 'plugins/*/agents/*.md'`; + [#4253](https://github.com/melodic-software/claude-code-plugins/issues/4253) for the filed pin + list; for the task rows, see + [model config: choose an effort level](https://code.claude.com/docs/en/model-config#choose-an-effort-level); + for how a frontmatter pin ranks against the session, see + [model config: set the effort level](https://code.claude.com/docs/en/model-config#set-the-effort-level) + and the `effort` field in + [subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields). +- **As of:** 2026-10-01. +- **Recheck trigger:** any new model on Claude Code's model page, or a pinned agent's `model` + changes, since a level name means a different depth on each model; the task rows change; a + checker pinned `medium` misses a defect a `high` pin caught; or a maintainer changes or drops a + named pin. +- **Source conflict:** the newer effort post, + (correlate only), and model config's + choose-an-effort-level section disagree on which level fits implementation work. We follow model + config, because a docs section outranks a blog post under the + [upstream-drift convention](conventions/upstream-drift/README.md#required-parts). As of: + 2026-10-01. Recheck trigger: either page is revised on that point. + +**Per-pin rows.** Each pinned agent outside `plugins/implementation` follows one row of model +config's effort-level table, read for the model its `model` alias resolves to. Review, verification +and verdict lanes follow the `high` row; well-specified mechanical work follows the `medium` row. + +| Agent | Model | Pin | Row | Why that row | +|---|---|---|---|---| +| `discovery` `explorer` | `sonnet` | `medium` | `medium` | Reads and indexes a scope it is handed | +| `discovery` `intent-tracer` | `opus` | `high` | `high` | Reconstructed rationale feeds decisions | +| `discovery` `research-verifier` | `opus` | `high` | `high` | Verdict on a research artifact | +| `discovery` `researcher` | `opus` | `high` | `high` | Research that feeds decisions | +| `planning` `plan-reviewer` | `opus` | `medium` | `medium` | A review lane held to `medium` by its [recorded exception](#named-agent-bar), not by the review rule | +| `plugin-quality` `auditor` | `opus` | `high` | `high` | Audit verdict | +| `review` `architecture-guardian` | `opus` | `high` | `high` | Review verdict | +| `review` `ci-log-auditor` | `sonnet` | `high` | `high` | Audit verdict on a CI run | +| `review` `code-reviewer` | `sonnet` | `high` | `high` | Review verdict | +| `review` `doc-drift-detector` | `sonnet` | `high` | `high` | Drift verdict | +| `review` `ecosystem-specialist` | `sonnet` | `medium` | `medium` | Runs a repository's declared build, test and lint commands | +| `review` `security-reviewer` | `opus` | `high` | `high` | Security verdict | +| `songwriting` `object-writer` | `opus` | `high` | `high` | Creative generation, which no row names; the `high` choice is our judgment | + +- **Pointer:** for the rows, see + [model config: choose an effort level](https://code.claude.com/docs/en/model-config#choose-an-effort-level); + for the model each alias resolves to, see + [model config: model aliases](https://code.claude.com/docs/en/model-config#model-aliases); for + each model's default level, see + [model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level). +- **As of:** 2026-10-02. +- **Recheck trigger:** a row named in the table changes, the default effort of the model an agent's + alias resolves to changes, or an agent's `model` or `effort` changes. + +**Override levers.** We name two levers for a user who wants a pinned agent at another level. +`CLAUDE_CODE_EFFORT_LEVEL` sets one level for a whole session and replaces every pin. A Workflow +script's `agent()` call passes `opts.effort`, and `opts.model`, for that call alone. A +`maxEffortLevel` setting or an organization effort cap also limits every pin. + +- **Pointer:** for the variable, see + [environment variables](https://code.claude.com/docs/en/env-vars#variables); for how a pin ranks + against the variable and a cap, see + [model config: set the effort level](https://code.claude.com/docs/en/model-config#set-the-effort-level). +- **As of:** 2026-10-01. +- **Recheck trigger:** the variable stops replacing a frontmatter pin, or the Agent tool gains a + per-invocation `effort` parameter. + +Our Workflow probe: an explicit `opts.effort` or `opts.model` on an `agent()` call overrode the +named agent's frontmatter pin for that call, and omitting them kept the pin. No docs section +covers per-call effort for `agent()`; for the call itself, see +[workflows: what the saved script looks like](https://code.claude.com/docs/en/workflows#what-the-saved-script-looks-like). + +- **Pointer:** probes `wf_1a471686-8a2` and `wf_5196f26b-e1f`, run on Claude Code 2.1.284 and + 2.1.285; no artifact is stored in this repository. +- **As of:** 2026-10-01. +- **Recheck trigger:** a docs page starts covering per-call effort for a workflow `agent()` call or + for the Agent tool. + +**Where per-task effort is set.** We set a task's effort through Workflow's per-call effort +option. An Agent-tool dispatch of a named agent runs at that agent's pin, and a generic one at the +session level. A lane's `--effort` sets the session level, so it covers the orchestrator's own +turns and its generic dispatches; it does not move a named agent's pin. Workflow scripts this +repository ships never pass an effort below a named agent's pin, and omit effort on a call to a +named agent to keep the pin. A call that names no agent passes an explicit level. + +- **Pointer:** for the `--effort` flag, see + [CLI reference: CLI flags](https://code.claude.com/docs/en/cli-reference#cli-flags); for a + subagent's `effort` field and its rank over the session level, see + [subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields). + No docs page covers per-call effort for a Workflow `agent()` call as of 2026-10-02; our Workflow + probe above is the record. +- **As of:** 2026-10-02. +- **Recheck trigger:** a docs page starts covering it, the Agent tool gains an effort parameter, or + either section above changes how the flag or the field ranks. + +**A skill's pin reaches the subagents it dispatches.** We treat a skill's frontmatter `effort` pin +as applying to the main turns while that skill is active and to the subagents it dispatches. + +- **Pointer:** our probe, two headless sessions on Claude Code 2.1.285. Session + `ec6654fe-076a-4c3d-970c-363339cb63b3` ran a throwaway plugin skill pinned `effort: high` under + `--effort low` and ran at `high` on three main turns and two subagent turns; control session + `908560cb-3f30-4157-a672-5e00112bba0a` ran without the pin and stayed at `low` throughout. The + throwaway plugin no longer exists, so the session ids are the artifact. For the skill field, see + [skills: frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference); + no docs page covers the pin's reach into dispatched subagents. - **As of:** 2026-09-29. -- **Recheck:** a checker pinned `medium` misses a defect its `high` pin caught, the Agent tool gains - a per-invocation `effort` parameter, a maintainer lowers or drops a named pin, or a plugin ships a `userConfig` effort key that actually reaches the worker. - -**Effort is one dial of two, and the other is not an effort value.** The `thinking` parameter decides -whether Claude reasons in thinking blocks; `effort` decides how hard the whole response works, -"which in adaptive mode includes how often and how deeply it thinks". Upstream states the resulting -trap outright, "Don't pass `adaptive` as an `effort` value: `adaptive` is a thinking mode, not an -effort level", and a frontmatter `effort` field is exactly where that trap is reachable, because the -two dials share vocabulary. The second consequence bounds what any pin can promise, in upstream's own -words: "**You need a hard ceiling on spend:** use `max_tokens`. Effort is soft guidance; `max_tokens` -is a strict limit." Read what that limit bounds before reaching for it. `max_tokens` is a request -parameter capping one response's output, and it "includes all thinking Claude generates in the current -turn", so it binds per response and constrains neither input and cache reads nor the further -requests an agentic lane makes. **And no documented frontmatter field reaches it.** Those fields set -the model and the effort level, and a subagent adds `maxTurns`, which bounds agentic turns rather -than tokens and has no skill-frontmatter counterpart; neither field list carries a token cap, because -the parameter belongs to the API request that the lane-pin surface does not assemble. So the rule -this section can actually state is narrower than the quote: a lane wanting to spend less lowers -`effort` knowing it is guidance, and a hard cap has to be imposed by whoever builds the request -([thinking and effort](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-effort), -[subagent frontmatter](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields), and -[skill frontmatter](https://code.claude.com/docs/en/skills#frontmatter-reference), all verified -2026-08-03; recheck trigger: the accepted `effort` value set changes on the model-config or effort -page, or either documented frontmatter field list gains a token cap). Checking the value set mechanically stays deferred: a lint rule's source of truth is -the harness's own accepted-value list, which this section deliberately does not restate. - -Session-level effort is the consumer's own knob, out of plugin scope: `low` through `xhigh` -persist via the `effortLevel` setting, while `max` and `ultracode` are session-only, and `max` is -durable only through the `CLAUDE_CODE_EFFORT_LEVEL` environment variable. Plugins never set -session effort. +- **Recheck trigger:** the Claude Code CLI version moves past 2.1.285, or a docs page starts + covering pin propagation to subagents. + +**Configurability gap.** No per-agent user setting exists: a user cannot move one named agent's pin +without editing its definition, and we have not confirmed that a plugin `userConfig` value can +reach an agent's `effort` field. We record this as a gap against +[configuration ownership](#configuration-ownership-and-scope), not as a design choice. + +- **Pointer:** for plugin options, see + [plugins reference: user configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration); + for the agent field, see + [subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields). +- **As of:** 2026-10-01. +- **Recheck trigger:** Claude Code adds a per-agent effort setting, or a docs page or a probe shows + a `userConfig` value reaching a subagent's `effort`. + +**Effort is one dial of two, and the other is not an effort value.** We keep the `thinking` mode and +the `effort` level apart: `adaptive` is a thinking mode, never an `effort` value, and a frontmatter +`effort` field is where that mix-up is reachable, because the two share vocabulary. A pin also +promises less than a spend ceiling. We treat effort as guidance and the request's `max_tokens` as +the hard ceiling, and that ceiling binds one response, not input, cache reads, or the further +requests an agentic lane makes. No documented skill or subagent frontmatter field reaches it: +`maxTurns` bounds agentic turns, not tokens, and has no skill counterpart. So a lane wanting to +spend less lowers `effort` knowing it is guidance, and a hard cap is imposed by whoever builds the +request. + +- **Pointer:** for how thinking and effort relate, see + [thinking and effort](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-effort); + for the frontmatter fields, see + [subagent frontmatter](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields) + and [skill frontmatter](https://code.claude.com/docs/en/skills#frontmatter-reference). +- **As of:** 2026-08-03. +- **Recheck trigger:** the accepted `effort` value set changes on the model-config or effort page, + or either frontmatter field list gains a token cap. + +Checking the value set mechanically stays deferred: a lint rule's source of truth is the harness's +own accepted-value list, which this section deliberately does not restate. + +Session-level effort is the consumer's own knob, out of plugin scope: plugins never set session +effort. A plugin may advise a session level for a phase of work, but it never changes the user's +session level or saved level (Pointer: for how a level is set and saved, see +[model config: set the effort level](https://code.claude.com/docs/en/model-config#set-the-effort-level); +for the ultracode setting, see +[model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level). +As of: 2026-10-01. Recheck trigger: any new model on Claude Code's model page, or a new way to set +or save a level). + +One exception, scoped to frontmatter: a skill or agent may pin `effort`, because the pin holds only +while that skill or agent is active and never changes the user's session or saved level. + +- **Pointer:** for the skill field, see + [skills: frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference); + for the agent field, see + [subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields). +- **As of:** 2026-10-02. +- **Recheck trigger:** either field starts persisting a level past the active skill or agent, or + stops overriding the session level. + +### Effort floor + +Code-changing or verifying work runs at `medium` effort or above, on every model that supports +effort. Where this repository sets effort, in an agent's or a skill's frontmatter, a lane that +changes code or verifies a change never pins below `medium`; `low` is only for chat-like exchanges +and read-only mechanical work. The floor is stated as a level, not a model, because an alias can +resolve to a model with a different default, or with no effort support, depending on the provider +and the release. + +- **Pointer:** for which models support effort and each one's default, see + [model config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level); + for what each level fits, see + [model config: choose an effort level](https://code.claude.com/docs/en/model-config#choose-an-effort-level); + for what each alias resolves to per provider, see + [model config: model aliases](https://code.claude.com/docs/en/model-config#model-aliases). +- **As of:** 2026-10-01. +- **Recheck trigger:** any new model on Claude Code's model page, or the effort section changes + which models support effort. ### Declared patterns @@ -1432,10 +1575,10 @@ cannot assume a plugin layout. The complete categorized index of plugin-relevant official pages is [`docs/official-docs.md`](official-docs.md); `https://code.claude.com/docs/llms.txt` is the -authoritative self-updating master list. The Claude Code pages this document rests on, each -re-fetched 2026-08-10 and confirmed to still carry the topics named beside it (the -`melodic-software/standards` entry below is not a Claude Code page and was not re-checked on that -date): +authoritative self-updating master list. Each list below names the pages this document rests on +and the topics we read each one for. Recheck trigger for both lists: a page moves, or stops +covering a topic named beside it. The first list is as of 2026-08-10 (the +`melodic-software/standards` entries are not Claude Code pages and carry no date): - [Create plugins](https://code.claude.com/docs/en/plugins): plugin structure incl. `bin/` and plugin `settings.json`, namespaces, testing, and migration. @@ -1454,7 +1597,7 @@ date): - `melodic-software/standards` engineering philosophy and cross-platform review criteria: repository design and verification policy. -Verified 2026-07-17: +As of 2026-07-17: - [Plugin dependencies](https://code.claude.com/docs/en/plugin-dependencies): the `dependencies` array, automatic installation, and version constraints. diff --git a/docs/upstream/claude-code-mods/sources.md b/docs/upstream/claude-code-mods/sources.md index 0add4896f1..e64a755528 100644 --- a/docs/upstream/claude-code-mods/sources.md +++ b/docs/upstream/claude-code-mods/sources.md @@ -17,35 +17,38 @@ Asked "what about mods?", do this in order: ## How to read the table -Everything below was fetched or re-fetched on **2026-09-19**, at Claude Code **2.1.278**, with -`anthropics/claude-code` `mods/` at **`92ec78f2`**. Where a row's fetched date differs, the row says -so. "Used by" names the ADR, the runbook, or the report sidecar under `research-2026-09-19/` that -relies on the source. Trust tiers run first-party first; a claim's own basis label -(`OBSERVED`, `SOURCE`, `BINARY`, `STAFF`, `COMMUNITY`, `INFERRED`) is carried in the report, not -here. +Each row is a pointer: the link, the topic we used it for (in our words; the source's own text is +read at the link, never stored here), the date it was last read, and what in this repository +relies on it. Everything below was fetched or re-fetched on **2026-09-19**, at Claude Code +**2.1.278**, with `anthropics/claude-code` `mods/` at **`92ec78f2`**. Where a row's date differs, +the row says so. Counts and hit totals are our own probe results. "Used by" names the ADR, the +runbook, or the report sidecar under `research-2026-09-19/` that relies on the source. Trust tiers +run first-party first; a claim's own basis label (`OBSERVED`, `SOURCE`, `BINARY`, `STAFF`, +`COMMUNITY`, `INFERRED`) is carried in the report, not here. Recheck trigger for every row: a +[go-no-go.md](go-no-go.md) run re-fetches it, and the row's date is refreshed with the outcome. ## First-party source tree and binary -| Link | What it establishes | Fetched | Used by | +| Link | What we used it for | As of | Used by | |---|---|---|---| | | The published source of the four shipped mods, `types/`, `tsconfig.json`. The surface link for the gate-run row. | 2026-09-19 | ADR 0035, `plugin-philosophy.md`, `research-what-mods-are.md` | -| | Definition of a mod, the four mods and their seating, the testing kit, noun contracts, and the early-access statement "may change between releases without notice". | 2026-09-19 | ADR 0035 (criterion 5), go-no-go.md, `research-what-mods-are.md` | +| | What a mod is, the four mods and their seating, the testing kit, noun contracts, and the early-access stability statement criterion 5 checks for. | 2026-09-19 | ADR 0035 (criterion 5), go-no-go.md, `research-what-mods-are.md` | | | The same file pinned at the research commit, so a later diff is exact. | 2026-09-19 | go-no-go.md | | | Raw form, for the greppable criterion-5 check. | 2026-09-19 | go-no-go.md | | | Raw form pinned at the research commit. | 2026-09-19 | `research-what-mods-are.md` | -| | The declarations: 12,990 lines, first line `// Written by Claude Code 2.1.277.`, the event and noun catalog, the five tiers, `HookBudget`, `Registration.catch`. | 2026-09-19 | ADR 0035, `research-api-surface.md`, `research-what-mods-are.md` | +| | The declarations: the event and noun catalog, the tiers, `HookBudget`, `Registration.catch`. Our count: 12,990 lines; its first line names Claude Code 2.1.277 as the writer. | 2026-09-19 | ADR 0035, `research-api-surface.md`, `research-what-mods-are.md` | | | Raw form, for the regeneration check (first line names the CLI that wrote it). | 2026-09-19 | go-no-go.md | -| | Raw form pinned at the research commit; the counted catalog (38 + 54 events, 33 `classic.*`, 19 nouns). | 2026-09-19 | `research-api-surface.md` | -| | The `instructionFiles` option and its three modes, naming `claude-md-or-agents-md` as the default and stating that a project whose own `CLAUDE.md` the engine already loaded keeps the plugin out. | 2026-09-19 | ADR 0035 | -| | `sec-default` constrains by ordering, not policy; `prependPlugins` seating; `next.to` refused outside a managed tier. | 2026-09-19 | `research-security-and-semantics.md`, `research-what-mods-are.md` | +| | Raw form pinned at the research commit; our counted catalog (38 + 54 events, 33 `classic.*`, 19 nouns). | 2026-09-19 | `research-api-surface.md` | +| | The `instructionFiles` option, its modes and default, and when the plugin stays out of a project that already loads its own `CLAUDE.md`. | 2026-09-19 | ADR 0035 | +| | How `sec-default` constrains, `prependPlugins` seating, and where `next.to` is allowed. | 2026-09-19 | `research-security-and-semantics.md`, `research-what-mods-are.md` | | | The same, pinned. | 2026-09-19 | `research-security-and-semantics.md` | | | The manifest shape a mod-bearing `hooks.json` takes (the `modules` key). | 2026-09-19 | `research-authoring-and-testing.md`, `research-repo-fit.md` | -| | A second manifest instance; `telemetry` is narrowed to internal builds. | 2026-09-19 | `research-authoring-and-testing.md` | +| | A second manifest instance, and which builds `telemetry` targets. | 2026-09-19 | `research-authoring-and-testing.md` | | | The `engine.create` noun-fold pattern a mod uses to add a noun to `$`. | 2026-09-19 | `research-what-mods-are.md` | -| | Commit history of the tree: first commit `d9c456d7` 2026-09-09, 83 commits at `92ec78f2`. | 2026-09-19 | `research-api-surface.md` | +| | Commit history of the tree. Our count: first commit `d9c456d7` 2026-09-09, 83 commits at `92ec78f2`. | 2026-09-19 | `research-api-surface.md` | | | The same history through the API, for the "has `mods/` moved?" check. | 2026-09-19 | go-no-go.md | | | The paginated form used to count the 83 commits. | 2026-09-19 | `research-api-surface.md` | -| | The `mods/` commit "sparing six of the files a hooks module may link". | 2026-09-19 | `research-contradictions-and-corrections.md` | +| | The `mods/` commit on which files a hooks module may link. | 2026-09-19 | `research-contradictions-and-corrections.md` | | | Release `v2.1.278`, published 2026-09-19T03:10:40Z: the version every `OBSERVED` and `BINARY` claim is pinned to. | 2026-09-19 | go-no-go.md | | | The same through the API. | 2026-09-19 | `research.md` | | | A single release lookup used while dating the version floor claims. | 2026-09-19 | `research-enablement-and-distribution.md` | @@ -65,107 +68,107 @@ community posts and sit in the Community table. All five are marked **seed**. `p disclosed inference, not by confirmed metadata; everything attributed to that account is from GitHub, never from X. -| Link | What it establishes | Fetched | Used by | +| Link | What we used it for | As of | Used by | |---|---|---|---| -| | **Seed.** Boris Cherny, 2026-09-14: "Claude Mods are landing now. Someone already built a Tetris-in-Claude mod". | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | **Seed.** Thariq, 2026-09-18: "AGENTS.md support is built off of Claude Code mods, our upcoming way to customize the Claude Code harness." | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | Official account, 2026-09-03: the Function Hooks announcement, "It hasn't shipped yet". | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | bcherny, 2026-09-03: "an early look at how we're thinking about making Claude Code way more extensible". | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | The first post of the 2026-09-18 AGENTS.md thread: support starting in 2.1.277. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | **Seed.** Boris Cherny, 2026-09-14: the mods landing announcement and a community Tetris mod. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | **Seed.** Thariq, 2026-09-18: AGENTS.md support and its relation to mods. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | Official account, 2026-09-03: the Function Hooks announcement and its shipped status at that date. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | bcherny, 2026-09-03: the early look at extensibility. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | The first post of the 2026-09-18 AGENTS.md thread, and the version it names. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | An earlier Claude Code post by the same staff account, swept for mods mentions. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | The 10-page design document attached to issue #91870, byline "Alice Poteat · August 2026 · Anthropic". | 2026-09-19 | `research-security-and-semantics.md` | -| | The cheat sheet attached to the Community Update, enumerating affordances on v267/v268. Served as SVG despite the `.png` markdown. | 2026-09-19 | `research-surfaces.md` | +| | The 10-page design document attached to issue #91870, and its byline. | 2026-09-19 | `research-security-and-semantics.md` | +| | The cheat sheet attached to the Community Update, listing affordances on v267/v268. We observed it served as SVG despite the `.png` markdown. | 2026-09-19 | `research-surfaces.md` | ### Comments in issue #91870 cited by permalink -Each is a verbatim quote's permalink. All but one are `poteat`'s; the exception is noted in the -GitHub issues table below, because its author is not staff. The issue body is edited in place, so -re-read the body as well as the comments. +Each row is a permalink to a comment the report relies on. All but one are `poteat`'s; the +exception is noted in the GitHub issues table below, because its author is not staff. The issue +body is edited in place, so re-read the body as well as the comments. -| Link | What it establishes | Fetched | Used by | +| Link | What we used it for | As of | Used by | |---|---|---|---| -| | `$` mediates everything; no ambients, so an admin can audit, allowlist, deny or log any event. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | The onion model: the plugin registered first owns subsequent hooks on an event instance. | 2026-09-19 | `research-what-mods-are.md` | -| | `tool.call` is planned to cover MCP and non-MCP tool calls alike. | 2026-09-19 | `research-api-surface.md` | -| | What the model sees when a hook rewrites; calling `next` twice is supported. | 2026-09-19 | `research-api-surface.md` | -| | The performance claim: in-process on Bun, "a p99 of 50μs per hook". | 2026-09-19 | `research-security-and-semantics.md` | +| | How `$` mediates events, and what that lets an admin audit, allowlist, deny, or log. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | The onion model of hook ownership by registration order. | 2026-09-19 | `research-what-mods-are.md` | +| | The planned coverage of `tool.call` across MCP and non-MCP tools. | 2026-09-19 | `research-api-surface.md` | +| | What the model sees when a hook rewrites, and calling `next` more than once. | 2026-09-19 | `research-api-surface.md` | +| | The in-process performance claim. | 2026-09-19 | `research-security-and-semantics.md` | | | First of four recorded positions on fail-open and re-dispatch; also the surface roadmap. | 2026-09-19 | `research-security-and-semantics.md`, `research-surfaces.md` | -| | Second position: leaning to route-around on failure. | 2026-09-19 | `research-security-and-semantics.md` | -| | Packaging unchanged: "dependencies, versions, etc. will all go unchanged"; the static scan; OTEL intent. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | Second position on failure handling. | 2026-09-19 | `research-security-and-semantics.md` | +| | Packaging, the static scan, and OTEL intent. | 2026-09-19 | `research-enablement-and-distribution.md` | | | First mention of `/plugin-types`. | 2026-09-19 | `research-authoring-and-testing.md` | -| | Isolation is "a boundary" but not part of the contract; and `deny` after `next(e)` does not un-run. | 2026-09-19 | `research-security-and-semantics.md` | -| | Third position: the three skip cases; a declarative fail-closed policy declined. | 2026-09-19 | `research-security-and-semantics.md` | -| | "`validate` only syntactically checks your plugin." | 2026-09-19 | `research-authoring-and-testing.md` | -| | A `tool.call` hook sees tool calls made inside a subagent. | 2026-09-19 | `research-api-surface.md` | -| | `$.process.run` and `$.http.fetch` as the non-TypeScript escape hatches. | 2026-09-19 | `research-api-surface.md` | -| | The `classic.*` bridge events and the five tiers `[prepend] [user] [append] [builtin] [core]`. | 2026-09-19 | `research-what-mods-are.md` | -| | Fourth position: the `.catch` proposal that the Sep 9 cheat sheet then ships. | 2026-09-19 | ADR 0035, `research-security-and-semantics.md` | -| | `prompt.submit` is for text only, not command execution. | 2026-09-19 | `research-api-surface.md` | -| | `/plugin-types` lists the available `$` affordances once the flag is enabled. | 2026-09-19 | `research-authoring-and-testing.md` | -| | Why `tool.check` guards can run concurrently and `tool.call` guards cannot. | 2026-09-19 | `research-security-and-semantics.md` | -| | No Emacs-advice-style sugar; only `turn.step` may yield. | 2026-09-19 | `research-api-surface.md` | -| | 2026-09-16 restatement of `on(...).catch(...)` as the fail-closed spelling. | 2026-09-19 | `research-security-and-semantics.md` | -| | AGENTS.md shipped as a built-in mod. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | 2026-09-19: "mods cannot dictate their own registration order, full-stop"; the four-name tier list. | 2026-09-19 | `research-what-mods-are.md` | -| | The comment bcherny's "latest community update" link resolves to. Authored by `sezaakgun`, `author_association: NONE`: a community Tetris demo, not an Anthropic update. Tiering by "linked from a staff post" misfiles it. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | Isolation's place in the contract, and `deny` after `next(e)`. | 2026-09-19 | `research-security-and-semantics.md` | +| | Third position: the skip cases, and the declarative fail-closed policy question. | 2026-09-19 | `research-security-and-semantics.md` | +| | What `validate` checks. | 2026-09-19 | `research-authoring-and-testing.md` | +| | Whether a `tool.call` hook sees tool calls made inside a subagent. | 2026-09-19 | `research-api-surface.md` | +| | `$.process.run` and `$.http.fetch` as non-TypeScript escape hatches. | 2026-09-19 | `research-api-surface.md` | +| | The `classic.*` bridge events and the tier list. | 2026-09-19 | `research-what-mods-are.md` | +| | Fourth position: the `.catch` proposal the Sep 9 cheat sheet then ships. | 2026-09-19 | ADR 0035, `research-security-and-semantics.md` | +| | The intended scope of `prompt.submit`. | 2026-09-19 | `research-api-surface.md` | +| | What `/plugin-types` lists once the flag is enabled. | 2026-09-19 | `research-authoring-and-testing.md` | +| | Concurrency of `tool.check` guards versus `tool.call` guards. | 2026-09-19 | `research-security-and-semantics.md` | +| | Advice-style sugar, and which event may yield. | 2026-09-19 | `research-api-surface.md` | +| | 2026-09-16 restatement of the fail-closed spelling with `on(...).catch(...)`. | 2026-09-19 | `research-security-and-semantics.md` | +| | AGENTS.md as a built-in mod. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | 2026-09-19: whether mods can set their own registration order, and the tier names. | 2026-09-19 | `research-what-mods-are.md` | +| | The comment bcherny's community-update link resolves to. Authored by `sezaakgun`, `author_association: NONE`: a community Tetris demo, not an Anthropic update. Tiering by "linked from a staff post" misfiles it. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | How all 203 comments were fetched and grepped (for `desktop`, for install commands). | 2026-09-19 | `research-surfaces.md`, `research-enablement-and-distribution.md` | ## Official docs and changelog, checked for absence Every row here was checked for the absence of any mods or function-hooks mention. The surface rows -also carry the positive facts the surface matrix rests on. Criterion 2 of the go check re-runs the +also point at the pages the surface matrix rests on. Criterion 2 of the go check re-runs the corpus sweeps; they are first-mention detectors, never availability checks. The `docs/en/*` pages were read as page bodies inside `llms-full.txt` by one lane and individually fetched by another, so per-page fetch provenance differs between lanes. -| Link | What it establishes | Fetched | Used by | +| Link | What we used it for | As of | Used by | |---|---|---|---| -| | All 197 English documentation pages, 9,590,632 bytes: **0 hits** for `function hook`, `hooks module`, `plugin-types`, `prependPlugins`, `appendPlugins`, `engine.create`, `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS`, `sec-default`, `next.to(`, `"modules"`. Live controls (`plugin-dir` 67) prove the sweep works. | 2026-09-19 | ADR 0035 (criterion 2), go-no-go.md, `plugin-philosophy.md` | +| | Our sweep of all 197 English documentation pages, 9,590,632 bytes: **0 hits** for `function hook`, `hooks module`, `plugin-types`, `prependPlugins`, `appendPlugins`, `engine.create`, `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS`, `sec-default`, `next.to(`, `"modules"`. Live controls (`plugin-dir` 67) prove the sweep works. | 2026-09-19 | ADR 0035 (criterion 2), go-no-go.md, `plugin-philosophy.md` | | | The curated index. Recorded so a later agent does not grep it by mistake: always grep `llms-full.txt`. | 2026-09-19 | go-no-go.md | | | How the 197-page count was established. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | 7,158 lines, `## 2.1.278` down to `## 0.2.21`: **0 hits** for the same terms, and word-bounded `mods?` → 0. | 2026-09-19 | ADR 0035, `plugin-philosophy.md` | +| | Our sweep of 7,158 lines, `## 2.1.278` down to `## 0.2.21`: **0 hits** for the same terms, and word-bounded `mods?` → 0. | 2026-09-19 | ADR 0035, `plugin-philosophy.md` | | | Raw form, for the greppable criterion check. | 2026-09-19 | go-no-go.md | -| | The same pinned at the research commit; the 2.1.277 AGENTS.md entry that never says "mod". | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | The plugin manifest reference: no `modules` key documented; the `npm ci`-at-plugin-root rule. | 2026-09-19 | `research-repo-fit.md` | -| | The plugin system as documented, with no mods surface. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | The same pinned at the research commit; the 2.1.277 AGENTS.md entry, which our grep found does not say "mod". | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | The plugin manifest reference: our sweep found no `modules` key; the plugin-root dependency install rule. | 2026-09-19 | `research-repo-fit.md` | +| | The plugin system as documented; our sweep found no mods surface. | 2026-09-19 | `research-enablement-and-distribution.md` | | | Classic hooks as documented; the contract a `classic.*` bridge event mirrors. | 2026-09-19 | `research-security-and-semantics.md` | | | Distribution as documented; the one bare `plugin test` string in the corpus is the English word "testing" on this page. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS` is not listed. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | Managed settings as documented; `prependPlugins` / `appendPlugins` are absent. | 2026-09-19 | `research-security-and-semantics.md` | -| | Which surfaces fetch server-managed settings, and that Cowork never does. | 2026-09-19 | `research-surfaces.md` | -| | Desktop Code tab runs "the same underlying engine with a graphical interface" and shares `~/.claude` config; on Windows the app inherits user and system environment variables, the route the Desktop probe used. | 2026-09-19 | `research-surfaces.md`, experiments.md | -| | The VS Code extension bundles a private CLI copy. | 2026-09-19 | `research-surfaces.md` | -| | JetBrains runs `claude` in the integrated terminal, so it is a `terminal` surface. | 2026-09-19 | `research-surfaces.md` | -| | Web and mobile Code clients onto cloud sessions; `/plugin` unavailable. | 2026-09-19 | `research-surfaces.md` | +| | Our check that `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS` is not listed. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | Managed settings as documented; our check that `prependPlugins` / `appendPlugins` are absent. | 2026-09-19 | `research-security-and-semantics.md` | +| | Which surfaces fetch server-managed settings (surface matrix input). | 2026-09-19 | `research-surfaces.md` | +| | How the Desktop Code tab relates to the CLI engine and `~/.claude` config, and the Windows environment inheritance the Desktop probe used. | 2026-09-19 | `research-surfaces.md`, experiments.md | +| | How the VS Code extension runs the CLI (surface matrix input). | 2026-09-19 | `research-surfaces.md` | +| | How JetBrains runs `claude`, which makes it a `terminal` surface. | 2026-09-19 | `research-surfaces.md` | +| | Web and mobile Code clients on cloud sessions, and `/plugin` there. | 2026-09-19 | `research-surfaces.md` | | | Cloud sessions and their plugin delivery path. | 2026-09-19 | `research-surfaces.md` | -| | SDK plugin loading is `{ type: "local", path }` only. | 2026-09-19 | `research-surfaces.md` | -| | "Skills in Cowork and cloud sessions": which plugin components reach the account-side surfaces. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | How the SDK loads plugins. | 2026-09-19 | `research-surfaces.md` | +| | Which plugin components reach Cowork and cloud sessions. | 2026-09-19 | `research-enablement-and-distribution.md` | | | Swept in the 197-page term sweep; no mods mention. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | The early-access and feature-flag error entries, located and quoted. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | `claude plugin eval`'s own honesty about sandboxing, quoted as the comparison a mods warning lacks. | 2026-09-19 | `research-security-and-semantics.md` | -| | The canonical announcement URL; 308-redirects to the claude.com blog post below. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | The plugins launch announcement, published 2025-10-09: four plugin component types, no mods. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | The 2026-09-16 Cowork/Chat merge announcement; it never names Claude Code. | 2026-09-19 | `research-surfaces.md` | -| | Plugins in Cowork, synced through claude.ai rather than `~/.claude`. | 2026-09-19 | `research-surfaces.md` | +| | The early-access and feature-flag error entries. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | How `claude plugin eval` discloses its sandboxing, the comparison a mods warning lacks. | 2026-09-19 | `research-security-and-semantics.md` | +| | Not a pointer (correlate only): the canonical announcement URL, which we observed 308-redirecting to the blog post below. For the topic, see [Plugins](https://code.claude.com/docs/en/plugins). | 2026-09-19 | `research-enablement-and-distribution.md` | +| | Not a pointer (correlate only): the plugins launch announcement, published 2025-10-09; our check found no mods mention. For the topic, see [Plugins](https://code.claude.com/docs/en/plugins). | 2026-09-19 | `research-enablement-and-distribution.md` | +| | Not a pointer (correlate only): the 2026-09-16 Cowork/Chat merge announcement; our check found it does not name Claude Code. No docs page covered the merge as of 2026-09-19. | 2026-09-19 | `research-surfaces.md` | +| | Not a pointer (correlate only): plugins in Cowork and how they sync. For the topic, see [Use plugins in Claude](https://support.claude.com/en/articles/13837440-use-plugins-in-claude) below. | 2026-09-19 | `research-surfaces.md` | | | Cowork's own changelog: no mods mention. | 2026-09-19 | `research-surfaces.md` | | | The MCP bundle manifest specification, read as the adjacent packaging format. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | A synced plugin's "skills, agents, hooks, MCP servers, and LSP servers all load" in Cowork. | 2026-09-19 | `research-surfaces.md` | -| | Cowork runs its sessions on Claude Code. | 2026-09-19 | `research-surfaces.md` | +| | Which components of a synced plugin load in Cowork. | 2026-09-19 | `research-surfaces.md` | +| | What Cowork runs its sessions on. | 2026-09-19 | `research-surfaces.md` | | | Consumer-facing release notes: no mods mention. | 2026-09-19 | `research-surfaces.md` | | | Named by the `.mcpb` page's cross-reference. Reached only as a search snippet, never fetched. | not recorded | `research-surfaces.md` | ## GitHub issues and pull requests -| Link | What it establishes | Fetched | Used by | +| Link | What we used it for | As of | Used by | |---|---|---|---| -| | The roadmap thread ("Mods - make Claude 10x more extensible"), opened 2026-09-03. OPEN, 203 comments. The body is edited in place and carries the self-dated "Sep 9, 2026" Community Update and the public `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1` acknowledgement. | 2026-09-19 | ADR 0035, go-no-go.md, `research-roadmap-and-staff-statements.md` | -| | **The blocking defect.** Registering any `tool.call` hook on Bash breaks `Agent(isolation: "worktree")`; a passthrough `next(e)` is enough. OPEN, filed 2026-09-06 on 2.1.263 (macOS), one independent Windows reproduction at 2.1.272 in the thread, and a second Windows reproduction run locally on 2026-09-19 at 2.1.278 (`experiments.md` E2); no staff reply. | 2026-09-19 | ADR 0035 (criterion 3), go-no-go.md, experiments.md, `plugin-philosophy.md`, `research-security-and-semantics.md` | +| | The mods roadmap thread, opened 2026-09-03. OPEN, 203 comments. The body is edited in place and carries the self-dated "Sep 9, 2026" Community Update and the public `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1` acknowledgement. | 2026-09-19 | ADR 0035, go-no-go.md, `research-roadmap-and-staff-statements.md` | +| | **The blocking defect.** Registering any `tool.call` hook on Bash breaks `Agent(isolation: "worktree")`; a passthrough `next(e)` is enough. OPEN, filed 2026-09-06 on 2.1.263 (macOS), one independent Windows reproduction at 2.1.272 in the thread, and our second Windows reproduction run locally on 2026-09-19 at 2.1.278 (`experiments.md` E2); no staff reply. | 2026-09-19 | ADR 0035 (criterion 3), go-no-go.md, experiments.md, `plugin-philosophy.md`, `research-security-and-semantics.md` | | | Plugin-native `PreToolUse` hooks auto-discovered via `hooks/hooks.json` reported not enforced in interactive sessions. OPEN, 0 comments. Gates the story that `hooks` and `modules` in one manifest both fire. | 2026-09-19 | go-no-go.md, `research-security-and-semantics.md` | | | The pull request that added `mods/`, opened 2026-09-09T22:31:07Z and self-merged at 22:34:33Z: the evidence behind the disclosed staff-by-inference call on `poteat`. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | Generated-type incompleteness. OPEN. | 2026-09-19 | go-no-go.md, `research-api-surface.md` | | | Generated-type drift. OPEN. | 2026-09-19 | go-no-go.md, `research-api-surface.md` | | | A `session.compact` hook's compaction undone on resume. OPEN. | 2026-09-19 | go-no-go.md, `research-security-and-semantics.md` | -| | The more ambitious runtime goal a staff comment links to when discussing raised node limits and WebAssembly. | not recorded (quoted inside a staff comment) | `research-api-surface.md` | +| | The runtime goal a staff comment links to when discussing node limits and WebAssembly. | not recorded (linked from a staff comment) | `research-api-surface.md` | | | The MCP bundle manifest, compared against the plugin manifest as a distribution alternative. | 2026-09-19 | `research-enablement-and-distribution.md` | ## Community @@ -173,24 +176,24 @@ per-page fetch provenance differs between lanes. Third-party evidence. Nothing here outranks a first-party source; where a community claim and a first-party source disagreed, the first-party source won. -| Link | What it establishes | Fetched | Used by | +| Link | What we used it for | As of | Used by | |---|---|---|---| | | **Seed.** A reply in the bcherny post's chain, by the author of the `Mindful-Claude` mod. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | **Seed.** A note tweet quoting bcherny: "Mods are still in early access, and their APIs may change." Carries the circulated prompt template telling Claude to read documentation that does not exist. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | **Seed.** A third-party adopter's account of modifying Claude Code's internal functions, and the promise to distribute mods through aitmpl.com. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | The no-auth X-to-Markdown converter every X quote was fetched and re-fetched through; it resolved 9 of 9 posts. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | **Seed.** A note tweet relaying bcherny on early access and API stability, carrying a circulated prompt template that points Claude at documentation that does not exist. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | **Seed.** A third-party adopter's account of modifying Claude Code's internal functions, and the plan to distribute mods through aitmpl.com. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | +| | The no-auth X-to-Markdown converter every X post was fetched and re-fetched through; it resolved 9 of 9 posts. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | Attempted unroll of the bcherny seed. **Failed**: HTTP 200 landing page, no post content. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | Attempted unroll of the dani_avila7 seed. **Failed**, same shape. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | Attempted unroll of the trq212 seed. **Failed**, same shape. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | -| | The substantial independent migration report (2026-09-17, Windows, 2.1.273/2.1.274): "Claude Code never calls `register()`", which is the silent-inert failure mode, and the version-canary and denial-test lessons. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | A third-party write-up; names 2.1.273 as the version its inspected declarations identify, which is not a floor. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | The substantial independent migration report (2026-09-17, Windows, 2.1.273/2.1.274): the silent-inert failure mode where the module is never registered, and the version-canary and denial-test lessons. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | A third-party write-up; the version it names is the one its inspected declarations identify, which is not a floor. | 2026-09-19 | `research-enablement-and-distribution.md` | | | A third-party write-up stating no version floor. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | A mods component library; claims a `>= 2.1.259` floor, uncorroborated. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | A mods component library; its version floor claim is uncorroborated. | 2026-09-19 | `research-enablement-and-distribution.md` | | | The same library's repository, with an `npx` installer: evidence that distribution is happening outside any marketplace. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | A third-party adopter shipping a hooks module with tests; its `hooks.json` documents the deliberate "activates only when the variable is exactly 1" defense, as does the same author's `compact-adviser`. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | A third-party mod; the only source naming the install path (`claude plugin marketplace add` / `claude plugin install`) and a "2.1.269 or later" floor. Neither is corroborated by staff. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | A third-party adopter shipping a hooks module with tests; its `hooks.json` documents an exact-value activation guard, as does the same author's `compact-adviser`. | 2026-09-19 | `research-enablement-and-distribution.md` | +| | A third-party mod; the only source naming an install path and a version floor. Neither is corroborated by staff. | 2026-09-19 | `research-enablement-and-distribution.md` | | | A third-party mod repository, counted in the adopter census. | 2026-09-19 | `research-enablement-and-distribution.md` | -| | The Tetris-in-Claude mod bcherny's post points at. | not recorded (quoted from comment 5666255143) | `research-roadmap-and-staff-statements.md` | +| | The Tetris-in-Claude mod bcherny's post points at. | not recorded (linked from comment 5666255143) | `research-roadmap-and-staff-statements.md` | | | A third-party reverse-engineered changelog, dating `claude plugin test` as absent in 2.1.270 and present in 2.1.271. | 2026-09-19 | `research-authoring-and-testing.md` | | | Hacker News attention: the mods story is 3 points, 0 comments. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | | | The same sweep on the engineering term. | 2026-09-19 | `research-roadmap-and-staff-statements.md` | @@ -199,8 +202,8 @@ first-party source disagreed, the first-party source won. | | Secondary coverage of the same merge; does not carry the two-mode claim. | 2026-09-19 | `research-surfaces.md` | | | The one secondary outlet of three that carries the two-mode claim, which no first-party source supports. | 2026-09-19 | `research-surfaces.md` | | | Indicative of the 2026-09-17 Projects beta date. Reached as a search snippet, never fetched, so not an independent corroborator. | not recorded | `research-surfaces.md` | -| | Node's own statement that `node:vm` "is not a security mechanism", the support for this corpus's `INFERRED` non-containment reading. | 2026-09-19 | `research-security-and-semantics.md` | -| | The essay a staff comment cites when refusing a numeric priority system. | not recorded (quoted inside a staff comment) | `research-repo-fit.md` | +| | Node's own position on `node:vm` as a security mechanism, the support for this corpus's `INFERRED` non-containment reading. | 2026-09-19 | `research-security-and-semantics.md` | +| | The essay a staff comment cites when refusing a numeric priority system. | not recorded (linked from a staff comment) | `research-repo-fit.md` | ## Not indexed, and why diff --git a/docs/upstream/claudedevs-cost-performance.md b/docs/upstream/claudedevs-cost-performance.md index a5f3f72c3b..1b94603ae2 100644 --- a/docs/upstream/claudedevs-cost-performance.md +++ b/docs/upstream/claudedevs-cost-performance.md @@ -35,42 +35,59 @@ cache-diagnostics UI; a second real need for API-cost tooling in this marketplac ## Source and verification -- Post: `https://x.com/ClaudeDevs/status/2097369738968195513` (2026-09-08). Canonical mirror: - `https://claude.com/blog/reducing-cost-and-improving-performance-with-claude-platform`. +- Main docs page covering the article's topic: + [Optimizing for cost and intelligence](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence). + The article itself is never a pointer: (correlate with `https://x.com/ClaudeDevs/status/2097369738968195513`, 2026-09-08) + and its mirror (correlate with `https://claude.com/blog/reducing-cost-and-improving-performance-with-claude-platform`). Content parity between the two confirmed 2026-09-09; the blog carries the seven figures. - Reply-chain coverage is a bounded gap: no converter reachable from the discovery session could enumerate X replies (Thread Reader has no unroll of this post), and a web search surfaced no official follow-up posts. Checked: threadreaderapp.com, web search. Unchecked: the live X reply timeline (needs X auth). -- Claim verification ran 2026-09-09 against live primaries (a 24-row bounded ledger, coverage +- Our verification ran 2026-09-09 against live primaries (a 24-row bounded ledger, coverage gate exit 0): every prompt-caching, effort, mid-conversation-system-message, batching, and - Admin API mechanics claim in the article verified against the platform docs pages the - article itself links, and the skill-behavior claims verified source-as-spec from the + Admin API mechanism the article names was checked against the platform docs page each row + points at, and the skill-behavior rows were checked source-as-spec against the anthropics/skills clone (HEAD `41bbe19`, 2026-09-03) plus the skill bundled inside Claude - Code 2.1.263. Recheck trigger for every row below: divergence at re-fetch of the named - basis. -- Three verification findings qualify adoption everywhere below: - 1. **hillclimb repo lag.** `/claude-api hillclimb` (and `build-eval`) ship in the bundled - skill inside the Claude Code binary but are absent from the public anthropics/skills repo - the article links (HEAD 2026-09-03) and from the skill's platform-docs page. A reader - following the article's GitHub link will not find them. Recheck: the repo or docs page - gains the subcommands. - 2. **Claude Console diagnostics UI unverified.** The API half of the cache-diagnostics claim - is fully verified (four `*_changed` miss reasons); no fetched doc describes the Console - request-comparison UI in the article's Figure 1. Checked: cache-diagnostics doc, - usage-cost-api doc, two web searches. Unchecked: the Console product itself (needs a - login). - 3. **Benchmark numbers are vendor-internal.** Figures 3 to 7 (14.6 percent cost cut, +5.3 - points, FrontierCode Diamond, CursorBench 3.2, the four cost-optimize benchmarks) are - single-pool Anthropic measurements with no published artifact to reproduce from. Posture - recommended by the research run: adopt mechanisms, cite numbers only as vendor-reported. - 4. **Beta boundaries.** Per-message effort changes (Fable 5.1, Mythos 5.1, Opus 5), the - cache diagnostics API, and turn-scoped system messages are betas with headers; plain - mid-conversation system messages are GA on six models (not Sonnet 5). Adopted guidance - carries the qualifiers. - 5. **Corroboration verdict (fresh-context verifier, 2026-09-09).** Every accepted claim is - HIGH confidence and five live spot checks matched current sources verbatim; four claims - rest on a single evidence pool because no independent second pool exists publicly: + Code 2.1.263. +- Five verification findings qualify adoption everywhere below: + 1. **hillclimb repo lag.** Our extraction found `/claude-api hillclimb` (and `build-eval`) in + the bundled skill inside the Claude Code binary, and an exhaustive grep found them absent + from the public anthropics/skills repo the article links (HEAD 2026-09-03) and from the + skill's platform-docs page. A reader following the article's GitHub link will not find + them. Pointer: our binary extraction and clone grep, recorded in the claude-api row of + [`docs/native-surfaces/records.json`](../native-surfaces/records.json), and [In Claude Code (bundled)](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/claude-api-skill#in-claude-code-bundled). + As of: 2026-09-09. Recheck trigger: the repo or docs page gains the subcommands. + 2. **Claude Console diagnostics UI unverified.** The API half of the cache-diagnostics topic + is verified against + [Cache miss reason types](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics#cache-miss-reason-types); + no fetched doc covers the Console request-comparison UI the article shows. Checked: + cache-diagnostics doc, usage-cost-api doc, two web searches. Unchecked: the Console + product itself (needs a login). As of: 2026-09-09. Recheck trigger: a docs page starts + covering the Console request-comparison UI, or someone checks the Console itself. + 3. **Benchmark numbers are vendor-internal.** The article's benchmark figures are single-pool + Anthropic measurements with no published artifact to reproduce from, so no figure is + recorded here. Posture recommended by the research run: adopt mechanisms, cite numbers + only as vendor-reported, read at the source. + 4. **Release-status boundaries.** Every adopted line touching per-message effort, the cache + diagnostics API, or mid-conversation system messages names the feature's release status and + the platforms and models it is limited to, as read at the pointer when the line is written, + never from memory and never assuming beta. Re-read 2026-10-01: the three features no longer + share one status, so a line that calls all three beta is stale and is corrected when next + touched. + - **Pointer**: for each feature's status and limits, see + [Per-message effort (beta)](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta), + [Cache diagnostics](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics) + and + [Mid-conversation system messages](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) + (on both pages the status and availability notes sit under the page title, in no section + of their own). + - **As of**: 2026-10-01 + - **Recheck trigger**: any of the three pages changes the feature's release status, its + supported platforms, or its supported models. + 5. **Corroboration verdict (fresh-context verifier, 2026-09-09).** Every accepted row is + HIGH confidence and five live spot checks matched current sources; four rows rest on a + single evidence pool because no independent second pool exists publicly: automatic-caching breakpoint movement (docs plus a restatement page; substance re-confirmed live), cost-optimize behavior (skill source only; the skill's docs page does not document the command), hillclimb-bundled and hillclimb-absent (binary extraction and @@ -79,15 +96,17 @@ cache-diagnostics UI; a second real need for API-cost tooling in this marketplac ## Row schema -Each row is a four-part record per -[docs/conventions/upstream-drift/README.md](../conventions/upstream-drift/README.md): the -claim, the basis it was derived against, the as-of date, and a recheck trigger. Columns: +Each row is an [upstream-drift](../conventions/upstream-drift/README.md#required-parts) record. +**Topic** names the article's practice in our words; **Ours** and **Verdict** are our decision +and its reasoning; **Pointer** is the exact docs section, or our own probe where no docs page +covers the topic; **As of** is when the pointer was last read. No row restates the article or +the page. Columns: -| Article claim / practice | Ours | Verdict | Reasoning, basis, as-of | +| Topic | Ours | Verdict | Pointer | As of | -Shared basis shorthand used below: "article" is the source snapshot above; "verified" means -the 2026-09-09 research run confirmed the claim against the named primary; explore evidence -paths were independently re-verified against the repo the same day. +Shared recheck trigger for every row: a re-fetch of the row's pointer no longer supports the +verdict. A TRACK verdict names its own trigger in its cell. "Explore evidence" paths were +independently re-verified against the repo on 2026-09-09. ## Lane T1: prompt-cache management @@ -101,19 +120,19 @@ authoring guidance has no incumbent. Decided at interview, 2026-09-10: the adopted API-side practices land as **one new prompt-caching reference chapter in the playbooks plugin**, beside the model-adaptation -chapters: four-part pointer rows citing each docs anchor, beta qualifiers carried, -session-side coverage cross-referenced. One work item covers the chapter. - -| Article claim / practice | Ours | Verdict | Reasoning, basis, as-of | -|---|---|---|---| -| Monitor cache hit rate; diagnose misses via the cache diagnostics API (miss reasons: messages / system / tools / model changed) | `harness-ops:observability` covers Claude Code sessions only | ADOPT API half (chapter row, beta-qualified) + TRACK Console half; also ADOPT one boundary-pointer line in the observability skill's cache-health context (decided 2026-09-10) | Verified against `platform.claude.com/docs/en/build-with-claude/cache-diagnostics`, 2026-09-09. Console UI unverified (finding 2); TRACK trigger: a Console-access check or a docs page confirming the request-comparison UI | -| Keep volatile values (timestamps, IDs) out of the prefix; stable-first request layout; tool definitions render first and any change breaks cache | `extract-ssot` anti-patterns record states the byte-identical-prefix rule (verified 2026-08-04); no authoring-rule surface for request-building code | ADOPT (chapter rows; decided 2026-09-10) | Verified against `prompt-caching#structuring-your-prompt`, 2026-09-09 | -| defer_loading rarely used tools; tool search appends them without breaking cache | `context-budget` levers.json engages defer_loading for Claude Code MCP tools only | ADOPT (chapter row; decided 2026-09-10) | Verified against `tool-use-with-prompt-caching#defer-loading-and-cache-preservation`, 2026-09-09 | -| Apply system-prompt updates as mid-conversation messages (cache-preserving; certain models) | No coverage | ADOPT (chapter row; GA six-model list, not Sonnet 5; decided 2026-09-10) | Verified against `mid-conversation-system-messages`, 2026-09-09 | -| Batch model/effort changes into already-broken-cache moments (compaction) | PLUGIN-PHILOSOPHY cache caveat carries the session-side version | COVERED session-side; API-side sentence joins the chapter (decided 2026-09-10) | Article; corroborated by Cognition devin-fusion post, 2026-06-29 | -| Move breakpoints as conversation grows; automatic caching pins the last cacheable block | No coverage | ADOPT (chapter row; decided 2026-09-10) | Verified against `prompt-caching#automatic-caching`, 2026-09-09 | -| Pre-warm with `max_tokens: 0` plus explicit breakpoint at session start | No coverage | ADOPT (chapter row; decided 2026-09-10) | Verified against prompt-caching doc and skill source, 2026-09-09; the rejection list (streaming, extended thinking, structured outputs, forced tool_choice, batches) rides along | -| 5-minute TTL counts from request start; long tool calls expire the parent cache; use 1-hour TTL (2x write rate) | fable-5 `orchestration.md:97` carries the Claude Code subagent version (re-verified 2026-09-06) | COVERED session-side; API-side rows join the chapter (decided 2026-09-10) | Verified against `prompt-caching#ttl-support` and pricing page, 2026-09-09 | +chapters: pointer records citing each docs anchor, beta qualifiers carried, session-side +coverage cross-referenced. One work item covers the chapter. + +| Topic | Ours | Verdict | Pointer | As of | +|---|---|---|---|---| +| Cache hit-rate monitoring and miss diagnosis | `harness-ops:observability` covers Claude Code sessions only | ADOPT API half (chapter row, release status read at the pointer) + TRACK Console half; also ADOPT one boundary-pointer line in the observability skill's cache-health context (decided 2026-09-10). Console UI unverified (finding 2); TRACK trigger: a Console-access check or a docs page confirming the request-comparison UI | [Cache miss reason types](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics#cache-miss-reason-types) | 2026-09-09 | +| Prefix stability and request layout | `extract-ssot` anti-patterns record carries the byte-identical-prefix rule (as of 2026-08-04); no authoring-rule surface for request-building code | ADOPT (chapter rows; decided 2026-09-10) | [Structuring your prompt](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#structuring-your-prompt) | 2026-09-09 | +| Deferred loading of rarely used tools | `context-budget` levers.json engages defer_loading for Claude Code MCP tools only | ADOPT (chapter row; decided 2026-09-10) | [defer_loading and cache preservation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-use-with-prompt-caching#defer-loading-and-cache-preservation) | 2026-09-09 | +| System-prompt updates as mid-conversation messages | No coverage | ADOPT (chapter row; GA and model-list boundary carried; decided 2026-09-10) | [When to use a mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#when-to-use-a-mid-conversation-system-message) | 2026-09-09 | +| Timing model and effort changes to cache breaks (compaction) | PLUGIN-PHILOSOPHY cache caveat carries the session-side version | COVERED session-side; API-side sentence joins the chapter (decided 2026-09-10). No docs page covers this timing practice as of 2026-09-09 (correlate with Cognition's devin-fusion post, 2026-06-29); recheck trigger: a docs page starts covering it | [What invalidates the cache](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#what-invalidates-the-cache) | 2026-09-09 | +| Breakpoint placement as a conversation grows | No coverage | ADOPT (chapter row; decided 2026-09-10) | [Automatic caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#automatic-caching) | 2026-09-09 | +| Cache pre-warming at session start | No coverage | ADOPT (chapter row; decided 2026-09-10); the chapter points at the request shapes pre-warming rejects rather than listing them | [Pre-warming the cache](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#pre-warming-the-cache) and the bundled skill source | 2026-09-09 | +| Cache TTL across long tool calls | fable-5 `orchestration.md:97` carries the Claude Code subagent version (as of 2026-09-06) | COVERED session-side; API-side rows join the chapter (decided 2026-09-10) | [1-hour cache duration](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#1-hour-cache-duration) and [Prompt caching pricing](https://platform.claude.com/docs/en/about-claude/pricing#prompt-caching) | 2026-09-09 | ## Lane T2: prompt instruction anti-patterns @@ -122,9 +141,9 @@ run (Claude Code 2.1.258 against Fable 5.1) covered 74 plugins, 241 skills, 798 115k lines; 805 findings applied, 207 withheld, 694 files changed (ADR-0028), which makes it a repeating lane per model change. The `harness-config:audit-instructions` criteria catalog maps one-to-one onto -the article's six anti-pattern families: +the six anti-pattern families the article names: -| Article anti-pattern | Catalog row(s) in `audit-instructions/reference/criteria.md` | +| Anti-pattern family | Catalog row(s) in `audit-instructions/reference/criteria.md` | |---|---| | Verification rituals | I8-a, I8-b | | Emphasis boosters | I28-a, I6 | @@ -139,30 +158,32 @@ Native-first gate row, decided at interview 2026-09-10: the (bundled claude-api, subcommand where it fits the use case, run our own processes where they fit; no on-paper routing restriction. The verdict is baked where the model reads it: a `## Boundary, the bundled claude-api skill` section in the `audit-instructions` body (routing, mutation gate, -availability rule) with the four-part records in +availability rule) with its records in `plugins/harness-config/skills/audit-instructions/reference/bundled-claude-api.md`. Recheck fires with the store row's trigger (subcommand set changes, or the public repo / docs page gains hillclimb). -| Article claim / practice | Ours | Verdict | Reasoning, basis, as-of | -|---|---|---|---| -| Six anti-pattern families hobble frontier models; audit and remove them | `audit-instructions` catalog + executed fleet-wide prompt-audit | COVERED (Claude Code surfaces; decided 2026-09-10) | Explore mapping re-verified against criteria.md TOC and rows, 2026-09-09; anti-patterns verified documented in `optimizing-for-cost-and-intelligence` and skill `prompt-audit.md`, 2026-09-09; ADR-0028's repeat-per-model-change lane is the standing remedy | -| prompt-audit also covers prompts in application code calling the Claude API | No repo skill audits app-code prompts; `audit-instructions` scope is deliberately Claude Code surfaces | REJECT scope-widening (decided 2026-09-10) | The bundled prompt-audit owns the app-code surface; the skill's scope boundary is deliberate. prompt-audit.md Step 0-1 inventory includes request-building code, verified source-as-spec 2026-09-09. Revisit only if marketplace-native app-code coverage is wanted later | -| Manual thinking budgets rejected outright by the API on newer models | I17 family covers the instruction-surface version | COVERED (decided 2026-09-10) | Verified (400 rejection documented) 2026-09-09 | +| Topic | Ours | Verdict | Pointer | As of | +|---|---|---|---|---| +| Instruction anti-patterns on current models | `audit-instructions` catalog + executed fleet-wide prompt-audit | COVERED (Claude Code surfaces; decided 2026-09-10). Explore mapping re-verified against criteria.md TOC and rows; ADR-0028's repeat-per-model-change lane is the standing remedy | [Audit prompts against the current model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#audit-prompts-against-the-current-model) and the skill's `prompt-audit.md` | 2026-09-09 | +| Prompt audit of application-code prompts | No repo skill audits app-code prompts; `audit-instructions` scope is deliberately Claude Code surfaces | REJECT scope-widening (decided 2026-09-10). The bundled prompt-audit owns the app-code surface; the skill's scope boundary is deliberate. Our source-as-spec read found request-building code in its Step 0-1 inventory. Revisit only if marketplace-native app-code coverage is wanted later | [Claude API skill](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/claude-api-skill#what-the-skill-provides) and the skill's `prompt-audit.md` | 2026-09-09 | +| Manual thinking budgets on newer models | I17 family covers the instruction-surface version | COVERED (decided 2026-09-10) | [Configuring thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#configuring-thinking) | 2026-09-09 | ## Lane T3: effort calibration Mostly covered: PLUGIN-PHILOSOPHY "Effort tiers" lane rules, catalog rows I21/I22/I27, model-adaptation chapters (opus-5 "move down liberally", fable-5-1 "recall at low effort"), -`${CLAUDE_EFFORT}` consumed by 7 skills, 13 agents pinned `effort: high`. Missing: cross-model -economics and sweep tooling. - -| Article claim / practice | Ours | Verdict | Reasoning, basis, as-of | -|---|---|---|---| -| Effort miscalibration cuts both ways (over-thinking degrades quality; under-thinking answers from partial evidence) | PLUGIN-PHILOSOPHY Effort tiers; opus-5 chapter overthinking guidance; fable-5-1 low-effort recall caveat | COVERED, plus a sharpening ADOPT (decided 2026-09-10) | Explore evidence re-verified 2026-09-09; effort semantics verified against the effort doc. Work item: fold the article's sharpest phrasings (deliberation only helps while there is evidence to find; the answer looks finished but is built on partial information) into the existing surfaces | -| Test stronger models at lower effort (Fable 5.1 low matches Fable 5 high at a third of the cost; cache reads $0.25/M vs $1.00/M) | Nowhere; adaptation chapters deliberately carry no pricing | ADOPT (decided 2026-09-10) | Land as a pricing-free section in the fable-5-1 model-adaptation chapter plus a one-line pointer in PLUGIN-PHILOSOPHY Effort tiers; numbers cited vendor-reported; pricing stays pointer-resolved through the claude-api skill. Pricing verified against the pricing page and API release notes 2026-09-01 entry, fetched 2026-09-09 | -| Sweep effort levels on a non-saturated eval; flat curve means not thinking-bound | `evals` plugin has zero effort content | ADOPT (decided 2026-09-10) | Land as an effort-axis note in the evals plugin citing the bundled hillclimb per the Lane M posture (bundled-only, public-repo lag noted). Verified against `optimizing-for-cost-and-intelligence#tune-effort`, 2026-09-09 | -| Only select models change effort mid-conversation without breaking cache | PLUGIN-PHILOSOPHY cache caveat + criteria I17-b carry the session-side version | COVERED session-side (decided 2026-09-10) | The API-side model list (Fable 5.1, Mythos 5.1, Opus 5, beta header; Fable 5 returns 400) lands only inside whatever T1/T3 adoptions get written, per the Lane M beta posture; no separate surface. Verified against `effort#change-effort-mid-conversation-beta`, 2026-09-09 | +`${CLAUDE_EFFORT}` consumed by 7 skills, and every named agent's effort pin, `high` or `medium` +(see the pinned-agents record under [Effort tiers](../plugin-philosophy.md#effort-tiers) and the +[Effort floor](../plugin-philosophy.md#effort-floor)). Missing: cross-model economics and sweep +tooling. + +| Topic | Ours | Verdict | Pointer | As of | +|---|---|---|---|---| +| Effort miscalibration in both directions | PLUGIN-PHILOSOPHY Effort tiers; opus-5 chapter overthinking guidance; fable-5-1 low-effort recall caveat | COVERED, plus a sharpening ADOPT (decided 2026-09-10). Explore evidence re-verified 2026-09-09. Work item: fold the article's two sharpest phrasings on miscalibration, in our words, into the existing surfaces | [How effort works](https://platform.claude.com/docs/en/build-with-claude/effort#how-effort-works) | 2026-09-09 | +| A stronger model at lower effort | Nowhere; adaptation chapters deliberately carry no pricing | ADOPT (decided 2026-09-10). Land as a pricing-free section in the fable-5-1 model-adaptation chapter plus a one-line pointer in PLUGIN-PHILOSOPHY Effort tiers; numbers cited vendor-reported; pricing stays pointer-resolved through the claude-api skill | [Compare models on cost per task](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) and [Model pricing](https://platform.claude.com/docs/en/about-claude/pricing#model-pricing) | 2026-09-09 | +| Effort sweeps on a non-saturated eval | `evals` plugin has zero effort content | ADOPT (decided 2026-09-10). Land as an effort-axis note in the evals plugin citing the bundled hillclimb per the Lane M posture (bundled-only, public-repo lag noted) | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) | 2026-09-09 | +| Effort changes mid-conversation and the cache | PLUGIN-PHILOSOPHY cache caveat + criteria I17-b carry the session-side version | COVERED session-side (decided 2026-09-10). The API-side model list is read at the pointer only inside whatever T1/T3 adoptions get written, per the Lane M beta posture; no separate surface | [Per-message effort (beta)](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) | 2026-09-09 | ## Lane T4: API cost optimization and profiling @@ -171,13 +192,13 @@ measurement (`context-budget`), and the adjacent `rate-limit-guard` (subscriptio cost). Batch API, output bounding as a cost lever, and the usage/cost Admin API are uncovered. `cost-optimize` and `hillclimb` have zero in-repo references. -| Article claim / practice | Ours | Verdict | Reasoning, basis, as-of | -|---|---|---|---| -| cost-optimize profiles spend (Admin API, else logged `usage` objects, else code estimate), ranks levers, measures against an eval | No incumbent for API-application profiling | TRACK on the bundled cost-optimize, plus one mention in the new playbooks chapter as the automation for its levers (decided 2026-09-10) | Verified source-as-spec (`shared/cost-optimization.md`), 2026-09-09; nuance: it proposes rather than silently applies. New-plugin question deferred to a second real need | -| hillclimb searches cost/performance over models and effort with train/test split | No incumbent; `evals` owns eval design without a cost axis | Cited per the Lane M posture: bundled-only, public-repo lag noted (decided 2026-09-10); the evals effort-axis note carries the citation | Verified from the bundled skill source extracted from the binary, 2026-09-09; absent from public repo HEAD. Recheck: the repo or docs page gains the subcommand | -| Batch unattended work (50 percent discount, stacks with cache multipliers) | Absent (sole mention is a routines.md disclaimer) | ADOPT (chapter row; decided 2026-09-10) | Verified against the pricing page batch section, 2026-09-09 | -| Bound output to save cost (SWE-bench case: concise-output constraint) | In tension with prompt-audit Group 1f, which removes numeric output ceilings from skill bodies | Recorded scope-disjoint (decided 2026-09-10): output bounding is an API-request cost lever, never a skill-body instruction pattern; one sentence in the chapter says so | Article; tension identified by explore, 2026-09-09 | -| Usage and Cost Admin API for org spend profiling | Absent | ADOPT (chapter row; decided 2026-09-10) | Verified against `manage-claude/usage-cost-api`, 2026-09-09 | +| Topic | Ours | Verdict | Pointer | As of | +|---|---|---|---|---| +| Spend profiling with the bundled cost-optimize | No incumbent for API-application profiling | TRACK on the bundled cost-optimize, plus one mention in the new playbooks chapter as the automation for its levers (decided 2026-09-10). Our source-as-spec read found it proposes rather than silently applies. New-plugin question deferred to a second real need. Recheck: a docs page starts covering the command | Our read of the bundled skill source (`shared/cost-optimization.md`); no docs page covers the command | 2026-09-09 | +| Model and effort search with the bundled hillclimb | No incumbent; `evals` owns eval design without a cost axis | Cited per the Lane M posture: bundled-only, public-repo lag noted (decided 2026-09-10); the evals effort-axis note carries the citation. Recheck: the repo or docs page gains the subcommand | Our extraction of the bundled skill source from the binary (finding 1); [In Claude Code (bundled)](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/claude-api-skill#in-claude-code-bundled) | 2026-09-09 | +| Batching unattended work | Absent (sole mention is a routines.md disclaimer) | ADOPT (chapter row; decided 2026-09-10) | [Batch processing pricing](https://platform.claude.com/docs/en/about-claude/pricing#batch-processing) | 2026-09-09 | +| Output bounding as a cost lever | In tension with prompt-audit Group 1f, which removes numeric output ceilings from skill bodies | Recorded scope-disjoint (decided 2026-09-10): output bounding is an API-request cost lever, never a skill-body instruction pattern; one sentence in the chapter says so. Tension identified by explore, 2026-09-09 | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) | 2026-09-09 | +| Org spend profiling through the Admin API | Absent | ADOPT (chapter row; decided 2026-09-10) | [Usage and Cost API: Cost API](https://platform.claude.com/docs/en/manage-claude/usage-cost-api#cost-api) | 2026-09-09 | ## Lane M: record and gating meta-decisions @@ -185,10 +206,10 @@ Decided at interview, 2026-09-10: | Question | Verdict | Reasoning, basis, as-of | |---|---|---| -| Record shape | DECIDED: keep this file's shape. Only the row-schema FORMAT is borrowed from aihero-course.md (four-part rows, verdict vocabulary); this source is unrelated to AI Hero and this record stands alone | Owner interview, 2026-09-10 | +| Record shape | DECIDED: keep this file's shape. Only the row-schema FORMAT is borrowed from aihero-course.md (per-row records, verdict vocabulary); this source is unrelated to AI Hero and this record stands alone | Owner interview, 2026-09-10 | | Native-overlap gate before any new skill: the article's guidance IS the bundled claude-api skill | DECIDED: run `/harness-ops:audit-native-overlap` against the four topics first and record its verdicts as gate rows; adoption scope is NOT pre-restricted on paper. The owner receives full information per topic and decides at each lane interview. Amended 2026-09-11: a registry row alone is not the deliverable; each non-`defer` verdict lands as a `## Boundary` section in the skill body with detail in a same-skill reference file, and the native-references convention (1.1.0) now requires the pair | Owner interview, 2026-09-10 and 2026-09-11; PLUGIN-PHILOSOPHY Native-first section; ADR-0028 precedent | | Vendor-internal numbers and beta features | DECIDED: adopt mechanisms only; cite figures as vendor-reported and unreproduced; every adopted line touching a beta feature carries its beta qualifier and GA/model-list boundary | Owner interview, 2026-09-10 | -| Citing hillclimb while the public repo lags | DECIDED: cite it as a bundled Claude Code command with a four-part record noting the public-repo lag; recheck trigger fires when the anthropics/skills repo or the skill's docs page gains the subcommand | Owner interview, 2026-09-10 | +| Citing hillclimb while the public repo lags | DECIDED: cite it as a bundled Claude Code command with an upstream-drift record noting the public-repo lag; recheck trigger fires when the anthropics/skills repo or the skill's docs page gains the subcommand | Owner interview, 2026-09-10 | ## Interview queue diff --git a/docs/upstream/opus-5-5-usage-guide.md b/docs/upstream/opus-5-5-usage-guide.md index 0cd2643b3f..2666f7574e 100644 --- a/docs/upstream/opus-5-5-usage-guide.md +++ b/docs/upstream/opus-5-5-usage-guide.md @@ -29,86 +29,88 @@ their own session are filed as issues (see [Decisions and follow-ups](#decisions ## Source and verification -- Guide: `https://claude.dev/blog/getting-the-most-out-of-opus-5-5/`, by Addy Osmani, published - 2026-09-22, read 2026-09-23. -- Primary docs used alongside it, read 2026-09-23: the "Prompting Claude Opus 5.5" guide on the - Claude platform docs, and the Claude Code model-config page (effort defaults, the `opus` alias, - `MAX_THINKING_TOKENS`, and the "Automatic model fallback" section). -- Vendor-reported claims (for example, that Opus 5.5 at its lowest effort caught more bugs than - Opus 5 at high effort) are recorded as vendor-reported, not as verified facts. +- Pointer for every row: [Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5) + on the Claude platform docs, read 2026-09-23. The usage guide by Addy Osmani, published + 2026-09-22: (correlate with `https://claude.dev/blog/getting-the-most-out-of-opus-5-5/`). +- Pointer for model behavior: the Claude Code + [Model configuration](https://code.claude.com/docs/en/model-config) page, read 2026-09-23 (effort + defaults, the `opus` alias, `MAX_THINKING_TOKENS`, and + [Automatic model fallback](https://code.claude.com/docs/en/model-config#automatic-model-fallback)). +- Vendor-reported performance claims are not recorded here; read them at the source. ## Row schema -Each row is a four-part record per -[docs/conventions/upstream-drift/README.md](../conventions/upstream-drift/README.md): the claim, the -basis it was derived against, the as-of date, and a recheck trigger. Shared basis for every row -below: the guide as read 2026-09-23. Shared recheck trigger: a revised Opus 5.5 guide, or the next -Opus release. +Each row is an [upstream-drift](../conventions/upstream-drift/README.md#required-parts) record. The +**Topic** column names the guide item in our words, with an ID other files cite; **Ours** and +**Verdict** are our decision. No row restates the guide. Shared for every row: -The guide's "Try this first" items and closing checklist restate the sections below, so they carry -no rows of their own: handing over the whole task maps to G1.1 and G2.1, deleting think-carefully -lines to G1.2, and reading what it needs from you first to G3.1. Each checklist line maps to G1.1, -G1.2, G1.4, G4.1, G2.1 to G2.3, G3.1 to G3.3, or the fallback row. +- **Pointer**: the pages in [Source and verification](#source-and-verification). +- **As of**: 2026-09-23 +- **Recheck trigger**: a revised Opus 5.5 prompting page or usage guide, or the next Opus release. + +The guide's opening and closing summary sections map onto the rows below and carry no rows of +their own. ## How to ask -| Guide item | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| -| G1.1 Say what "done" looks like, then let it run | Root `AGENTS.md` stop rule; `harness-config:audit-prompting-postures` P6 (finish line and both kinds of stop); `docs-hygiene:write-for-agents`; dispatch briefs in codebase-health, batch-simplify, coupling, mutation-testing, review:fanout, course-digest, architecture:improve; implementation and planning briefs | ADOPT | -| G1.2 Stop telling it to "think hard" | `audit-instructions` I8-f (scoped to Opus 5.5 targets, because the model-agnostic best-practices page still recommends thinking steers) and widened I8-c; `write-for-agents` defers to I8-f for the target model; the one live steer found (event-storming simulation) replaced; boris `autonomy.md` carries an amendment note | ADOPT | -| G1.3 Add to a running task | A user habit in the Claude Code UI, with no repository instruction to change | N/A | -| G1.4 Name the design styles to leave out | `audit-instructions` I26 extended; visualization, prototype, playgrounds, education (eli5, teach), and adhd:clarify name exclusions, extend the list when the user dislikes a choice, and render again | ADOPT | +| G1.1 The finish line in a request | Root `AGENTS.md` stop rule; `harness-config:audit-prompting-postures` P6 (finish line and both kinds of stop); `docs-hygiene:write-for-agents`; dispatch briefs in codebase-health, batch-simplify, coupling, mutation-testing, review:fanout, course-digest, architecture:improve; implementation and planning briefs | ADOPT | +| G1.2 Thinking steers in prompts | `audit-instructions` I8-f (scoped to Opus 5.5 targets, because the Opus 5.5 page and the model-agnostic [Thinking and reasoning](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices#thinking-and-reasoning) section disagree on thinking steers) and widened I8-c; `write-for-agents` defers to I8-f for the target model; the one live steer found (event-storming simulation) replaced; boris `autonomy.md` carries an amendment note | ADOPT | +| G1.3 Adding to a running task | A user habit in the Claude Code UI, with no repository instruction to change | N/A | +| G1.4 Excluded design styles | `audit-instructions` I26 extended; visualization, prototype, playgrounds, education (eli5, teach), and adhd:clarify name exclusions, extend the list when the user dislikes a choice, and render again | ADOPT | ## Steering a long run -| Guide item | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| -| G2.1 Tell it which stops you want | Root `AGENTS.md` rule (keep going with status in the same message; stop before destructive or outside-this-checkout actions; keep permission prompts on); postures P6; fable-5 `communication.md`; adhd:shape; knowledge:docpage-digest; implementation:implement-dispatch autonomous mode; autonomy `lane-stop-gate`; session-flow orchestrate and keep-going; the worker and merge lane launch prompts in `prompts/loops/loop-lane-prompts.md`. No existing confirmation gate was weakened. Left by decision: babysit-loop and work-loop (the loop-lane convention keeps those clauses in the launch prompts), continue-in-background (its resume wording is pinned by `save_point.py`) | ADOPT | -| G2.2 Split big work across subagents and check each result | postures P1; review:fanout, mcp-tools:audit, ai-slop rubric fan-out, docs-hygiene:audit-encapsulation, course-digest, discipline fan-out now check each worker's evidence before accepting it. Already present in bugs:scan, codebase-health, plugin-quality, map-corpus, architecture:improve, verification:confirm | ADOPT / COVERED | -| G2.3 Keep the task list in a file | Root `AGENTS.md`; postures P9 (an existing ledger counts). Already present in batch-simplify, coupling, docpage-digest, map-corpus, discovery, disk-hygiene, machine-health, unhobble, audit-pass | ADOPT / COVERED | +| G2.1 The stops a user wants | Root `AGENTS.md` rule (keep going with status in the same message; stop before destructive or outside-this-checkout actions; keep permission prompts on); postures P6; fable-5 `communication.md`; adhd:shape; knowledge:docpage-digest; implementation:implement-dispatch autonomous mode; autonomy `lane-stop-gate`; session-flow orchestrate and keep-going; the worker and merge lane launch prompts in `prompts/loops/loop-lane-prompts.md`. No existing confirmation gate was weakened. Left by decision: babysit-loop and work-loop (the loop-lane convention keeps those clauses in the launch prompts), continue-in-background (its resume wording is pinned by `save_point.py`) | ADOPT | +| G2.2 Subagent splits and per-result checks | postures P1; review:fanout, mcp-tools:audit, ai-slop rubric fan-out, docs-hygiene:audit-encapsulation, course-digest, discipline fan-out now check each worker's evidence before accepting it. Already present in bugs:scan, codebase-health, plugin-quality, map-corpus, architecture:improve, verification:confirm | ADOPT / COVERED | +| G2.3 A task list kept in a file | Root `AGENTS.md`; postures P9 (an existing ledger counts). Already present in batch-simplify, coupling, docpage-digest, map-corpus, discovery, disk-hygiene, machine-health, unhobble, audit-pass | ADOPT / COVERED | ## Checking the result -| Guide item | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| -| G3.1 Read what it needs from you first | Root `AGENTS.md` ("Blocked on me, Changed, Found"); postures P11; reports in bugs, codebase-health, mutation-testing, discovery, architecture, machine-health, batch-simplify, coupling, ai-briefing, harness-config:audit-pass, session-flow (keep-going, reconcile, clean-stop), source-control (babysit-prs, pull-request monitor and readiness), repo-hygiene batch runs, repo-fleet-hygiene apply, work-items drain mode, and the loop-lane launch prompts now lead with what waits on the user. A skill's own report template keeps its headings (for example "Needs you") | ADOPT | -| G3.2 Ask it to review the code | review:code-review and security-review findings carry file and line, why it is wrong, and how to show it fails; quality-gate PR mode leads with merge-blocking findings; harness-config:audit-automation-gaps. The CI lane keeps its "block or flag" bar by decision | ADOPT | -| G3.3 Mark what it couldn't confirm | dometrain grounding added. Already present in discovery research and trace-intent, github:audit, ai-briefing, postures P4 and P5 | ADOPT / COVERED | +| G3.1 What the reader must act on, first | Root `AGENTS.md` ("Blocked on me, Changed, Found"); postures P11; reports in bugs, codebase-health, mutation-testing, discovery, architecture, machine-health, batch-simplify, coupling, ai-briefing, harness-config:audit-pass, session-flow (keep-going, reconcile, clean-stop), source-control (babysit-prs, pull-request monitor and readiness), repo-hygiene batch runs, repo-fleet-hygiene apply, work-items drain mode, and the loop-lane launch prompts now lead with what waits on the user. A skill's own report template keeps its headings (for example "Needs you") | ADOPT | +| G3.2 Code review requests | review:code-review and security-review findings carry file and line, the reason it is a defect, and a way to demonstrate the failure; quality-gate PR mode leads with merge-blocking findings; harness-config:audit-automation-gaps. The CI lane keeps its "block or flag" bar by decision | ADOPT | +| G3.3 Unconfirmed claims | dometrain grounding added. Already present in discovery research and trace-intent, github:audit, ai-briefing, postures P4 and P5 | ADOPT / COVERED | ## In Claude apps -| Guide item | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| -| G4.1 Share the chart or screenshot itself | playwright reads the screenshot file for visual questions. Already present in computer-use and the knowledge video and course digests | ADOPT / COVERED | -| G4.2 Ask it to check a long document | review `doc-drift-detector` checks for self-contradicting numbers, dates, and names, quoting both locations | ADOPT | -| G4.3 Ask for the finished file | wizard already produces the finished script; no skill here returns an outline where a file is wanted | COVERED | -| G4.4 Say when answers are settled | `audit-instructions` I35 flags the instruction on analysis and agentic surfaces, where a later step can show an earlier mistake. No repository surface is a long-chat project | ADOPT (audit only) | +| G4.1 Visual inputs | playwright reads the screenshot file for visual questions. Already present in computer-use and the knowledge video and course digests | ADOPT / COVERED | +| G4.2 Long-document consistency checks | review `doc-drift-detector` checks for self-contradicting numbers, dates, and names, quoting both locations | ADOPT | +| G4.3 Finished files over outlines | wizard already produces the finished script; no skill here returns an outline where a file is wanted | COVERED | +| G4.4 Settled answers | `audit-instructions` I35 flags the instruction on analysis and agentic surfaces, where later steps may expose an earlier error. No repository surface is a long-chat project | ADOPT (audit only) | ## Flagged messages -| Guide item | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| | Fallback on a flagged message | `opus-5-5.md` records the fallback targets; fable-5 meta-rule 3 re-resolves the adaptation chapter against the model now answering; harness-ops known-issues covers the switch-back steps | ADOPT | -| G5.1 Don't ask it to show its reasoning in the reply | `audit-instructions` I10 widened to Opus 5.5 targets; `write-for-agents` never asks the model to reproduce its reasoning; prompts/loops ask for a one-line rationale instead of "your reasoning" | ADOPT | +| G5.1 Reasoning shown in the reply | `audit-instructions` I10 widened to Opus 5.5 targets; `write-for-agents` never asks the model to write out its hidden thinking; prompts/loops ask for a one-line rationale instead of "your reasoning" | ADOPT | ## Speed -| Guide item | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| | G6.1 Fast mode | Recorded in `opus-5-5.md`. A per-session user choice with extra cost, so no skill turns it on | ADOPT (chapter only) | ## Model currency -| Claim | Ours | Verdict | +| Topic | Ours | Verdict | |---|---|---| -| Opus 5.5 is the current Opus; `opus` resolves to it | New `plugins/playbooks/reference/model-adaptation/opus-5-5.md`; live docs (official-docs, plugin-philosophy, loop-lane README) updated | ADOPT | -| Opus 5 and Opus 4.8 chapters | Kept: per model-config (re-fetched 2026-09-28), flagged Fable 5.1, Fable 5, and Opus 5.5 requests re-run on Opus 5 (biology) or Opus 4.8 (cybersecurity). Each chapter states this at its top with Claim/Basis/As of/Recheck. #4349 stays open until that trigger fires | KEEP, TRACK #4349 | +| The current Opus and the `opus` alias (pointer: [Model aliases](https://code.claude.com/docs/en/model-config#model-aliases), as of 2026-09-23) | New `plugins/playbooks/reference/model-adaptation/opus-5-5.md`; live docs (official-docs, plugin-philosophy, loop-lane README) updated | ADOPT | +| Opus 5 and Opus 4.8 chapters | Kept while they are fallback targets for a flagged request (pointer: [Automatic model fallback](https://code.claude.com/docs/en/model-config#automatic-model-fallback), as of 2026-09-28). Each chapter carries its own record of this at its top. #4349 stays open until that section no longer names them | KEEP, TRACK #4349 | ## Decisions and follow-ups Decisions made by the owner on 2026-09-23: -- `review:code-review` keeps its "block or flag" bar; each finding gains the guide's evidence. +- `review:code-review` keeps its "block or flag" bar; each finding gains file and line, the reason + it is a defect, and a way to demonstrate the failure. - The visualization chrome keeps its off-white background and monospace labels as the house look; its exclusion list names only the habits the chrome does not set. - The `audit-instructions` think-carefully row (I8-f) stays scoped to Opus 5.5 targets. diff --git a/docs/upstream/prompting-sonnet-5-5-guide.md b/docs/upstream/prompting-sonnet-5-5-guide.md deleted file mode 100644 index c25b9e6b22..0000000000 --- a/docs/upstream/prompting-sonnet-5-5-guide.md +++ /dev/null @@ -1,61 +0,0 @@ -# Upstream source: "Prompting Claude Sonnet 5.5" - -## Contents - -- [Status](#status) -- [Source and verification](#source-and-verification) -- [Guide sections](#guide-sections) -- [Model currency](#model-currency) -- [Left to other work](#left-to-other-work) - -This is the provenance record for applying Anthropic's Sonnet 5.5 prompting guide across this -marketplace. It follows the record pattern set by [opus-5-5-usage-guide.md](opus-5-5-usage-guide.md). -Each guide section has one row. A row names the surfaces that carry the guidance, gives evidence -that they already did, or says why the section changes nothing in repository instructions. The -guidance itself stays at the guide: each surface states our rule in our words and points at the -live section. - -## Status - -Applied under #5685. The model-adaptation chapter is -`plugins/playbooks/reference/model-adaptation/sonnet-5-5.md`; meta-rule 3 in the `fable-5` skill -routes to it. - -## Source and verification - -- Guide: , - raw `.md` read 2026-10-01 (27,412 bytes, MD5 `2bcb67cc9f72b68e8823f197034c06d6`). -- Read the same day for model currency: (111,803 - bytes) and (17,143 bytes). -- Vendor-reported figures in the guide (the cost cut from the reviewer-subagent paragraph, the - accuracy gain from the think-first line) are recorded as vendor-reported, not as verified facts. -- Recheck trigger: a re-fetch of the guide diverging from a row below, or a later Sonnet release. - -## Guide sections - -| Guide section | Ours | Verdict | -|---|---|---| -| Calibrate effort | The chapter's effort and thinking sections. `audit-instructions` I21 already covers it generically (models released after Opus 5.5 start at their own default). The `effort:` pins on Sonnet-bound agents were set against Sonnet 5, and the guide says to sweep rather than carry a level over | ADOPT (chapter), TRACK #5682 | -| Steer initiative and scope | The chapter's low-effort and initiative sections. `AGENTS.md` ("When to stop and when to keep going") already states the keep-going and stop rule. `audit-prompting-postures` P2 and P6 cite the section | ADOPT (chapter, postures), already present (`AGENTS.md`) | -| Running without up-front thinking | `audit-instructions` I8-c widened to `sonnet-5-5`. I17 is not extended: the disable rejection for this model is sourced on the what's-new page, not the guide, and its remediation differs from I17's "leave thinking on" because `between_tools` exists | ADOPT (audit), KEEP I17 | -| Reasoning tasks with JSON output | API-side; the chapter carries it. I8-f is deliberately not widened, since the guide prescribes a think-first line for these tasks | ADOPT (chapter only) | -| User-facing progress updates | The chapter. No skill or agent body tells a model to hold findings for the final reply (searched for "hold", "only at the end"; no match). `audit-instructions` I8-e stays scoped to `sonnet-5`, and exact-match scoping keeps it off 5.5 | ADOPT (chapter), no site elsewhere | -| Tool use in chat and knowledge work | The chapter. No skill or agent text discourages tool use (searched "strictly necessary", "minimize tool calls"; no match). I28 covers the over-eager direction only | ADOPT (chapter), no site to fix | -| Mid-turn user messages | Harness-side. The chapter says nothing changes for the model and that tool-result text stays data | ADOPT (chapter only) | -| Verification on coding tasks | The chapter's low-effort section. `audit-prompting-postures` P5 cites the section. The `implementer` binds `opus` at `high`; the Sonnet-bound agents (four review agents and the explorer) investigate rather than edit source | ADOPT (chapter, postures) | -| Tolerant tool-call handling | Harness-side; the chapter tells the model to copy declared names exactly | ADOPT (chapter only) | -| Tools for complex visual inputs | The chapter. A shared crop and zoom capability is its own effort | ADOPT (chapter), TRACK #5681 | -| Safeguard refusals | The chapter's safeguards section. `audit-instructions` I10 widened to `sonnet-5-5`. `fable-5` meta-rule 3 names the Sonnet 5.5 fallback targets. The loop-lane known gap now covers the fast tier | ADOPT | - -## Model currency - -| Claim | Ours | Verdict | -|---|---|---| -| Sonnet 5.5 is the current Sonnet; `sonnet` resolves to it on the Anthropic API | `docs/official-docs.md` row, loop-lane alias binding (fast tier, reliable cutoff), `docs/plugin-philosophy.md` tier table | ADOPT | -| The Sonnet 5 chapter stays | Sonnet 5 is a legacy model and the cybersecurity fallback target for Sonnet 5.5. The 5.5 chapter lists what it reverses, narrows and leaves open | KEEP | -| Queue the guide and the what's-new page for a digest | `plugins/knowledge/skills/docpage-digest/context/anthropic-docs-queue.md` | ADOPT | - -## Left to other work - -- Retuning the Sonnet-bound agents' effort pins is the effort eval sweep, #5682. -- The shared crop and zoom capability is #5681. diff --git a/plugins/ai-slop/.claude-plugin/plugin.json b/plugins/ai-slop/.claude-plugin/plugin.json index 20d4308148..79f9812f26 100644 --- a/plugins/ai-slop/.claude-plugin/plugin.json +++ b/plugins/ai-slop/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "ai-slop", - "version": "0.12.4", + "version": "0.12.5", "description": "Detects and removes AI-writing tells (slop) in checked-in markdown prose: em dashes, emoji formatting, AI vocabulary, negative parallelisms, chatbot phrases, filler, stacked hedging, citation artifacts, model-era phrases, and the rest of a catalog distilled from Wikipedia's Signs of AI writing plus a repo-owned, evidence-graded inventory of current-generation model vocabulary. Read-only audit by default with a deterministic detector plus a judgment rubric; an explicit fix action rewrites findings behind a semantic-diff guard. Findings conform to the detector-findings convention so the review fanout fix relay can consume them.", "author": { "name": "Melodic Software", diff --git a/plugins/ai-slop/CHANGELOG.md b/plugins/ai-slop/CHANGELOG.md index 58a2a2d1a8..4e1d328053 100644 --- a/plugins/ai-slop/CHANGELOG.md +++ b/plugins/ai-slop/CHANGELOG.md @@ -1,5 +1,16 @@ # Changelog +## [0.12.5] - 2026-10-02 + +### Changed + +- **The `audit` catalog stores no text from its source page.** The attribution record now points at + the pinned Wikipedia revision for the source inventory, and the postures derived from the page's + Caveats section and the notes on the em-dash and similar rules are restated as this catalog's own + decisions. No rule, default or severity changes: the em-dash rule stays zero-tolerance by default. +- The feeling-instead-of-mechanism rubric item is restated in our own example and tests, and the + archiewood/claudeisms ratios now sit behind a pointer to that repository instead of in the catalog. + ## [0.12.4] - 2026-10-02 ### Fixed diff --git a/plugins/ai-slop/skills/audit/reference/catalog.md b/plugins/ai-slop/skills/audit/reference/catalog.md index 369c3d3b28..f607c561ed 100644 --- a/plugins/ai-slop/skills/audit/reference/catalog.md +++ b/plugins/ai-slop/skills/audit/reference/catalog.md @@ -46,9 +46,14 @@ The "Cursor unslop additions" section was inspired by ## Upstream-drift record -- **Claim**: this catalog's tell inventory derives from the source page revision cited above. -- **Basis**: the revision-pinned URL in the attribution block. -- **As of**: 2026-08-17. +This catalog's tell inventory derives from the pinned source revision named in the attribution +block. Each tell is this repository's own firing rule, reworded and classified for use outside +Wikipedia; the page's text is read at the pointer, not stored here. + +- **Pointer**: for the source inventory, see revision + [1369699198](https://en.wikipedia.org/w/index.php?title=Wikipedia:Signs_of_AI_writing&oldid=1369699198) + of . +- **As of**: 2026-08-17 - **Recheck trigger**: each `ai-slop` release and each fleet audit. Per-revision rechecking was rejected: the page was measured at 50+ edits/week (2026-08-17), so a per-revision trigger would fire continuously. @@ -98,28 +103,28 @@ placement gate are defined at the top of that section. ## False-positive posture (source Caveats) -Mined from the pin's Caveats section (byte-identical on the live head; extraction closed -2026-08-25). Three source statements bind how this catalog's verdicts are read: - -- **The signs are descriptive, not prescriptive.** The source: "do not merely treat these signs - as the problems to be fixed; that could just make detection harder." This plugin's fix flow - is therefore framed as house style (better prose on its own merits), never detector evasion; - the rewrite guide's non-evasion posture carries the operational test. -- **Expert false-positive rate.** The source's calibration figure: an experienced LLM-output - patroller who tags 10 pages has probably made one false positive. A deterministic subset of - those signs run over a technical corpus is not better calibrated than the experts; verdicts - are evidence for a rewrite decision, never proof of provenance, and accusatory framing - ("this is AI-written") is outside this plugin's vocabulary. -- **Combination over isolation.** Individual signs are weak alone; the source repeats per-sign - that combination strengthens a verdict. Density thresholds, the minimum-hits floor, and the - rubric's counter-sign tempering are this catalog's mechanical forms of that instruction. +Derived from the pin's Caveats section (byte-identical on the live head; extraction closed +2026-08-25), which is read at the pointer in the upstream-drift record. Three postures bind how +this catalog's verdicts are read: + +- **The signs are descriptive, not prescriptive.** This plugin's fix flow is framed as house + style (better prose on its own merits), never detector evasion; the rewrite guide's + non-evasion posture carries the operational test. +- **Expert false-positive rate.** The Caveats section gives a false-positive rate for + experienced human reviewers. A deterministic subset of those signs run over a technical corpus + is not better calibrated than those reviewers; verdicts are evidence for a rewrite decision, + never proof of provenance, and accusatory framing ("this is AI-written") is outside this + plugin's vocabulary. +- **Combination over isolation.** This catalog treats an individual sign as weak alone and a + combination as stronger. Density thresholds, the minimum-hits floor, and the rubric's + counter-sign tempering are its mechanical forms of that posture. ## Quotation exemption (policy-level) -Stated once here and inherited by every rule; the design follows Wikipedia's MOS "principle of -minimal change" for quoted material (quotations are not the repo's own prose to restyle) and the -detector implements it mechanically. No rule scans fenced code, inline code spans, or -ignore-marked lines, whatever its class. Each rule carries a class: +Stated once here and inherited by every rule; the design follows the Wikipedia Manual of Style's +principle of minimal change for quoted material (quotations are not the repo's own prose to +restyle) and the detector implements it mechanically. No rule scans fenced code, inline code +spans, or ignore-marked lines, whatever its class. Each rule carries a class: - **wording**: the rule judges prose the repo AUTHORS. It never scans quoted material: blockquote lines and double-quoted spans are removed from its input. @@ -317,7 +322,7 @@ then-current 1,361-file tracked-markdown corpus: - applicability: general-prose - v1: script - Three constructions: "not just X, but also Y"; "not X, but Y" (including "isn't X; it's Y"); - "X rather than Y" (noted by the source as characteristic of Grok output). + "X rather than Y" (the source ties this one to a particular model family). ### rule-rule-of-three: Rule of three @@ -340,9 +345,8 @@ then-current 1,361-file tracked-markdown corpus: - applicability: general-prose - v1: recorded-only - Synonym-cycling to avoid repeating a word a human would simply repeat. -- Demoted from the active rubric 2026-08-25: the live source page moved this sign to its - Historical indicators (base rate collapsed in current model output), and this catalog - follows the upstream demotion rather than keeping an era-bound tell active. +- Demoted from the active rubric 2026-08-25, following the live source page's own move of this + sign to its historical section, rather than keeping an era-bound tell active. ## Style @@ -397,21 +401,17 @@ then-current 1,361-file tracked-markdown corpus: outside code fences and inline code flags. Documents that require em dashes opt out per-document via config path-lists or the in-file marker; the rule is never threshold-calibrated and is excluded from the `recorded-only` demotion path. -- The source page's Style section (catalog pin and the 2026-08-21 recheck) treats this as a - **valid sign**, not an ineffective one. The same section carries the qualifier *"This sign - is most useful when taken in combination with other indicators, not by itself."* That is a - corroboration note on a kept tell, not a listing under **Ineffective indicators** (checked - explicitly; see that section). The shipped default stays zero-tolerance: this plugin is a - house-style detector, not a Wikipedia AI-authorship tribunal. A consuming repo that wants the - source's combination reading disables the rule or uses `em_dash_allowed_paths` (or the - generalized `rule_allowed_paths`). -- **Spacing qualifier (mined 2026-08-25 from the same pinned section):** the source - distinguishes SPACED em dashes (`—` with a space on each side) as the stronger AI tell, - while unspaced em dashes are the typographically informed human convention; it cites - reporting (The Economist, 2026-07-30, wiki-cited, not independently verified here) that - among current models only Claude still over-uses them. The shipped rule stays - character-level zero-tolerance as house style, and records the spacing discriminator here - for any consuming repo calibrating a softer setting. +- The source page's Style section (catalog pin and the 2026-08-21 recheck) lists this as a + **valid sign**, not an ineffective one, with its own qualifier about weighing it alongside + other signs. That qualifier is a corroboration note on a kept tell, not a listing under + **Ineffective indicators** (checked explicitly; see that section). The shipped default stays + zero-tolerance: this plugin is a house-style detector, not a Wikipedia AI-authorship tribunal. + A consuming repo that wants the source's combination reading disables the rule or uses + `em_dash_allowed_paths` (or the generalized `rule_allowed_paths`). +- **Spacing qualifier (mined 2026-08-25 from the same pinned section):** the source's section + separates spaced from unspaced em dashes as tells; read the distinction there. The shipped rule + stays character-level zero-tolerance as house style. A consuming repo calibrating a softer + setting can match only spaced em dashes (`—` with a space on each side). - **Zero-tolerance is a house-style choice, not a detection claim.** The false-accusation literature the source's Caveats cite is one more reason this rule's verdict is "this repo does not use em dashes", never "this text is AI-written". @@ -480,7 +480,7 @@ then-current 1,361-file tracked-markdown corpus: - v1: script - Assistant-frame residue, both halves of the source section (extraction completed 2026-08-25; the original ERE covered roughly one of the section's six words-to-watch families and missed - even the source's own example "as of my last knowledge update"): + even one of the source's own examples, now in the cutoff list below): - Cutoff half: "as of my knowledge cutoff", "as of my last (knowledge) update", "up to my last training update", "I cannot browse", "as an AI (language) model". - Source-gap (RAG-era) half: "while specific details are limited/scarce", "not widely @@ -593,7 +593,7 @@ then-current 1,361-file tracked-markdown corpus: Fetch gap closed 2026-08-21 (see the upstream-drift record). The source section is Wikipedia talk-page comments, so every tell classifies `wikipedia-specific` / `recorded-only`. They have -no general-prose analogue worth a script rule. Quoted from the catalog pin (revision +no general-prose analogue worth a script rule. Distilled from the catalog pin (revision 1369699198, parse section 62) and confirmed on the live page (revision 1370403579). One of the seven tells already has a slug under Edit summaries: downplaying AI use by @@ -776,29 +776,21 @@ pin (revision 1369699198, parse section 80, 2026-08-16) and the live recheck (re 1370403579, retrieved 2026-08-21 from ). -Quoted from the pin (CC BY-SA 4.0; ellipses mark dropped citation/example markup): - - -> False accusations of AI use can drive away new editors and foster an atmosphere of -> suspicion. […] Here are several somewhat commonly used indicators that are ineffective -> in LLM detection—and may even indicate the opposite. - - -- **Perfect grammar**: skilled human writers also produce this. -- **Combination of casual and formal registers**, or language that sounds both "clinical" - and "emotional": technical-field casual writing, mixed registers, or multi-editor pages. -- **"Bland" or "robotic" prose**: LLM output has *specific* traits; "robotic" is not one. -- **"Fancy", "academic", or "formal" prose**: in the page's own wording, LLMs favor *specific - words*, and "the correlation does not extend to all formal, academic, or 'fancy'-sounding - prose." `rule-ai-vocabulary` is the specific-word rule, not a formality detector. -- **Transition words (in isolation)**: older output overused a few (`Additionally`, - `Consequently`, `Notably`); "this is not a strong tell." The shipped vocabulary list - already dropped `additionally` for legitimate technical use; there is no standalone - transition-words rule. -- **Unsourced content**: most uncited articles predate LLMs; modern chatbots also cite. -- **Bizarre wikitext**: random HTML/VisualEditor artifacts are *not* the LLM markup tells - already catalogued under Markup. -- **Correct wikitext**: correct formatting is normal. +The eight, named only; the page's reasons for each are read at the pointer in the +[upstream-drift record](#upstream-drift-record), not stored here. Each line states what this +catalog does about it: + +- **Perfect grammar**: no rule. +- **Combination of casual and formal registers**: no rule. +- **Bland or robotic prose**: no rule. +- **Fancy, academic, or formal prose**: no rule. `rule-ai-vocabulary` is the specific-word rule, + not a formality detector. +- **Transition words (in isolation)**: no standalone transition-words rule. The shipped + vocabulary list already dropped `additionally` for legitimate technical use. +- **Unsourced content**: no rule. +- **Bizarre wikitext**: no rule; these are not the LLM markup tells already catalogued under + Markup. +- **Correct wikitext**: no rule. None of those eight is a shipped script rule, a shipped rubric tell, or a Cursor-addition slug. No drop or re-scope follows. @@ -906,9 +898,9 @@ either layer, and those rows say so. that" (for "because"), "it is important to note that", "it is worth noting that", "it should be noted that" (all deletable). Fires per occurrence; each hit has a mechanical rewrite. - **Recorded divergence from the source (2026-08-25):** the source's Syntax counter-sign list - names "isolated wordy constructions such as 'in order to'" among signs of HUMAN writing (an - uncited bullet, and the study its neighboring bullet cites does not measure this - construction). This rule keeps flagging it deliberately: the plugin's goal is concise house + counts the "in order to" construction among signs of HUMAN writing (an uncited bullet, and the + study its neighboring bullet cites does not measure this construction). This rule keeps + flagging it deliberately: the plugin's goal is concise house style, not authorship attribution, and "in order to" -> "to" is de-verbosing every style authority endorses. The divergence is a house-style choice, recorded rather than hidden. @@ -962,8 +954,9 @@ either layer, and those rows say so. "harness" naming an actual test harness is not a tell. Calibration kept this out of the script layer (see the calibration record's second pass). Replacements live in `rewrite-guide.md`. - **Model-era cues (2026-08, from the "Model-era additions" section)**: "load-bearing" and - "seam" join the cue list as the flagship 2026 Claude-family metaphor words (load-bearing - measured at >7,500x its Stack Overflow base rate in Claude Code output; seam at 62x). Their + "seam" join the cue list as the flagship 2026 Claude-family metaphor words (both measured far + above their Stack Overflow base rate in Claude Code output; the figures are at the + archiewood/claudeisms pointer under "Model-era additions"). Their literal boundary is BROAD, deliberately: a Feathers seam in refactoring/testing prose, a load-bearing wall, and a load-bearing invariant or instruction NAMED as such deliberately in architecture prose are all terms of art, not tells. The tell is the reflexive metaphor where @@ -1004,11 +997,10 @@ either layer, and those rows say so. - detectability: judgment - applicability: general-prose - v1: rubric -- A sentence naming a feeling about the thing ("SQL you can read", "the database stays close at - hand") where the reader needs the mechanism or the number ("`.toSQL()` returns the exact string - sent to the database"). Two tests: can the sentence be restated as a concrete instruction, - fact, or number (if not, cut it); and could it appear unchanged in another project's docs (if - so, it says nothing about this one). +- A sentence that names how the thing feels where the reader needs its mechanism or a number + ("the config feels lightweight" where "the config file has four keys" would serve). Report it + when it gives the reader nothing to act on or check (no step, fact, or number), or when it is + generic enough to fit any project's docs. ## Model-era additions (repo-owned) @@ -1105,12 +1097,11 @@ README's "Updating the model-era inventory". - era: 2025-2026 - models: Claude family (Claude Code agentic output specifically) - evidence: measured (single pool) -- The archiewood/claudeisms measurement: 145 words at >=20x frequency versus a 1.72B-word - Stack Overflow comment corpus, measured on ~5 weeks of the author's own Claude Code - Opus 4.7/4.8 transcripts. Author-declared confounds: workflow skew, an installed style - plugin, system-prompt priming, no stemming. Representative ratios: gating 759x, dedup 647x, - decisive 637x, verdict 631x, scaffolds 570x, settles 434x, handoff 358x, genuinely 224x, - errored 215x, drift 183x, pre-existing 183x, silently 36x, verbatim 31x, canonical 27x. +- The archiewood/claudeisms measurement: one author's Claude Code transcripts ranked against a + Stack Overflow comment corpus. The word list, ratios and the author's own caveats are at + (README and `claudeisms.csv`). We class it + single-pool: one author, one workflow. As of: 2026-10-01. Recheck trigger: the README's ranked + tables change, or a second independent frequency pool appears. - ONE of these ships in the default vocabulary list: `pre-existing` passed the measured quiet-gate test on this corpus (2026-08-27: 61 files contain the word, the density gate fired on none of them, the same measurement that admitted "leverage"), so it joined the @@ -1125,23 +1116,25 @@ README's "Updating the model-era inventory". ### Model-era record -- **Claim**: the entries above reflect the community-documented model-vocabulary layer as of - the dates below, and neither upstream inventory carries it. -- **Basis**: per-entry sources; upstream absence verified against the live Wikipedia page and - the Cursor skill head. +This section holds the model-vocabulary layer this repository tracks from community sources, and +keeps it here because neither upstream inventory carried it when checked. + +- **Pointer**: the per-entry sources named in each entry and in the record below; for the + absence check, the live Wikipedia page and the Cursor skill head named in the record. +- **As of**: 2026-08-26 - **Recheck trigger**: each `ai-slop` release, each new frontier-model generation, and, for `rule-model-era-vocabulary`, whether a second independent frequency pool has landed (the cluster's promotion condition, which no other trigger would look for). - **Record (2026-08-26, initial)**: layer established from the Hacker News thread 48905248 - (609 points), archiewood/claudeisms (two-measurement corroboration for "load-bearing": - >7,500x lower-bound ratio, and Marek Suppa's independent 188-in-1.7M-words count), + (609 points), archiewood/claudeisms (two-measurement corroboration for "load-bearing": its + lower-bound ratio and Marek Suppa's independent count), anthropics/claude-code issue 53454 (maintainer-reproduced), Velitchkov's cliché catalog, crystl.dev's hacker-idiom catalog, and jola.dev's filter hook. Wikipedia "Signs of AI writing" head revision 1371415133 (fetched 2026-08-26) and Cursor unslop head (last commit 2026-08-02) both carry none of it. Harness confound recorded on the metaphor cues: the - version-tracked Piebald-AI system-prompt mirror carries "give brief updates when you find - something load-bearing or change direction" verbatim, so "load-bearing" in Claude Code - output is partly prompt-primed rather than purely model-weight; the frequency spike aligns + version-tracked Piebald-AI system-prompt mirror uses the word "load-bearing" in its + progress-update instruction, so "load-bearing" in Claude Code output is partly prompt-primed + rather than purely model-weight; the frequency spike aligns with the Opus 4.6 release date and the word appears in non-Code output, so the weights-side claim stays alive at MEDIUM. A harness prompt change can therefore collapse a phrase's base rate overnight. Attribution notes exist so a recheck knows which entries die that way. diff --git a/plugins/architecture/.claude-plugin/plugin.json b/plugins/architecture/.claude-plugin/plugin.json index 9c7f4c924d..0cfd3e870f 100644 --- a/plugins/architecture/.claude-plugin/plugin.json +++ b/plugins/architecture/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "architecture", - "version": "0.18.1", + "version": "0.18.2", "description": "Uses Ousterhout's deep-module lens to scan an existing codebase for module-level architecture friction: shallow modules, seam leaks, and locality gaps. Presents candidates as a self-contained HTML report, then runs an interview loop on the selected candidate before handing off for planning. Also charts a repository and the systems it references as a C4 system landscape plus an application-portfolio table, committing the result as a record that later runs check for drift, and below that altitude cites a build-declaration graph and draws component, container, context, sequence, event, data, and deployment views from tested extractors. Records an architecture decision into the repository's existing ADR convention.", "author": { "name": "Melodic Software", diff --git a/plugins/architecture/CHANGELOG.md b/plugins/architecture/CHANGELOG.md index 29dd118234..bd9fff001a 100644 --- a/plugins/architecture/CHANGELOG.md +++ b/plugins/architecture/CHANGELOG.md @@ -3,6 +3,15 @@ All notable changes to the `architecture` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.18.2] - 2026-10-02 + +### Changed + +- `record-decision` states its handling of the ADR template repository's license as our decision: + the README content is treated as CC BY-NC-SA 4.0 and each bundled template as carrying its own + license, and the skill names the license when it declines to paste. The record points at the + repository's `LICENSE.md` and stores none of its text. + ## [0.18.1] - 2026-10-02 ### Fixed diff --git a/plugins/architecture/skills/record-decision/SKILL.md b/plugins/architecture/skills/record-decision/SKILL.md index 61342499bf..1534391f2a 100644 --- a/plugins/architecture/skills/record-decision/SKILL.md +++ b/plugins/architecture/skills/record-decision/SKILL.md @@ -98,10 +98,12 @@ Templates for the human to read, cited by URL: License: **CC BY-NC-SA 4.0** (`https://creativecommons.org/licenses/by-nc-sa/4.0/`). -- **Claim**: README-authored content in that repository is licensed CC BY-NC-SA 4.0; the bundled - templates carry their own licenses, stated per template. -- **Basis**: `https://raw.githubusercontent.com/joelparkerhenderson/architecture-decision-record/main/LICENSE.md`. -- **As of**: 2026-09-06. +This skill treats the repository's README content as CC BY-NC-SA 4.0 and each bundled template +as carrying its own license, and names the license when it declines to paste. + +- **Pointer**: for the repository's license terms, see + `https://raw.githubusercontent.com/joelparkerhenderson/architecture-decision-record/main/LICENSE.md`. +- **As of**: 2026-09-06 - **Recheck trigger**: that `LICENSE.md` changes, or the repository moves. Rule beside it: catalog templates are cited for the human to read, never pasted into this skill, into diff --git a/plugins/attribution/.claude-plugin/plugin.json b/plugins/attribution/.claude-plugin/plugin.json index b20ee8f815..12ceb29481 100644 --- a/plugins/attribution/.claude-plugin/plugin.json +++ b/plugins/attribution/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "attribution", - "version": "0.8.5", + "version": "0.9.0", "description": "Finds prose in tracked markdown that restates content an external source owns (vendor docs, blogs, articles) without adequate attribution, confirms the source, and refactors the copy into a pointer, a citation, or a dated stamped record. Documentation attribution, not software supply chain. Nomination and judgment are LLM work; the scripts do only reasoning-free work (corpus scoping, breadcrumb extraction, stamp expiry, fingerprint compare of two concrete texts). Read-only audit by default; explicit fix and sweep actions apply dispositions behind a semantic-diff guard and live pointer verification. Findings conform to the detector-findings convention.", "author": { "name": "Melodic Software", diff --git a/plugins/attribution/CHANGELOG.md b/plugins/attribution/CHANGELOG.md index 8bc98e15b1..cea018f47e 100644 --- a/plugins/attribution/CHANGELOG.md +++ b/plugins/attribution/CHANGELOG.md @@ -1,5 +1,26 @@ # Changelog +## [0.9.0] - 2026-10-02 + +### Changed + +- **The stamped record a repair writes is now the links-only shape.** `condense-to-stamped-record` + writes the surface's own decision in its own words, a pointer to the exact source section, an + as-of date and a recheck trigger, and stores no source text, quoted or paraphrased. It replaces + the claim, basis URL, as-of date and trigger shape. The README, `SKILL.md`, `dispositions.md`, + `nomination.md`, `persist-findings.md`, the evals and the restated-fact finding's suggested fix + name the new shape, and a pointer-shaped record is read by `check-stamps.sh` through its `As of` + and `Recheck trigger` bullets, which a new test case covers. +- **The rubric's record test reads the new shape.** A passage is conforming when a whole record + (pointer, as-of date, observable recheck trigger) sits beside the decision it records, and text + inside a record that matches the source is still judged like any other passage. The R4 criterion, + the conforming-record carve-out and the refutation prompt use that test. The org rules the rubric + applies, and the fetch-route record in `source-fetch.md`, are restated as pointer records with an + as-of date of 2026-10-01. +- **The fetch rule "No verbatim quote, no claim" becomes "No read, no verdict".** A verdict rests on + the text read and names the matched span in the run's report; the span never enters the file + being repaired. + ## [0.8.5] - 2026-10-02 ### Fixed diff --git a/plugins/attribution/README.md b/plugins/attribution/README.md index 3a781e74b1..a2571d36cb 100644 --- a/plugins/attribution/README.md +++ b/plugins/attribution/README.md @@ -11,9 +11,9 @@ signing, or SLSA. A copied paragraph starts accurate and silently stops being accurate the next time the upstream page changes. Nothing in the repository records that it drifted. Citing the source and fetching -it at read time removes the drift risk entirely; where a surface must restate a volatile -specific to function, a four-part stamped record (claim, basis URL, as-of date, recheck trigger) -keeps the restatement honest and re-checkable. +it at read time removes the drift risk entirely; where a surface must act on a volatile specific +without the source, a stamped record (the surface's own decision in its own words, a pointer, an +as-of date, a recheck trigger) keeps it honest and re-checkable without copying the source. ## Skills @@ -68,11 +68,14 @@ own line, and the closing marker is required: An existing `provenance:source` fence is still recognized: the breadcrumb extractor reads any URL-carrying HTML comment fence and matches no marker name. -A stamped record is prose, not a marker, and carries all four parts: +A stamped record is prose, not a marker. It states the surface's own decision and stores no +source text: ```markdown -The runner accepts three values (per https://example.com/docs/page, as of 2026-08-27; -recheck when the CLI's major version changes). +The pipeline passes `--mode strict` to the runner. +- **Pointer**: for the runner's accepted modes, see https://example.com/docs/page#modes. +- **As of**: 2026-08-27 +- **Recheck trigger**: the CLI's major version changes. ``` There is no per-instance suppression marker, deliberately. Allowances are categorical: vendored @@ -153,7 +156,7 @@ a conforming stamped record, plus finding the authoritative source and condensin `code-tidying:audit-comment-residue`. - The stamped-record format and the fetch route are owned by the upstream-drift convention. This plugin implements checks against them and carries an operational restatement in - `skills/audit/reference/source-fetch.md` as a four-part record citing that convention, because + `skills/audit/reference/source-fetch.md` as a stamped record citing that convention, because the plugin ships to consumers who do not have that repository. ## Untrusted content diff --git a/plugins/attribution/skills/audit/SKILL.md b/plugins/attribution/skills/audit/SKILL.md index fd166a33a0..50137f9770 100644 --- a/plugins/attribution/skills/audit/SKILL.md +++ b/plugins/attribution/skills/audit/SKILL.md @@ -1,5 +1,5 @@ --- -description: "Audit tracked markdown for prose restating content an external source owns without a pointer or a stamped record, and convert copies into links, citations, or four-part stamped records. Breadcrumb-first, then budgeted search. Two rubrics: copy, and restated fact (a default or limit, any wording). Evidence-gated tiers; only fingerprint-confirmed copies are fix-eligible. Only a unanimous restated fact that survives refutation relays, report-only. Also flags verification stamps past their expiry window. Use when: 'find copied content', 'is this copied from the docs', 'check our docs for copied text', 'replace copies with links', 'find stale verification stamps', 'audit provenance', 'where did this paragraph come from', or before publishing prose that restates an upstream page. Read-only by default; explicit 'fix' applies dispositions behind a semantic-diff guard and live pointer checks, and 'sweep' adds per-file closure. Empty target audits tracked markdown." +description: "Audit tracked markdown for prose restating content an external source owns without a pointer or a stamped record, and convert copies into links, citations, or stamped pointer records. Breadcrumb-first, then budgeted search. Two rubrics: copy, and restated fact (a default or limit, any wording). Evidence-gated tiers; only fingerprint-confirmed copies are fix-eligible. Only a unanimous restated fact that survives refutation relays, report-only. Also flags verification stamps past their expiry window. Use when: 'find copied content', 'is this copied from the docs', 'check our docs for copied text', 'replace copies with links', 'find stale verification stamps', 'audit provenance', 'where did this paragraph come from', or before publishing prose that restates an upstream page. Read-only by default; explicit 'fix' applies dispositions behind a semantic-diff guard and live pointer checks, and 'sweep' adds per-file closure. Empty target audits tracked markdown." argument-hint: "[audit|fix|sweep] [target]" user-invocable: true disable-model-invocation: false @@ -33,12 +33,13 @@ as such in the audit's declined/limits section, never read it as an empty config ## Purpose Find prose in tracked markdown that restates content an external source owns, and convert it -into a pointer, a quoted citation, or a four-part stamped record. +into a pointer, a quoted citation, or a stamped record. The harm being reduced is drift, not plagiarism. A copied paragraph starts accurate and stops being accurate the next time the upstream page changes, with nothing in the repository recording that it did. Citing the source and fetching it at read time removes that risk; a stamped record -keeps it honest where a surface must restate a specific to function. +(the surface's own decision, a pointer, an as-of date and a recheck trigger) keeps it honest where +a surface must act on a specific without the source. Detection is LLM-led and breadcrumb-first. The deterministic scripts do only reasoning-free work (path filtering, breadcrumb extraction, date arithmetic, fingerprint comparison of two concrete diff --git a/plugins/attribution/skills/audit/context/persist-findings.md b/plugins/attribution/skills/audit/context/persist-findings.md index 3b29d34f53..199c64186f 100644 --- a/plugins/attribution/skills/audit/context/persist-findings.md +++ b/plugins/attribution/skills/audit/context/persist-findings.md @@ -291,7 +291,7 @@ and never carry a row forward from a previous run. is repaired by re-deriving the record against its live basis; a trigger-less stamp is repaired by writing the observable event that obliges re-derivation; a restated fact is report-only, so its row names the two dispositions (a pointer at the point of use, or a - four-part record when the surface must work offline) and never a fix invocation, because `fix` + stamped record when the surface must work offline) and never a fix invocation, because `fix` and `sweep` never reach the class. - **`Tier`** is LOOKED UP from the rule's crosswalk row, never chosen per finding, then mapped to the consuming project's severity vocabulary when it defines one. diff --git a/plugins/attribution/skills/audit/evals/evals.json b/plugins/attribution/skills/audit/evals/evals.json index 7d0d99c146..614563add4 100644 --- a/plugins/attribution/skills/audit/evals/evals.json +++ b/plugins/attribution/skills/audit/evals/evals.json @@ -85,10 +85,10 @@ "name": "offline-load-bearing-condenses-not-points", "prompt": "/attribution:audit fix a fingerprint-confirmed copy that sits in a reference file a subagent reads mid-dispatch, with no network available at read time. Choose and justify the disposition.", "narration": true, - "expected_output": "The disposition is condense-to-stamped-record, not convert-to-pointer, because the surface must work when the source is unreachable. The record carries all four parts: claim, basis URL, as-of date, and a recheck trigger naming an observable event. The as-of date is written in ISO 8601 so the stamp checker can parse it later.", + "expected_output": "The disposition is condense-to-stamped-record, not convert-to-pointer, because the surface must work when the source is unreachable. The record carries the surface's own decision in its own words, a pointer to the source section, an as-of date, and a recheck trigger naming an observable event, and stores no source text. The as-of date is written in ISO 8601 so the stamp checker can parse it later.", "expectations": [ "Chooses condense-to-stamped-record and rejects a bare convert-to-pointer for an offline-load-bearing surface", - "Writes all four parts, including a recheck trigger that names an observable event rather than a date or a review cadence", + "Writes the decision, the pointer, the as-of date and a recheck trigger that names an observable event rather than a date or a review cadence, with no source text in the record", "Uses an ISO 8601 date so check-stamps.sh can parse the record on a later run", "Runs the semantic-diff guard blind to the rewrite rationale and reverts any flagged hunk before closing the file" ] diff --git a/plugins/attribution/skills/audit/reference/dispositions.md b/plugins/attribution/skills/audit/reference/dispositions.md index c09e30521f..2ecbec08cc 100644 --- a/plugins/attribution/skills/audit/reference/dispositions.md +++ b/plugins/attribution/skills/audit/reference/dispositions.md @@ -16,7 +16,7 @@ Three edit. Two do not. |---|---|---| | `convert-to-pointer` | Replaces the restatement with a link to the source | The reader can follow the link at the moment they need the fact | | `trim-to-citation` | Keeps a short quoted excerpt, attributed, and drops the rest | A specific span is worth quoting verbatim and the surrounding restatement is not | -| `condense-to-stamped-record` | Condenses to a four-part record: claim, basis URL, as-of date, recheck trigger | The surface must state the fact to function even when the source is unreachable | +| `condense-to-stamped-record` | Condenses to a stamped record: the surface's own decision in its own words, a pointer, an as-of date, a recheck trigger | The surface must act on the fact even when the source is unreachable | | `leave-with-reason` | Records why the passage stays | A carve-out applies, or a review veto fired, or the human decided | | `neutral-not-found` | Records that no source was identified | Budgets were exhausted; every searched surface is named | @@ -44,16 +44,22 @@ hot path is a stronger candidate for condensing than for pointing, because the f paid repeatedly. That is a disposition argument. It is never an allowance argument: "this is read often" does not make a copy acceptable, it makes a stamped record the right repair. -## The four-part record, when condensing +## The stamped record, when condensing -A stamped record carries all four parts or it is not one: +A stamped record stores no source text, quoted or paraphrased. It carries all of these parts or +it is not one: -1. **The claim**: what exactly is being asserted, narrow enough to check. -2. **The basis**: the specific URL, with anchor where one exists. "Verified" with no stated - basis is not re-checkable. -3. **The as-of date**: when the derivation happened. +1. **The decision**: what the surface does, in its own words, narrow enough to check. Where the + surface must act on a value offline, the value appears as the surface's own setting or firing + rule, never as a sentence describing what the source says. +2. **The pointer**: the specific URL, with the anchor of the exact section where one exists. A + record with no pointer is not re-checkable. +3. **The as-of date**: when the decision was last derived from the source. 4. **The recheck trigger**: the observable event that obliges re-deriving it. +The labels the marketplace's upstream-drift convention uses are `**Pointer**`, `**As of**` and +`**Recheck trigger**`, each on its own bullet under the decision. + A date alone is not a trigger. "Recheck periodically" is not a trigger. A trigger names an event someone could notice: a major version bump, a named page changing, a deprecation landing. If you cannot name one, that is a signal the passage wanted `convert-to-pointer` instead. A claim @@ -74,11 +80,12 @@ The dispositions below are what the report recommends, not what a run performs. | Disposition | Recommend it when | |---|---| | `convert-to-pointer` | The reader can follow a link at the moment they need the fact, and the surface does not have to work offline. | -| `condense-to-stamped-record` | The surface must state the fact to function without the source. The offline-load-bearing rule above still applies: it selects the stamped record over a bare pointer whatever the finding's tier. | +| `condense-to-stamped-record` | The surface must act on the fact without the source. The offline-load-bearing rule above still applies: it selects the stamped record over a bare pointer whatever the finding's tier. | | `leave-with-reason` | A carve-out applies (a conforming pointer or record, owned content, a distilling-file surface whose own attribution enumerates the fact, quoted and cited text), or the human decided. | -A recommended stamped record carries the four parts above: the claim, the basis URL, an ISO 8601 -as-of date `check-stamps.sh` can parse, and a trigger naming an event someone could notice. When +A recommended stamped record carries the parts above: the surface's decision in its own words, the +pointer, an ISO 8601 as-of date `check-stamps.sh` can parse, and a trigger naming an event someone +could notice. When no such trigger exists, recommend `convert-to-pointer`. ## Guards, all of which must pass before an edit is kept @@ -116,8 +123,8 @@ A pointer that was live at edit time can die later. That is a foreseen state wit repair, not a defect in the disposition. - A dead target demotes to a **stamped record**, if the fact is still knowable and still needed: - claim, the now-dead basis URL marked as such, the original as-of date, and a trigger naming - the recovery of a live basis. + the surface's decision, the now-dead pointer marked as such, the original as-of date, and a + trigger naming the recovery of a live source. - Where an archived snapshot of the original page exists, demote instead to an **archived-snapshot citation**, pointing at the archive and saying it is an archive. - Never silently re-expand the pointer back into a copy. The copy is what the fix removed, and diff --git a/plugins/attribution/skills/audit/reference/nomination.md b/plugins/attribution/skills/audit/reference/nomination.md index c48523b127..c31fef3603 100644 --- a/plugins/attribution/skills/audit/reference/nomination.md +++ b/plugins/attribution/skills/audit/reference/nomination.md @@ -195,7 +195,7 @@ well-headed file. Restated-fact rubric: one reads for who owns the fact and whether its owner could change it without this repository noticing; one reads for what a reader who acts on the passage does wrong if the owner has changed the fact; one reads for whether anything in the file records the fact's -basis, as-of date and recheck trigger, naming the part each nearby citation lacks. That third +pointer, as-of date and recheck trigger, naming the part each nearby citation lacks. That third stance is deliberately not "is a source cited": the rubric rejects that reading, because a link beside a stated value cites the value and records neither when it was checked nor what obliges a recheck. @@ -231,7 +231,7 @@ recheck. > span of text that decided it. A grade without a quoted span is not a grade. If the text you > would need to quote is not in front of you, grade it UNKNOWN and say what you would need. You > are not asked whether the fact is still true, or whether the passage's words match a source's: -> the rubric asks whether an external owner's fact is stated without a whole four-part record. +> the rubric asks whether an external owner's fact is stated without a whole stamped record. > > [lens sentence, when lens diversity is on] > @@ -289,7 +289,7 @@ One adversary per finding, in a fresh context, never one of the panel's judges. > [framing block above] > > A panel has unanimously judged that the passage below restates a fact an external source owns -> without a whole four-part record. Assume the panel is wrong and try to show it. Default to +> without a whole stamped record. Assume the panel is wrong and try to show it. Default to > refute: the finding survives only if you tried every attack below and each one failed against > text you can quote. You have the local passage, the containing file, the source text where one > was fetched, and the criterion grades with their quoted evidence. You do not have the judges' @@ -299,9 +299,9 @@ One adversary per finding, in a fresh context, never one of the panel's judges. > establishes it. (b) R1: the fact is general practice, this repository's own, or has no external > owner. (c) R2: the fact is stable across the owner's revisions, or is not concrete. (d) R3: the > passage is history, an example the text labels illustrative, or otherwise not asserted as -> current. (e) R4: a pointer that stands in place of stating the fact, or a whole four-part record -> whose claim names this very fact, is anywhere in the file; quote it and name the four parts. A -> link beside a stated value is not one. +> current. (e) R4: a pointer that stands in place of stating the fact, or a whole stamped record +> (pointer, as-of date, observable recheck trigger) that names this very fact, is anywhere in the +> file; quote it and name each part. A link beside a stated value is not one. > > Quote a span for every attack you rely on. Where the material leaves a question open, that is > a REFUTED: say what would settle it. diff --git a/plugins/attribution/skills/audit/reference/rubric.md b/plugins/attribution/skills/audit/reference/rubric.md index 95adaf2d8e..2be5285a26 100644 --- a/plugins/attribution/skills/audit/reference/rubric.md +++ b/plugins/attribution/skills/audit/reference/rubric.md @@ -59,9 +59,10 @@ drift is handled by its sync path, not by this audit. ### 2. Conforming stamped records -A passage carrying all four parts, claim, basis URL, as-of date, and recheck trigger, is already -the sanctioned fallback for a restatement that has to exist. It is not a copy to be found; it is -the end state a copy is converted into. +A passage carrying a whole stamped record, a pointer to the source, an as-of date, and a recheck +trigger beside the decision it records, is already the sanctioned end state a copy is converted +into. It is not a copy to be found. Text inside the record that matches the source is still judged +like any other passage. **Conforming is the whole test.** A dated sentence with no trigger is not carved out; it is a `rule-trigger-less-stamp` candidate where the repository has enabled that check, and a plain @@ -102,10 +103,10 @@ without the source in hand?" test is not enough to treat a sibling README as fir like any other upstream restatement (pointer, quote, or stamped record), or close it only when a breadcrumb or nomination already names the sibling repo as the source. -**Claim:** sibling-org repos are external at audit time unless the passage cites that sibling as its -source. **Basis:** `judgment`; no audited pass or upstream source shows the carve-out behavior yet. -**As of:** 2026-09-28. **Recheck trigger:** a consumer policy file declares same-org siblings -owned, or a golden case is added that turns on this boundary. +Sibling-org repos are external at audit time unless the passage cites that sibling as its source. +**Pointer:** none; this is `judgment`, and no audited pass or upstream source shows the carve-out +behavior yet. **As of:** 2026-09-28. **Recheck trigger:** a consumer policy file declares same-org +siblings owned, or a golden case is added that turns on this boundary. ### 5. Distilled-product architectures @@ -289,8 +290,8 @@ Two consequences that judges get wrong if they are not stated: Apply this section only when the dispatch names `restated-fact`. It decides whether a passage **restates a fact an external source owns**, in any wording, without conforming to the -upstream-drift shape: a pointer at the point of use, or a four-part record (claim, basis URL, -as-of date, observable recheck trigger). It never asks whether the passage's words correspond to +upstream-drift shape: a pointer at the point of use, or a whole stamped record (pointer, as-of +date, observable recheck trigger). It never asks whether the passage's words correspond to a source's words: a paraphrase, a summary, or a table restates a fact as fully as a copied sentence does. @@ -316,8 +317,8 @@ Categorical, as for the copy rubric: each names a class of surface, never a pass wanted kept. 1. **Conforming pointer or record.** The passage names the source in place of stating the fact, - or carries all four parts of a conforming record (see "A conforming record has four parts" - below). Conforming is the whole test. A link beside a stated value cites the value and does not + or carries every part of a conforming record (see "A conforming record's parts" below). + Conforming is the whole test. A link beside a stated value cites the value and does not record when it was checked or what obliges a recheck, and a dated sentence with no observable trigger is missing a part; neither is carved out. 2. **Owned content.** Facts this repository owns, in its own vocabulary. The direction test and @@ -364,16 +365,17 @@ to act on?* true at a named point and is not a claim about now, and neither is an example the text labels illustrative. -**R4-no-conforming-shape.** *Is the fact stated without a whole four-part record?* A record -elsewhere in the file covers the fact only where its claim names it. A whole record FAILS this +**R4-no-conforming-shape.** *Is the fact stated without a whole record (pointer, as-of date, +observable recheck trigger)?* A record elsewhere in the file covers the fact only where it names +the fact's topic. A whole record FAILS this criterion and clears the candidate. Quote the nearest citation or stamp and name the part it lacks; where there is none, say so. - **PASS, worked.** `The default is 30 seconds ([docs]()).` A basis, with no as-of date and no trigger. - **FAIL, worked.** `The default is 30 seconds. Verified against ; recheck when the - vendor changelog lists a change to the timeout.` All four parts, and the trigger is an event a - reader can check. + vendor changelog lists a change to the timeout.` A pointer, a date, and a trigger that is an + event a reader can check. ### Restated-fact verdict and tier @@ -384,29 +386,28 @@ lexical evidence, and unanimity does not manufacture any. A fetched source caps `source-fetched-similar` whatever the fingerprint showed. Every other tier is mapped from evidence by the table above, by fixed rule, never from a judge's confidence. -## Restated external rules, as four-part records +## Org rules this rubric applies, as stamped records -Each entry restates a rule this catalog does not own, because a judge applying the rubric offline -cannot follow a pointer. Each is source-pinned so the restatement can be re-derived. +Each entry is how this rubric applies a rule it does not own, stated here because a judge applying +the rubric offline cannot follow a pointer. Each is pinned so it can be re-derived. -**Prefer the pointer over the snapshot.** *Claim:* upstream bodies are read on demand; citing a -source and fetching it at read time is preferred over storing a snapshot of it, and a time-bound -external claim in durable content carries a recheck trigger. *Basis:* -`melodic-software/standards`, `conventions/engineering/documentation-and-citations.md`, as cited -by `docs/conventions/upstream-drift/README.md` "Boundary" in the marketplace repository. *As of:* -2026-08-28. *Recheck trigger:* any revision of that org standard, or of the upstream-drift +**Prefer the pointer over the snapshot.** This rubric treats a pointer read on demand as the end +state and a stored snapshot as a candidate, and expects a time-bound external claim in durable +content to carry a recheck trigger. *Pointer:* `melodic-software/standards`, +`conventions/engineering/documentation-and-citations.md`, as cited by +`docs/conventions/upstream-drift/README.md` "Boundary" in the marketplace repository. *As of:* +2026-10-01. *Recheck trigger:* any revision of that org standard, or of the upstream-drift convention's Boundary section that cites it. -**A conforming record has four parts.** *Claim:* a record deriving a fact from a source this -repository does not own carries the claim, the basis (a specific URL or probe), the as-of date, -and the recheck trigger, the observable event that obliges re-derivation. A date alone does not -qualify as a trigger. *Basis:* `docs/conventions/upstream-drift/README.md` "Required parts" and -"The observability bar" in the marketplace repository. *As of:* 2026-08-28. *Recheck trigger:* -any change to that convention's required parts, or the org standard broadening the accepted -trigger forms in a way this repository adopts. - -**A date is never authority.** *Claim:* a dated verification stamp records when a claim last -matched its source and confers no standing authority; a stale stamp reads identically to a fresh -one, so what obliges re-derivation is the trigger, not the date. *Basis:* -`docs/conventions/upstream-drift/README.md` "A date is never authority" in the marketplace -repository. *As of:* 2026-08-28. *Recheck trigger:* any change to that section. +**A conforming record's parts.** This rubric treats a record as conforming when it carries a +pointer to the source (a specific URL or probe), an as-of date, and a recheck trigger, the +observable event that obliges re-derivation, beside the decision it records. A date alone does not +qualify as a trigger. *Pointer:* `docs/conventions/upstream-drift/README.md` "Required parts" and +"The observability bar" in the marketplace repository. *As of:* 2026-10-01. *Recheck trigger:* any +change to that convention's required parts, or the org standard broadening the accepted trigger +forms in a way this repository adopts. + +**A date is never authority.** This rubric reads a dated stamp as the last time the record was +derived from its source, never as standing authority: what obliges re-derivation is the trigger, +not the date. *Pointer:* `docs/conventions/upstream-drift/README.md` "A date is never authority" in +the marketplace repository. *As of:* 2026-10-01. *Recheck trigger:* any change to that section. diff --git a/plugins/attribution/skills/audit/reference/source-fetch.md b/plugins/attribution/skills/audit/reference/source-fetch.md index 6790dcda2d..9647ccda60 100644 --- a/plugins/attribution/skills/audit/reference/source-fetch.md +++ b/plugins/attribution/skills/audit/reference/source-fetch.md @@ -28,20 +28,21 @@ subagent reads the corpus without seeing this file. route, and the marketplace repository is where the full argument, the measured incidents, and the issue links live. This plugin ships to consumers who do not have that repository, so a bare pointer cannot serve at run time. What follows is the operational subset, restated deliberately -and carried as a four-part record so the restatement stays honest. - -**Claim:** a candidate source is read through the raw-markdown channel first, checked for -wholeness and for page identity before its body is trusted, and an absence is assertable only -against a page whose identity was checked. **Basis:** -`docs/conventions/upstream-drift/README.md` "Reading the basis: the fetch route" in the -melodic-software/claude-code-plugins repository, which carries the measured incidents behind -each rule. **As of:** 2026-08-28. **Recheck trigger:** any change to that section, or a fetch +and carried as a stamped record so the restatement stays honest. + +This skill reads a candidate source through the raw-markdown channel first, checks it for +wholeness and for page identity before trusting its body, and asserts an absence only against a +page whose identity was checked. **Pointer:** `docs/conventions/upstream-drift/README.md` +"Reading the basis: the fetch route" in the melodic-software/claude-code-plugins repository, +which carries the measured incidents behind each rule. **As of:** 2026-10-01. **Recheck +trigger:** any change to that section, or a fetch in a live run that behaves in a way the rungs below do not describe: a new channel, a redirect where the doc says none occurs, or an identity check the doc's two tests do not settle. ## Three rules that bind every read -- **No verbatim quote, no claim.** A verdict about a source states the quoted span it matched. +- **No read, no verdict.** A verdict about a source rests on the text read and names the span it + matched in the run's report; the span never enters the file being repaired. This is not a formality here: a summarizer's paraphrase written down as page text turns the fingerprint comparison into a measurement of the summarizer rather than of the copy. The fingerprint module compares two concrete diff --git a/plugins/attribution/skills/audit/scripts/check-stamps.test.sh b/plugins/attribution/skills/audit/scripts/check-stamps.test.sh index c13d080965..49bc5e93dc 100755 --- a/plugins/attribution/skills/audit/scripts/check-stamps.test.sh +++ b/plugins/attribution/skills/audit/scripts/check-stamps.test.sh @@ -108,6 +108,15 @@ mkdir -p "$DIR" echo 'Nothing here says when to look again.' # 4 } >"$DIR/no-trigger.md" +{ + echo '# Pointer record' # 1 + echo '' # 2 + echo 'The pipeline passes strict mode to the runner.' # 3 + echo '- **Pointer**: for its modes, see https://example.com/docs#modes.' # 4 + echo '- **As of**: 2025-06-01' # 5 + echo '- **Recheck trigger**: the CLI ships a new major version.' # 6 +} >"$DIR/pointer-record.md" + run() { bash "$CHECK" --as-of "$AS_OF" "$@"; } # --- Usage ----------------------------------------------------------------------- @@ -335,17 +344,17 @@ assert_eq "no slack May line becomes a finding" \ # real stamp date this script does not parse, and says so in its own words. { - echo '# Year shapes' # 1 - echo '' # 2 - echo 'Dead code: variables set but never read (SC2034), and more.' # 3 - echo ' "cache_read_input_tokens": 2000' # 4 - echo 'STE-100 verified real (Issue 9, 2025, 53 rules/900 words).' # 5 - echo 'Read @/work/20260901T100000Z-handoff-widget.md, then go on.' # 6 - echo 'As of 2026-07 (official billing docs), surfaces varied.' # 7 - echo 'Checked as of 2024 and not revisited since.' # 8 - echo 'last_verified: 2026-01-01 against the vendor page.' # 9 - echo 'Read on 2026-01-02 from the vendor page.' # 10 - echo 'We _read_ the source on 2025-01-01 as a check.' # 11 + echo '# Year shapes' # 1 + echo '' # 2 + echo 'Dead code: variables set but never read (SC2034), and more.' # 3 + echo ' "cache_read_input_tokens": 2000' # 4 + echo 'STE-100 verified real (Issue 9, 2025, 53 rules/900 words).' # 5 + echo 'Read @/work/20260901T100000Z-handoff-widget.md, then go on.' # 6 + echo 'As of 2026-07 (official billing docs), surfaces varied.' # 7 + echo 'Checked as of 2024 and not revisited since.' # 8 + echo 'last_verified: 2026-01-01 against the vendor page.' # 9 + echo 'Read on 2026-01-02 from the vendor page.' # 10 + echo 'We _read_ the source on 2025-01-01 as a check.' # 11 } >"$DIR/year-shapes.md" OUT="$(run "$DIR/year-shapes.md" 2>/dev/null)" @@ -384,6 +393,12 @@ OUT="$(run --trigger-less "$DIR/with-trigger.md" 2>/dev/null)" assert_eq "a stated recheck trigger clears the surface" \ "$(echo "$OUT" | jq -r '[.findings[] | select(.rule | test("trigger-less"))] | length')" "0" +OUT="$(run --trigger-less "$DIR/pointer-record.md" 2>/dev/null)" +assert_eq "a pointer record's As of bullet is read as a stamp" \ + "$(echo "$OUT" | jq -r '.findings[] | select(.rule == "attribution/audit/rule-stamp-expired") | .line')" "5" +assert_eq "a pointer record's Recheck trigger bullet clears the surface" \ + "$(echo "$OUT" | jq -r '[.findings[] | select(.rule | test("trigger-less"))] | length')" "0" + # --- Config cascade -------------------------------------------------------------- mkdir -p "$CLAUDE_PROJECT_DIR/.claude" diff --git a/plugins/attribution/skills/audit/scripts/emit-findings.sh b/plugins/attribution/skills/audit/scripts/emit-findings.sh index 8fc1987da8..430e31a1a5 100755 --- a/plugins/attribution/skills/audit/scripts/emit-findings.sh +++ b/plugins/attribution/skills/audit/scripts/emit-findings.sh @@ -589,7 +589,7 @@ function rule_action(slug) { if (slug == "rule-trigger-less-stamp") return "Not auto-applicable: state the observable event that obliges re-derivation (upstream-drift required part 4)" if (slug == "rule-restated-upstream-fact") - return "Not auto-applicable: report-only, no fix pass reaches it; replace the restatement with a pointer at the point of use, or with a four-part record (claim, basis URL, as-of date, observable recheck trigger) when the surface must work offline" + return "Not auto-applicable: report-only, no fix pass reaches it; replace the restatement with a pointer at the point of use, or with a stamped record (the decision in its own words, a pointer, an as-of date, an observable recheck trigger) when the surface must work offline" return "Review by hand" } diff --git a/plugins/attribution/skills/audit/scripts/emit-findings.test.sh b/plugins/attribution/skills/audit/scripts/emit-findings.test.sh index 0098915db7..7568aea4f2 100755 --- a/plugins/attribution/skills/audit/scripts/emit-findings.test.sh +++ b/plugins/attribution/skills/audit/scripts/emit-findings.test.sh @@ -1665,7 +1665,7 @@ assert_match "the restated row ranks last" "$RS_ROW" '^\| 4 \|' assert_contains "the copy row keeps Confidence high" "$RS_COPY_ROW" "| high |" assert_contains "the copy row still names the fix flow" "$RS_COPY_ROW" '/attribution:audit fix' assert_contains "the restated remedy is a pointer" "$RS_ROW" "a pointer at the point of use" -assert_contains "or a four-part record" "$RS_ROW" "four-part record (claim, basis URL, as-of date, observable recheck trigger)" +assert_contains "or a stamped record" "$RS_ROW" "stamped record (the decision in its own words, a pointer, an as-of date, an observable recheck trigger)" assert_contains "and it says the row is report-only" "$RS_ROW" "Not auto-applicable: report-only" assert_not_contains "the restated remedy never names the fix flow" "$RS_ROW" '/attribution:audit fix' assert_not_contains "nor the sweep flow" "$RS_ROW" '/attribution:audit sweep' diff --git a/plugins/computer-use/.claude-plugin/plugin.json b/plugins/computer-use/.claude-plugin/plugin.json index 6f93683e89..23d8cbee43 100644 --- a/plugins/computer-use/.claude-plugin/plugin.json +++ b/plugins/computer-use/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "computer-use", - "version": "0.1.7", + "version": "0.1.8", "description": "Operating knowledge for Claude Code's built-in computer-use MCP server, the desktop screen-control surface. `/computer-use:diagnose` resolves a symptom to a cause instead of retrying: why every screenshot is downscaled to a fixed pixel budget and why zoom (not a bigger display) is the way back to detail, how to read a capture or input failure, and the per-OS quirks that make a synthesized key or menu behave unlike a human's. `/computer-use:setup` verifies the prerequisites the surface cannot verify for itself and reports the environment settings that end a session mid-run.", "author": { "name": "Melodic Software", diff --git a/plugins/computer-use/CHANGELOG.md b/plugins/computer-use/CHANGELOG.md index c0bc4a9e0c..aa89ce8d2d 100644 --- a/plugins/computer-use/CHANGELOG.md +++ b/plugins/computer-use/CHANGELOG.md @@ -3,6 +3,16 @@ All notable changes to the `computer-use` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.1.8] - 2026-10-02 + +### Changed + +- **`diagnose` states its screenshot downscale and zoom findings as our decisions with pointers.** + `reference/screenshots-and-zoom.md` no longer restates the documentation: the fixed target size, + the zoom action and the remedy of larger text in the app are our decisions, each with a pointer to + the docs section, an as-of date and a recheck trigger. The measured pixel-count figure is marked + as our own probe. + ## [0.1.7] - 2026-10-02 ### Fixed diff --git a/plugins/computer-use/skills/diagnose/reference/screenshots-and-zoom.md b/plugins/computer-use/skills/diagnose/reference/screenshots-and-zoom.md index ebada4cfc0..0af47edb8b 100644 --- a/plugins/computer-use/skills/diagnose/reference/screenshots-and-zoom.md +++ b/plugins/computer-use/skills/diagnose/reference/screenshots-and-zoom.md @@ -3,42 +3,46 @@ Why every screenshot arrives smaller than your screen, why that is not tunable, and the one mechanism that recovers detail. -**Recheck trigger:** re-verify the quotes and figures below if the `computer-use` CLI page's -downscaling section starts documenting a setting to change the target size, if the -`computer-use-tool` platform page changes its "full resolution" zoom wording or its -implementation-best-practices resolution guidance, or if the linked best-practices blog post -revises its recommended resolutions or its "single highest impact optimization" claim. - ## Downscaling is automatic and has no setting -Claude Code downscales **every** screenshot before it reaches the model. Official wording -(verified 2026-08-10, [computer use from the CLI](https://code.claude.com/docs/en/computer-use)): +We treat the downscale of **every** screenshot as fixed: no setting changes the target size, so +when a screenshot leaves text or buttons unreadably small, the remedy is larger text in the app, +never a lower display resolution. -> There is no setting to change the target size. If on-screen text or controls are too small for -> Claude to read after downscaling, increase their size in the app rather than changing your -> display resolution. +- **Pointer**: for the downscale and the absence of a size setting, see + [Screenshots are downscaled automatically](https://code.claude.com/docs/en/computer-use#screenshots-are-downscaled-automatically). +- **As of**: 2026-08-10 +- **Recheck trigger**: that section documents a setting that changes the target size. ## The target is a pixel budget, not a scale factor -This is the part that surprises people. Two very different displays converge on the same -megapixel count: +This is the part that surprises people. We read the target as a fixed pixel count, about 1.2 +megapixels, with the aspect ratio preserved, rather than a fixed scale factor: | Source display | Delivered image | Megapixels | Linear scale | |---|---|---|---| | 2560x1440 (Windows, measured 2026-08-10) | 1456x816 | 1.19 | 1.76x | -| 3456x2234 (upstream's MacBook example) | 1372x887 | 1.22 | 2.52x | -Same budget, different ratios, aspect ratio preserved. The practical consequence: **a smaller -monitor does not buy a sharper screenshot**. It buys the same ~1.2MP with less on it. That is -occasionally worth doing for a dense UI, but it is a trade of coverage for density, never a -quality win. +The worked example in the downscaling section above, on a larger display, is the second data point +we compared against. The practical consequence: **a smaller monitor does not buy a sharper +screenshot**. It buys the same ~1.2MP with less on it. That is occasionally worth doing for a dense +UI, but it is a trade of coverage for density, never a quality win. + +- **Pointer**: the measurement above is our own probe; for the upstream worked example, see + [Screenshots are downscaled automatically](https://code.claude.com/docs/en/computer-use#screenshots-are-downscaled-automatically). +- **As of**: 2026-08-10 +- **Recheck trigger**: a capture on a measured display delivers a pixel count far from 1.2MP, or + that section's worked example changes. ## `zoom` re-captures at full resolution -`zoom` is not a crop of the downscaled image. Upstream calls it "view a specific region of the -screen **at full resolution**" ([computer use -tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool), verified -2026-08-10). +We treat `zoom` as a fresh capture of the region at full resolution, not a crop of the downscaled +image. + +- **Pointer**: for the `zoom` action, see + [Available actions](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool#available-actions). +- **As of**: 2026-08-10 +- **Recheck trigger**: the `zoom` row in that table stops describing a full-resolution capture. Local behavior matches: while capture is failing, `zoom` returns `Screenshot capture failed after 3 attempts` rather than a blurry crop. A crop of the @@ -55,31 +59,39 @@ one. Zoom is read-only inspection. ## Order of remedies for "Claude can't read this" 1. **`zoom` the region.** Free, immediate, no environment change. -2. **Increase the size in the app**: editor font size, browser zoom, app scaling. This is - upstream's own recommendation and it survives across screenshots. -3. **Keyboard instead of mouse** for genuinely tiny targets (tray icons, small checkboxes). - Upstream recommends this over trying to click them. +2. **Increase the size in the app**: editor font size, browser zoom, app scaling. It survives + across screenshots. +3. **Keyboard instead of mouse** for genuinely tiny targets (tray icons, small checkboxes), rather + than trying to click them. 4. **Do not lower display resolution.** Claude Code already downscales; dropping the source only removes information earlier. -## Upstream's own resolution guidance, and how it applies here - -The API-side computer use tool exposes display dimensions the caller chooses, and there the -guidance is concrete. The platform docs' implementation-best-practices section gives the -resolutions: 1024x768 or 1280x720 for general desktop work, and nothing above 1920x1080. It names -"resolution too low" as the cause of consistently poor accuracy -([computer use tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool), -verified 2026-08-10). The benchmarked blog post adds that pre-downscaling before sending is "the -single highest impact optimization" and that native unscaled resolution is the primary cause of -poor accuracy -([best practices](https://claude.com/blog/best-practices-for-computer-and-browser-use-with-claude), -verified 2026-08-10). - -**That knob does not exist on the Claude Code surface.** The harness owns the downscale and -already does the recommended thing. The guidance is still worth knowing because it explains -*why* the harness behaves this way, and because it tells you that native unscaled resolution is -the documented primary cause of poor click accuracy. Do not translate the API advice into a -display-settings change on a Claude Code machine. +- **Pointer**: for the in-app size remedy, see + [Screenshots are downscaled automatically](https://code.claude.com/docs/en/computer-use#screenshots-are-downscaled-automatically); + for keyboard use on hard targets, see + [Optimize model performance with prompting](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool#optimize-model-performance-with-prompting). +- **As of**: 2026-08-10 +- **Recheck trigger**: either section drops or reverses its remedy. + +## The API-side resolution guidance, and how it applies here + +The API-side computer use tool takes display dimensions the caller chooses, and its docs give +concrete resolution guidance and name low resolution as a cause of poor accuracy. Read it at the +pointer; we do not restate it. + +**That knob does not exist on the Claude Code surface.** The harness owns the downscale, and we +read its behavior as already following that guidance. The guidance is still worth knowing because +it explains *why* the harness downscales and why resolution affects click accuracy. Do not +translate the API advice into a display-settings change on a Claude Code machine. + +- **Pointer**: for the resolution guidance, see + [Size screenshots to fit image limits](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool#handle-coordinate-scaling-for-higher-resolutions) + and + [Diagnose click issues](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool#diagnose-click-issues) + (correlate with [best practices](https://claude.com/blog/best-practices-for-computer-and-browser-use-with-claude)). +- **As of**: 2026-08-10 +- **Recheck trigger**: either section changes its recommended resolutions, or the Claude Code + computer-use page documents a setting to change the target size. ## `save_to_disk` is not an escape hatch diff --git a/plugins/context-guard/.claude-plugin/plugin.json b/plugins/context-guard/.claude-plugin/plugin.json index 6dd7dbbcd8..0b4a7dbfe8 100644 --- a/plugins/context-guard/.claude-plugin/plugin.json +++ b/plugins/context-guard/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "context-guard", - "version": "0.8.5", + "version": "0.8.6", "description": "Per-session context-window observability plus the first shipped consumer: a statusline wrapper tees each session's context_window fields to a per-session snapshot file, a zone resolver classifies usage into smart/acceptable/dumb bands (percentage bands plus window-class token bands, conservative-min combination, zones.json SSOT with shipped defaults), a reader contract fixes how consuming sessions interpret the snapshots, and zone-crossing hooks report once per transition into a worse zone across two channels: the continuation menu to the operator, who owns that choice, and to the model only the zone determination plus the counter-steer that a zone word is not a decay signal (advisory by default; an optional blocking mode gates new mutating work on a fresh dumb-zone snapshot with handoff-writing exempt), with a PostCompact hook persisting an evidence-degraded marker.", "author": { "name": "Melodic Software", diff --git a/plugins/context-guard/CHANGELOG.md b/plugins/context-guard/CHANGELOG.md index a4513548ae..0402c13b92 100644 --- a/plugins/context-guard/CHANGELOG.md +++ b/plugins/context-guard/CHANGELOG.md @@ -5,6 +5,22 @@ All notable changes to the `context-guard` plugin. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [0.8.6] - 2026-10-02 + +### Changed + +- **`reader-contract.md` records its upstream dependencies as decisions plus pointers.** The + statusline `context_window` fields, the percentage shape, the version field and the + absolute-token degradation basis each state what the contract relies on, with a pointer to the + docs section or the Chroma context-rot report, an as-of date and a recheck trigger, and carry + none of the source wording. One recheck trigger covers every dated record in the file. No + behavior or zone change. +- The reader contract's folklore-number record no longer rests on the Opus 5 guide. It points at + the current models' context windows, with a trigger on those windows or the practitioner figure + changing. +- The 2.1.132 token-field floor now points at the changelog entry that names the fix, instead of + saying no upstream source names that version. + ## [0.8.5] - 2026-10-02 ### Fixed diff --git a/plugins/context-guard/reference/reader-contract.md b/plugins/context-guard/reference/reader-contract.md index 32c6e2cc8a..b9ce3c1ad6 100644 --- a/plugins/context-guard/reference/reader-contract.md +++ b/plugins/context-guard/reference/reader-contract.md @@ -28,9 +28,10 @@ staleness value, and the default zone bands. Inlined copies in consumers must st **byte-identical** to the values printed here; a consumer lane carries a drift check that grep-matches its inlined values against this file. -**Recheck trigger for every dated stamp in this file:** re-read the cited page and re-date the -stamp when any of these change. The statusline stdin schema, meaning the `context_window` field -names, the `used_percentage` formula, and the top-level `version` field. The auto-compact trigger, +**Recheck trigger for every dated record in this file:** re-read the record's pointer, re-derive +the decision, and re-date the record when any of these change. The statusline stdin schema, +meaning the `context_window` field names, the `used_percentage` formula, and the top-level +`version` field. The auto-compact trigger, meaning whether a default threshold is published as a number, and which models and environments compact before the model's context limit. The four surfaces in the tunable table below (`autoCompactWindow`, `CLAUDE_CODE_AUTO_COMPACT_WINDOW`, `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE`, @@ -82,16 +83,16 @@ concurrent sessions each own the file named by their `session_id`. "session_id": "abc123", "cli_version": "2.1.218", "context_window": { - "total_input_tokens": 15500, - "total_output_tokens": 1200, + "total_input_tokens": 42000, + "total_output_tokens": 3100, "context_window_size": 200000, - "used_percentage": 8, - "remaining_percentage": 92, + "used_percentage": 21, + "remaining_percentage": 79, "current_usage": { - "input_tokens": 8500, - "output_tokens": 1200, - "cache_creation_input_tokens": 5000, - "cache_read_input_tokens": 2000 + "input_tokens": 6000, + "output_tokens": 3100, + "cache_creation_input_tokens": 9000, + "cache_read_input_tokens": 27000 } } } @@ -106,12 +107,14 @@ concurrent sessions each own the file named by their `session_id`. - `cli_version`: the statusline payload's top-level `version` (the Claude Code version), copied only when it is a string; absent otherwise, never guessed. It gates the token shape (see "Version floor"), so an absent one is not a defect. It just leaves the percentage shape standing alone. -- `context_window`: copied **verbatim** from the statusline stdin schema - (, verified 2026-08-10), so upstream field additions +- `context_window`: copied **verbatim** from the statusline payload, so upstream field additions flow through without a plugin change. The key is absent when the session's statusline payload - carried none. Null states are upstream-documented and normal: `used_percentage` / - `remaining_percentage` may be `null` early in a session; `current_usage` is `null` before the - first API call **and again immediately after `/compact`** until the next response repopulates it. + carried none. A null `used_percentage`, `remaining_percentage`, or `current_usage` is a normal + state, not a defect; the capability table below says what each one does to the zone, and a null + `current_usage` after `/compact` is why that row resolves `unknown`. Pointer: for the field set + and when each field is null, see . + As of: 2026-08-10. Recheck trigger: a release note or that section changes the `context_window` + fields or the states in which they are null. - Treat all values as **untrusted data**: parse with a JSON parser; validate any value against its documented format before handing it to a lenient parser (the bundled resolver format-gates `captured_at` to strict ISO-8601 before date parsing, and requires the embedded `session_id` to @@ -154,16 +157,20 @@ always means "take the conservative route". The contract carries two zone shapes because the two underlying measures answer different questions. Never equate them without normalizing: -- **Percentage shape**: `context_window.used_percentage` against the percentage bands. Upstream - computes it from **input tokens only** (`input_tokens + cache_creation_input_tokens + - cache_read_input_tokens`, no output, per the statusline doc, verified 2026-07-26). It answers - *distance to compaction*, because compaction thresholds key off the same accounting. +- **Percentage shape**: `context_window.used_percentage` against the percentage bands. We rely on + it counting the input side only, and read it as *distance to compaction*, because compaction + thresholds key off the same accounting. Pointer: for how the percentage is computed, see + . As of: 2026-07-26. Recheck + trigger: that section changes the `used_percentage` formula. - **Token shape**: **occupancy**, defined as `total_input_tokens + total_output_tokens`, against the window-class token bands. Occupancy counts both directions because both occupy the window, - and the degradation evidence (Chroma context-rot report) tracks **absolute tokens in context, - not window fraction**. It answers *distance to quality loss*. That is also why the token bands - are absolute numbers selected by window class rather than percentages: 50% of a 1M window is a - materially different cognitive state than 50% of a 200k window. + and we treat quality loss as tracking **absolute tokens in context, not window fraction**. It + answers *distance to quality loss*. That is also why the token bands are absolute numbers + selected by window class rather than percentages: 50% of a 1M window is a materially different + cognitive state than 50% of a 200k window. Pointer: for the degradation evidence, see the Chroma + context-rot report, . As of: 2026-10-01. Recheck + trigger: Chroma revises or withdraws the report, or a newer study finds degradation tracking + window fraction. **Window-class selection:** use the band row whose class key is the **largest one ≤ `context_window_size`**. A window smaller than every configured class has no row, so the token @@ -181,19 +188,20 @@ misfire the token bands badly. Cumulative semantics are **not observable from th cumulative 170k in a 200k window is a perfectly plausible current occupancy, sits inside the window, and resolves `dumb` while the live context may be smart-zone. So the token shape requires an explicit version signal: the snapshot's `cli_version`, which the tee copies from the -statusline payload's top-level `version` field (the Claude Code version, statusline doc, verified -2026-08-10). **The token shape is computable only when `cli_version` is present, purely numeric -dotted, and ≥ 2.1.132**; absent, malformed, or older leaves the percentage shape to stand alone. - -> **Sourcing status of the 2.1.132 floor.** Claim: `total_input_tokens` / `total_output_tokens` -> mean current occupancy only from Claude Code 2.1.132. Basis: no current upstream source. The -> statusline page (`https://code.claude.com/docs/en/statusline.md`, complete raw page, re-checked -> 2026-08-10) states only the present-tense semantics this floor depends on: "Token counts -> currently in the context window, from the most recent API response" and "**Combined totals** -> (`total_input_tokens`, `total_output_tokens`): tokens currently in the context window". The -> floor is therefore a retained claim, a conservative lower bound kept deliberately: dropping it -> can only *widen* which payloads the token shape trusts, and the failure it guards is silent. -> Recheck trigger: re-source it before any change that relaxes it. +statusline payload's top-level `version` field, the Claude Code version (Pointer: +. As of: 2026-08-10. Recheck trigger: +that section renames or drops the `version` field). **The token shape is computable only when +`cli_version` is present, purely numeric dotted, and ≥ 2.1.132**; absent, malformed, or older +leaves the percentage shape to stand alone. + +> **Source of the 2.1.132 floor.** We do not trust `total_input_tokens` / +> `total_output_tokens` as current occupancy below Claude Code 2.1.132, the release whose +> changelog entry names the statusline token-count fix. +> +> - **Pointer**: [Changelog 2.1.132](https://code.claude.com/docs/en/changelog#2-1-132); for the +> fields' present meaning, . +> - **As of**: 2026-10-01 +> - **Recheck trigger**: any change that relaxes the floor, or the changelog entry moves or is reworded. **Plausibility guard (independent, retained):** **occupancy greater than `context_window_size` also marks the token shape not-computable**. That is corrupt or forged data, and it catches what @@ -201,11 +209,11 @@ a version field cannot (there is no writer authentication, so `cli_version` is u every other snapshot value). The bundled resolver implements both gates. **Band provenance:** all shipped band numbers are **declared judgment defaults with named -anchors**, not benchmark-derived constants. The 1M row's anchor is a named-staff informal range -(self-hedged "highly task-dependent"); the 200k row is declared judgment near practitioner -folklore values, but deliberately below them. Both rows carry equally low confidence; `zones.json` is the correction path, and the numeric agreement of -the 200k row's percentage translation with the shipped 50/75 percentage defaults is coincidence, -not validation. +anchors**, not benchmark-derived constants. The 1M row's anchor is an informal range a named staff +member gave and hedged as task-dependent; the 200k row is declared judgment near practitioner +folklore values, but deliberately below them. Both rows carry equally low confidence; +`zones.json` is the correction path, and the numeric agreement of the 200k row's percentage +translation with the shipped 50/75 percentage defaults is coincidence, not validation. ## Zone-crossing hooks (first shipped consumer) @@ -313,31 +321,37 @@ with a declared margin: if compaction triggers at 90% or above, the dumb band le points or more. The trigger is **model- and environment-dependent**, so no single band set is correct everywhere; `zones.json` is the correction path if compaction is ever observed earlier. -Two adjacent caveats, same fetch: the doc warns the statusline percentage "may differ from -`/context` output due to when each is calculated", so the value is as-of the last API response, not -the next request; and with `autoCompactEnabled: false` no compaction ever fires (the session -hard-stops at the window instead), which makes the dumb band the *only* tripwire, so it matters -strictly more, never less. +Two adjacent decisions. We read the statusline percentage as of the last API response, not the +next request, so it can trail `/context`. With auto-compact turned off, the dumb band is the +*only* tripwire, so it matters strictly more, never less. + +- **Pointer**: for how the statusline percentage relates to `/context`, see + ; for turning auto-compact off, see + . +- **As of**: 2026-09-28 +- **Recheck trigger**: a release note or either section changes when the percentage is computed or + what turning auto-compact off does. ### The trigger has no documented threshold, but it is operator-tunable No *default* threshold is published as a number (above), yet the point at which auto-compact fires -is a configured value the operator can read and set. **Four** surfaces govern it. Verified -2026-08-17 against two independent pools, the official -[settings reference](https://code.claude.com/docs/en/settings) and the shipped binary's own schema -strings (v2.1.233), then re-verified 2026-08-19 against the live settings, -[env-vars](https://code.claude.com/docs/en/env-vars), and -[model-config](https://code.claude.com/docs/en/model-config) pages: - -| Surface | Kind | What it does | -|---|---|---| -| `autoCompactWindow` | `settings.json` key | How full the window gets before auto-compact fires, **in tokens, `100000` to `1000000`** (binary schema: `.int().min(1e5).max(1e6).optional()`). **No numeric default**: unset means a window tuned for the model, deliberately not published as a number. Written by the `/autocompact` command; the `--autocompact` flag sets it for one launch and, unlike the command, is not preempted by a higher-priority settings scope. | -| `CLAUDE_CODE_AUTO_COMPACT_WINDOW` | environment variable | Same units and range; **highest precedence**: it overrides the command, the flag, and the setting while set. **Accepts a plain integer only**: the command and flag take `500k` / `1M` / a bare `500` meaning thousands, but the variable reads `500k` as `500` and clamps to the 100K minimum. | -| `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE` | environment variable | Sets the **percentage (1–100) of the auto-compact window** at which compaction triggers. **Can only lower the threshold**: values above the default percentage are ignored. Applies only in sessions that compact *before* the model's context limit, and to subagents as well as the main conversation. | -| `autoCompactEnabled` / `DISABLE_AUTO_COMPACT` | `settings.json` key (default `true`, shown in `/config` as **Auto-compact**) / environment variable | Turns auto-compact off entirely. (`DISABLE_COMPACT`, which disables *all* compaction including `/compact`, comes from the 2026-08-17 binary-strings pool; it is not listed on the env-vars page as of 2026-08-19, so treat it as unconfirmed by docs.) | - -Claude Code caps the window at the model's actual context window, so a configured value above it -does not extend anything. +is a configured value the operator can read and set. **Four** surfaces govern it. Each row states +what this plugin relies on; the pointer holds the units, ranges, forms, and precedence. + +| Surface | Kind | What this plugin relies on | Pointer | +|---|---|---|---| +| `autoCompactWindow` | `settings.json` key | A token count that moves the trigger. Unset gives no number we can read, so we never assume one. Normalize it into the percentage shape before comparing (below). | [settings-reference: `autoCompactWindow`](https://code.claude.com/docs/en/settings-reference#autocompactwindow); [model-config: Set the auto-compact window](https://code.claude.com/docs/en/model-config#set-the-auto-compact-window) | +| `CLAUDE_CODE_AUTO_COMPACT_WINDOW` | environment variable | Read as the effective window whenever it is set, ahead of the setting, the command, and the flag. | [env-vars: Variables](https://code.claude.com/docs/en/env-vars#variables) | +| `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE` | environment variable | Read as able only to move the trigger earlier, never later. | [env-vars: Variables](https://code.claude.com/docs/en/env-vars#variables) | +| `autoCompactEnabled` / `DISABLE_AUTO_COMPACT` | `settings.json` key / environment variable | Either one turning auto-compact off leaves the dumb band as the only tripwire. We treated `DISABLE_COMPACT` as unconfirmed by docs: it came from our 2026-08-17 probe of the shipped binary's strings (v2.1.233) and was absent from the env-vars page on 2026-08-19. | [settings-reference: `autoCompactEnabled`](https://code.claude.com/docs/en/settings-reference#autocompactenabled); [env-vars: Variables](https://code.claude.com/docs/en/env-vars#variables) | + +We read a configured window above the model's context window as the model's window: it extends +nothing. + +- **Pointer**: per row above. +- **As of**: 2026-08-19 +- **Recheck trigger**: a release note or one of those sections changes a surface's units, range, + precedence, or the set of surfaces itself. **Normalize before comparing: the trigger is not in occupancy.** The two zone shapes answer different questions and must never be equated (see "Occupancy and combination rule"), and the @@ -347,10 +361,12 @@ is input-token-based and answers *distance to compaction*, while the token bands A configured window is a fill threshold, so compare it against the percentage shape and let the occupancy bands move independently. -One consequence matters enough to state on its own, and it is the docs' own warning -(env-vars, verified 2026-08-19): **`used_percentage` always measures against the model's full -context window**, so once the auto-compact window is lowered, *the percentage no longer indicates -when compaction will run*. A consumer reading only the percentage will not see the trigger coming. +One consequence matters enough to state on its own: we read `used_percentage` as a share of the +whole window the model offers, so after the auto-compact window is lowered, **we never read the +percentage as a forecast of when compaction will run**. A consumer reading only the percentage +will not see the trigger coming. Pointer: for the `CLAUDE_CODE_AUTO_COMPACT_WINDOW` row, see +. As of: 2026-08-19. Recheck trigger: that row +changes what the percentage is measured against. **Tune bands below the effective trigger, never above it.** Whatever the trigger resolves to on a machine, the `dumb` band should be reached first. A zone reading exists so the session arrives at a @@ -373,25 +389,28 @@ is instrumentation, not prohibition: observable zones, then advisory injection, blocking gate with a grace budget, with auto-compact remaining the last-resort safety net beneath all of it (as-of 2026-08-17). -**On folklore numbers.** The vendored Boris playbook, §64, attributing the compromise to Thariq, -is a widely-cited practitioner anchor. It reports context rot setting in around 300–400k tokens on -1M-context models and suggests `CLAUDE_CODE_AUTO_COMPACT_WINDOW=400000`. Recorded here as a **named -anchor, never an adopted number**, and it comes with its own amendment: that calibration is -Opus 4.7-era, and the Opus 5 prompting guide (verified 2026-08-08) states the 1M window's -instruction following, tool calling, and reasoning "stay consistent throughout the window", which -removes the degradation premise for that specific figure. A lowered window remains a legitimate -cost and compaction-timing choice on its own terms. +**On folklore numbers.** The auto-compact window figure in the vendored Boris playbook, §64, is a +widely-cited practitioner anchor. We record it as a **named anchor, never an adopted number**: its +calibration predates the 1M-window models these sessions run on, and we do not treat its +context-degradation premise as holding on them. A lowered window remains a legitimate cost and +compaction-timing choice on its own terms. + +- **Pointer**: for the practitioner figure, see `/playbooks:boris` §64 ("Lower Your Auto-Compact + Threshold"); for the current models' context windows, see + [Latest models comparison](https://platform.claude.com/docs/en/about-claude/models/overview#latest-models-comparison). +- **As of**: 2026-10-01 +- **Recheck trigger**: the session models' context window changes, or §64's figure changes. ## Prompt-cache miss cause The statusline payload's `prompt_cache.last_miss_cause` names why the last cache miss happened. `plugins/context-guard/scripts/prompt-cache-cause.py` reads that object from a statusline JSON payload and prints the cause names. The tee snapshot still copies `context_window` and does not -copy `prompt_cache`; pass the live payload to the script. Claim: `last_miss_cause.causes` holds -names such as `tools_changed`, `system_prompt_changed`, `ttl_expired_5m`, and -`likely_server_side`, and the object is null when no cause was identified. Basis: -. As of: 2026-09-28. Recheck: that -section renames the object or its cause names. +copy `prompt_cache`; pass the live payload to the script. The script prints whatever cause names +`last_miss_cause.causes` carries, with no fixed list of its own, and prints `null` when the object +is null. Pointer: for the object and its cause names, see +. As of: 2026-09-28. Recheck trigger: +that section renames the object or its cause names. ## Zones (machine-scope tuning, optional) @@ -440,8 +459,9 @@ bash "/scripts/context-zone.sh" # prints one zone wo ## Session-id discovery (how a consumer learns its own id) A skill learns its session id via the **`${CLAUDE_SESSION_ID}` substitution** in skill markdown -content (, substitution table, verified 2026-08-10). The -skill body interpolates it into the snapshot path directly. +content. The skill body interpolates it into the snapshot path directly. Pointer: +. As of: 2026-08-10. +Recheck trigger: that table renames or drops `${CLAUDE_SESSION_ID}`. **Fallback:** when the substitution is unavailable (older Claude Code, non-skill context, or the literal string `${CLAUDE_SESSION_ID}` survives unexpanded), the consumer must not guess a session @@ -465,9 +485,8 @@ written into the session's own user settings was never invoked. This is not a degraded install and not a missing dependency. It is the absence of the only documented surface that **delivers per-session context-window occupancy to a local writer**: as of -**2026-08-21**, hook stdin carries no context, token, usage, or window field on any event, except -`PostToolUse` on the `Agent` tool, whose `tool_response` carries `totalTokens` and a `usage` -breakdown for the *subagent's* final API request and nothing about the main session's window. Two +**2026-08-21**, our channel inventory found no hook event whose stdin reports the main session's +window; the one token-bearing hook payload describes a *subagent's* request. Two other channels do carry live occupancy for the running session, the OpenTelemetry `claude_code.api_request` log event and the session transcript, and neither can be turned into a snapshot; `reference/cloud-headless-capture.md` records why in full. That file is the writer-side @@ -485,8 +504,8 @@ take the conservative route. What changes is how a consumer *reports* it: `claude_code.api_request` event carries live per-session token counts but no window size and no local sink, so a zone from it needs a fabricated denominator; the session transcript reachable through the documented `transcript_path` hook field carries the right numbers behind an entry - format its own docs call internal and version-unstable, where a field can keep its name and stop - meaning full-context occupancy with nothing to detect it; and the OpenTelemetry token metric is + format we do not treat as a stable interface, where a field can keep its name and stop meaning + full-context occupancy with nothing to detect it; and the OpenTelemetry token metric is a cumulative counter, not occupancy. A wrong zone is strictly worse than `unknown`: `unknown` routes to the conservative path, while a misread occupancy can read `smart` on a nearly full window. @@ -504,11 +523,10 @@ and managed settings, where `statusLine` is also a valid key. - **No `statusLine` in any scope** is structural, and offering statusline wiring as the remediation is wrong in an environment that runs no statusline. - **A `statusLine` configured but the status line disabled** is also structural, and the - remediation is policy or trust rather than wiring. Claude Code turns the status line off entirely - when managed settings set `disableAllHooks` or the folder is not trusted, and narrows the source - to managed settings when `allowManagedHooksOnly` is set. Under narrowing it runs a managed value - if one is deployed and otherwise skips yours *without warning*. This state looks exactly like a - broken install unless it is checked first. The dated record for both settings keys is + remediation is policy or trust rather than wiring. Check `disableAllHooks`, + `allowManagedHooksOnly`, and folder trust before anything else: either key, or an untrusted + folder, can disable or narrow the status line with no warning, so this state looks exactly like + a broken install unless it is checked first. The dated record for both settings keys is `cloud-headless-capture.md`, branch 3 of "Distinguishing structural absence from breakage". - **A `statusLine` configured, not disabled, in an environment that does not run a statusline** (cloud, headless `claude -p`, other terminal-less) is also structural: the command exists, is diff --git a/plugins/discovery/.claude-plugin/plugin.json b/plugins/discovery/.claude-plugin/plugin.json index 483c45563a..5ce3d73fba 100644 --- a/plugins/discovery/.claude-plugin/plugin.json +++ b/plugins/discovery/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "discovery", - "version": "0.25.24", + "version": "0.25.25", "description": "Structured discovery before changes: explore the local codebase, run disciplined multi-source external research, and reconstruct why a past decision was made from evidence outside the code. Each dispatches a purpose-built subagent by default so the reading stays out of the main conversation, with source tiers, falsification, recency gates, an intent-evidence tier, and a corpus-coverage ledger, and each persists EXPLORE.md / RESEARCH.md / INTENT.md index-plus-sidecar handoff artifacts.", "author": { "name": "Melodic Software", diff --git a/plugins/discovery/CHANGELOG.md b/plugins/discovery/CHANGELOG.md index f2dbe6994c..0782ca4429 100644 --- a/plugins/discovery/CHANGELOG.md +++ b/plugins/discovery/CHANGELOG.md @@ -1,5 +1,16 @@ # Changelog: discovery plugin +## [0.25.25] - 2026-10-02 + +### Changed + +- **`explorer` pins `effort: medium`**, the marketplace effort floor for code-changing or verifying + work, and gains a finish-then-stop paragraph: the run ends with the persisted EXPLORE.md set and + the bounded return payload. +- **`reference/parent-contract.md` states the new pin** and converts its harness-facts, credential + and permission-grant records to the links-only shape: our decision, a pointer to the exact + section, an as-of date and a recheck trigger. + ## [0.25.24] - 2026-10-02 ### Fixed diff --git a/plugins/discovery/agents/explorer.md b/plugins/discovery/agents/explorer.md index 89ef9aa4a9..d9f9cbbc3e 100644 --- a/plugins/discovery/agents/explorer.md +++ b/plugins/discovery/agents/explorer.md @@ -5,7 +5,7 @@ tools: "Read, Grep, Glob, Bash, Write, Skill, Agent" skills: - discovery:explore model: sonnet -effort: high +effort: medium maxTurns: 40 --- You are the discovery explorer: a fresh-context worker a main session dispatches so that the volume @@ -76,24 +76,26 @@ scope, testing conventions when the scope involves tests. Skip any that do not e path. Skipping this is what makes an otherwise-thorough exploration convention-blind, and convention-blind findings are how a downstream edit lands against the project's declared direction. -The dated record for that harness behavior: - -- **Claim.** A non-fork subagent inherits none of its parent's on-demand instruction surfaces. It - receives a path-scoped `.claude/rules/` file, or a nested `CLAUDE.md` and the `AGENTS.md` that - shim imports, only when it reads a path the surface covers, and the glob is matched against the - requested path, so even a read that finds no file fires it. -- **Basis.** First-party probe run inside a dispatched general-purpose subagent on the harness - `claude --version` reports as `2.1.268 (Claude Code)`, observing the `Contents of :` block - appended to `Read` tool results. -- **As of.** 2026-09-13. -- **Recheck trigger.** The consuming repository's Claude Code minor version moves past 2.1.268, or - a release note names subagent context inheritance, memory loading, or path-scoped rule - triggering, or a read of a covered path inside a subagent injects nothing. +The dated record for that harness behavior. We rely on a non-fork subagent inheriting none of its +parent's on-demand instruction surfaces: it receives a path-scoped `.claude/rules/` file, or a +nested `CLAUDE.md` and the `AGENTS.md` that shim imports, only when it reads a path the surface +covers, and the glob is matched against the requested path, so even a read that finds no file +fires it. + +- **Pointer**: for that behavior, see our own probe, recorded in pull request + [#4157](https://github.com/melodic-software/claude-code-plugins/pull/4157) ("The confirmed + harness claim"): run inside a dispatched general-purpose subagent on the harness + `claude --version` reports as `2.1.268 (Claude Code)`, it observed the `Contents of :` + block appended to `Read` tool results. +- **As of**: 2026-09-13 +- **Recheck trigger**: the consuming repository's Claude Code minor version moves past 2.1.268, a + release note names subagent context inheritance, memory loading, or path-scoped rule triggering, + or a read of a covered path inside a subagent injects nothing. ## Preload liveness: the first thing you do -A `skills:` entry that fails to resolve is skipped **silently**: Claude Code logs a warning to the -debug log and starts you anyway. An undisciplined run that still writes an artifact is +We treat a `skills:` entry that fails to resolve as skipped **silently**: you start anyway, without +the body, and nothing in your context says so. An undisciplined run that still writes an artifact is indistinguishable from a good one by every other signal, which is exactly the failure the token exists to prevent. The dated record for that harness behavior is [`${CLAUDE_PLUGIN_ROOT}/reference/parent-contract.md`](${CLAUDE_PLUGIN_ROOT}/reference/parent-contract.md), @@ -161,8 +163,8 @@ allowing nested spawning at your depth, which depends on the session's configure (`CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`). Both conditions must hold, which is why your dispatch prompt carries a nesting flag rather than leaving you to infer one, and why you check whether the tool is **actually there** rather than treating the flag as a guarantee. A spawn that comes back -denied is not an answer about depth: spawns are permission-classified before launch, so read the -error text. The dated record for that harness behavior is +denied is not an answer about depth: a permission rule can refuse a spawn before it launches, so +read the error text. The dated record for that harness behavior is [`${CLAUDE_PLUGIN_ROOT}/reference/parent-contract.md`](${CLAUDE_PLUGIN_ROOT}/reference/parent-contract.md), "Harness facts the dispatch design rests on". @@ -251,11 +253,10 @@ Two dimension-level notes where the preloaded text assumes a human turn or a mai `open_questions` entry. The rule exists to protect intentional deletions, and you cannot get the confirmation it wants. - **Plan mode.** The skill's plan-mode recommendation for high-blast-radius exploration applies to - the inline path only. `EnterPlanMode` is filtered out of every non-fork subagent - unconditionally, and `ExitPlanMode` is filtered from every non-fork subagent too, unless that - subagent's `permissionMode` is `plan`. Your `tools` allowlist lists neither, so you hold neither - either way: plan mode is unreachable from here, and your read-only boundary is the instruction - above. The dated record for that harness behavior is + the inline path only. We treat both plan-mode tools as withheld from a non-fork subagent in your + configuration, and your `tools` allowlist lists neither, so you hold neither either way: plan + mode is unreachable from here, and your read-only boundary is the instruction above. The dated + record for that harness behavior is [`${CLAUDE_PLUGIN_ROOT}/reference/parent-contract.md`](${CLAUDE_PLUGIN_ROOT}/reference/parent-contract.md), "Harness facts the dispatch design rests on". @@ -297,10 +298,11 @@ slice from what the resume returns**, so a payload you can still produce is wort more read. The disk carries the same signal without any payload: an index still marked `Run status: in progress` tells the parent's gate the run stopped short. -**Emit the payload block early and keep it current, as a second channel.** The harness marks -turn-limit output as partial and lets the parent resume you, but it does not document which text -that output carries, and a harness older than v2.1.246 may return none, which is why the disk -marker comes first. The re-emission is kept because it costs no turn of its own: emit it as text +**Emit the payload block early and keep it current, as a second channel.** We rely on a turn-limit +stop returning your output marked partial and on the parent being able to resume you, but which +text that output carries is not documented, and an older harness may return none, which is why the +disk marker comes first (record: the parent contract's "Harness facts the dispatch design rests +on"). The re-emission is kept because it costs no turn of its own: emit it as text on a turn you are already taking for a write, never on a turn by itself. As soon as the scope is resolved, write the block with `status: truncated`, `preload_token` echoed, `preload:` set, `scope_as_received` quoted, and the fields you do not have @@ -361,6 +363,12 @@ disjoint areas, never the six dimensions split across agents, and only when your says nesting is available. Without it, go sequential: slower, same coverage. Write the numbered gap-list before any fan-out either way. +**Your run ends with two things: the persisted `EXPLORE.md` set and the bounded return payload.** +When the outcome gate passes and the index reads `Run status: complete`, return the payload, and +the run is over. Exploration you judge worth doing beyond the scope you were sent, such as a +neighboring area or a deeper look at one already written up, goes into `open_questions` as a named +suggestion with a recommended default, and you do not explore it yourself. + **The parallel worker is the built-in `Explore` agent, and it is a scout.** Spawn one per disjoint area, never one per dimension, on either of the two triggers the preloaded skill body states under "When to fan out, and to what", and under the cap it sets there. Those thresholds live in that one diff --git a/plugins/discovery/reference/parent-contract.md b/plugins/discovery/reference/parent-contract.md index 0fe7741087..47c00e4741 100644 --- a/plugins/discovery/reference/parent-contract.md +++ b/plugins/discovery/reference/parent-contract.md @@ -138,8 +138,9 @@ named here rather than in the template above. Each worker definition pins a defa `explorer` runs on `sonnet`; `researcher`, `intent-tracer` and `research-verifier` run on `opus`. The default is still to **pass nothing**, and then the pin applies. Supply the parameter only to override the pin for a run whose scope earns a different model; it replaces the pin in either direction. Every producing worker -spends `maxTurns: 40` at `effort: high`, and the explorer's are spent almost entirely on reading; why 40 -stays is the harness-facts record "`maxTurns` is set per definition, so 40 is a checkpoint, not a +spends `maxTurns: 40`: `researcher` and `intent-tracer` at `effort: high`, and `explorer` at +`effort: medium`, the effort floor, because its turns go almost entirely to reading. Why 40 stays +is the harness-facts record "`maxTurns` is set per definition, so 40 is a checkpoint, not a completion budget". The pin outranks the consumer's `CLAUDE_CODE_SUBAGENT_MODEL`; `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` still overrides both the pin and the per-call parameter, which it blocks outright. Dated record: @@ -218,20 +219,22 @@ non-fork subagent starts with no history by design. So the operative rule is: > reports rather than repairs**, never as an empty scope to fill in, and never as a license to run > a general sweep. -That rule holds whichever way the harness renders the placeholder, which matters because **the -harness's behavior on this path is not documented in either direction.** Recorded as unsupported, -not as false. Nothing below establishes that a preloaded body renders the placeholder empty, and -nothing establishes that it does not: - -- (raw markdown, fetched 2026-08-11) scopes the placeholder - to invocation: "`$ARGUMENTS` | All arguments passed when invoking the skill." It states that - preload is a different path, "Subagents with preloaded skills work differently: the full skill - content is injected at startup", and says nothing about argument substitution on it. -- (raw markdown, same date) likewise: "The full content - of each listed skill is injected into the subagent's context at startup." No mention of arguments. -- The nearest documented analogue points the *other* way. The `context: fork` walkthrough on the - skills page shows the subagent "receives the skill content as its prompt (`"Research \$ARGUMENTS - thoroughly..."`)", the placeholder arriving as literal text, on a path that is not this one. +That rule holds whichever way the harness renders the placeholder, which matters because **we have +found no page that documents the harness's behavior on this path, in either direction.** Recorded as +unsupported, not as false: nothing we read establishes that a preloaded body renders the +placeholder empty, and nothing establishes that it does not. We checked the skills page's +substitution table, the subagents page's preload section, and the nearest documented analogue, the +skills page's `context: fork` walkthrough, which is a different path and is not evidence for this +one. + +- **Pointer**: for the placeholder, see + [skills: available string substitutions](https://code.claude.com/docs/en/skills#available-string-substitutions); + for preload, see + [subagents: preload skills into subagents](https://code.claude.com/docs/en/sub-agents#preload-skills-into-subagents); + for the analogue, see + [skills: run skills in a subagent](https://code.claude.com/docs/en/skills#run-skills-in-a-subagent). +- **As of**: 2026-10-01 +- **Recheck trigger**: either page starts describing argument substitution on the preload path. **Re-check both pages before restating any mechanism here.** Through 0.14.0 this plugin asserted a specific empty-string rendering of the placeholder on the preload path as settled fact, at five @@ -245,19 +248,22 @@ carries on the preload path; this one is about placeholder-shaped text **a calle into a dispatch prompt. Neither is evidence for the other. All four entry skills point here rather than each carrying its own copy. -**A `${CLAUDE_…}`-shaped token in a topic or scope may not arrive as you typed it.** Stated as what -was observed and what is documented, because the mechanism is neither: +**A `${CLAUDE_…}`-shaped token in a topic or scope may not arrive as you typed it.** - **Observed 2026-08-10:** an argument naming *another* plugin's `${CLAUDE_PLUGIN_DATA}` directory reached a dispatched discovery agent rewritten to **this** plugin's own path. The agent was asked a factually wrong question and answered it correctly. -- **Documented** (`plugins-reference`, `skills`, both fetched 2026-08-11): skill and agent content - is a substitution site for `${CLAUDE_PLUGIN_ROOT}`, `${CLAUDE_PLUGIN_DATA}` and - `${CLAUDE_PROJECT_DIR}` "anywhere the placeholder appears", and there is **no escape** for them. - "A backslash before any other `$` is left unchanged" covers `$ARGUMENTS` and declared argument - names, not these. -- **Not documented on any page:** whether argument-supplied text is itself scanned for those - placeholders. The ordering is unstated, so do not read the observation above as a mechanism. +- **What we now rely on:** skill and agent content is a substitution site for the plugin and + project `${CLAUDE_…}` variables, with no escape for them, and argument text is inserted before + those variables are replaced, which accounts for the observation above. Treat a `${CLAUDE_…}` + token in an argument as rewritten before the agent sees it. +- **Pointer**: for the substitution sites and the escape rule, see + [skills: available string substitutions](https://code.claude.com/docs/en/skills#available-string-substitutions); + for the order of argument insertion and variable replacement, see + [skills: pass arguments to skills](https://code.claude.com/docs/en/skills#pass-arguments-to-skills). +- **As of**: 2026-10-01 +- **Recheck trigger**: the skills page changes the order of argument insertion and variable + replacement, or adds an escape for the `${CLAUDE_…}` variables. Practically: name a path in plain words rather than passing a `${CLAUDE_…}` token and expecting it back. The `topic_as_received` / `scope_as_received` echo-back in the acceptance gate is what catches @@ -265,9 +271,6 @@ this whichever way the substitution actually runs, and it matters most under `/discovery:research-deep`, where one topic is copied into every envelope of an N-way fan-out, so check each dispatched agent's echo against the envelope it was sent, per topic, before synthesis. -**This caveat expires 2027-02-11.** Re-fetch both pages then. After that date it is an unverified -claim, not a fact. Say so rather than repeating it. - ## Credentials stay unread, stated once Every dispatched agent that holds a shell inherits a `Bash` pool (and, run in the background, a @@ -291,33 +294,34 @@ Record a capability you could not establish without reading a value as a gap in reveal a credential is a finding, never a step.** The same pool holds `curl`, so a page that steers the agent into a credential read also has an egress channel. -**Held by instruction; the operator's sandbox can enforce the file half.** +**Held by instruction; the operator's sandbox can enforce the file half.** We rely on two facts: no +subagent frontmatter can block one shell command while keeping the shell, because a +`disallowedTools` entry with a specifier removes the whole tool; and a `permissions.deny` Bash rule +in settings does block the command, for subagents as well as the main conversation. -- *Claim.* No subagent frontmatter can block one shell command while keeping the shell: a - `disallowedTools` entry with a specifier removes the whole tool. A `permissions.deny` Bash rule - in settings blocks the command and applies to subagents as well as the main conversation. -- *Basis.* [Create custom subagents](https://code.claude.com/docs/en/sub-agents): "A - `disallowedTools` entry with a specifier, such as `Bash(git push *)`, still removes the whole - tool from the subagent, not only the matching commands." and "To keep Bash and block specific - commands, add a Bash deny rule such as `Bash(git push *)` to `permissions.deny` in your - settings. The rule applies to the main conversation and to subagents." -- *As of.* Fetched 2026-09-19 (Claude Code 2.1.278); both spans re-verified on the page 2026-09-28. -- *Recheck trigger.* The page stops carrying either quoted span, or a release note names - `disallowedTools` specifier matching or subagent permission inheritance. +- **Pointer**: for both, see + [subagents: available tools](https://code.claude.com/docs/en/sub-agents#available-tools). +- **As of**: 2026-10-01 +- **Recheck trigger**: that section changes how a `disallowedTools` specifier or a settings deny + rule applies to a subagent, or a release note names `disallowedTools` specifier matching or + subagent permission inheritance. Command deny rules are a partial guardrail, not the boundary. `Bash(git credential *)` or `Bash(gh auth token*)` (each with its `PowerShell(...)` twin, because a background subagent keeps `PowerShell`) blocks that one spelling; `printenv`, a `python -c` or `node -e` reader, and every -other program that opens a file stay open, and the -[permissions page](https://code.claude.com/docs/en/permissions) calls Bash patterns that constrain -arguments fragile. A `Read(...)` deny does not cover a subprocess either. The stronger layer is the +other program that opens a file stay open, so we never treat an argument-constraining Bash pattern +as a boundary (Pointer: for why such patterns are unreliable, see +[permissions: Bash](https://code.claude.com/docs/en/permissions#bash). +As of: 2026-10-01. Recheck trigger: the permissions page documents argument matching that holds). +A `Read(...)` deny does not cover a subprocess either. The stronger layer is the operator's sandbox configuration, detailed and dated in the `harness-config` audit's [`required-permissions.md`](https://github.com/melodic-software/claude-code-plugins/blob/main/plugins/harness-config/skills/audit/reference/required-permissions.md) (a URL, because a marketplace install of `discovery` does not carry that plugin's files). A token held in an environment variable sits outside any file boundary and stays held by instruction. The plugin -cannot ship any of this: a plugin's `settings.json` takes only the `agent` and `subagentStatusLine` -keys ([plugins reference](https://code.claude.com/docs/en/plugins-reference), the `settings` field, -fetched 2026-09-29; recheck when that field lists another key). +cannot ship any of this, because a plugin's settings cannot carry permission rules (Pointer: for +the keys a plugin's settings may set, see +[plugins reference: `settings`](https://code.claude.com/docs/en/plugins-reference#settings). +As of: 2026-10-01. Recheck trigger: that field accepts another key). ## Read each file once, stated once @@ -413,136 +417,122 @@ findings returned in place of an artifact are a failed dispatch, not a fallback. Twelve harness behaviors this plugin's dispatch design depends on, each with one dated record here instead of an undated restatement at every site that relies on it. A skill, context file, or agent definition keeps its own one-sentence operative rule and cites this section by heading; none of -them repeats a basis. Records 1-6 were verified against Claude Code 2.1.263 with the pages -named, fetched 2026-09-06. Record 7 was verified against the skills and sub-agents pages -fetched 2026-09-08. Record 8 was verified against Claude Code 2.1.278 with the subagents page -fetched 2026-09-19. Record 9 was verified against the subagents page re-fetched 2026-09-27. -Record 10 was verified against Claude Code 2.1.280 with the sub-agents page fetched 2026-09-27. -Record 11 was verified against the sub-agents and CLI reference pages fetched 2026-09-27. -Record 12 is a first-party reproduction run on 2026-10-01, with the sub-agents page fetched the same day. - -**One shared recheck trigger covers all twelve:** any of the named pages stops carrying the quoted -span, a release note names subagent tool filtering, skill preloading, background execution, -subagent spawn permissions, effort substitution, built-in subagent capabilities, subagent -model resolution, per-invocation subagent parameters, turn-limit output or partial marking, or -`SendMessage` resume, or the CLI major -version moves. On any of those, re-fetch the page before -restating the record, and re-date this section rather than editing a claim in place. +them repeats a pointer. Each record states what we rely on in our words and points at the section +that carries the detail; none restates the page. + +- **As of**: 2026-10-01 for every record below unless the record names its own date, re-read that + day against the sub-agents, skills, permissions and CLI reference pages. +- **Recheck trigger**, shared by all twelve: a record's pointer stops supporting it, a release note + names subagent tool filtering, skill preloading, background execution, subagent spawn + permissions, effort substitution, built-in subagent capabilities, subagent model resolution, + per-invocation subagent parameters, turn-limit output or partial marking, or `SendMessage` + resume, or the CLI major version moves. On any of those, re-read the pointer before restating + the record, and re-date it rather than editing a claim in place. ### A preloaded skill that fails to resolve is skipped silently -*Claim.* A subagent's `skills:` preload that cannot resolve does not fail the dispatch; the agent -runs without the body it was supposed to carry, and the only trace is a debug-log warning. -*Basis.* [Create custom subagents](https://code.claude.com/docs/en/sub-agents): "If a listed skill -is missing or disabled, for example by your organization's policy, Claude Code skips it and logs a -warning to the debug log." The same page's field table gives the mechanism the preload uses: the -`skills` field injects "The full skill content", not only the description. *Why the plugin cares.* -A run whose discipline never loaded is indistinguishable from a good one by every other signal, -which is what the liveness token exists to catch. +*What we rely on.* A subagent's `skills:` preload that cannot resolve does not fail the dispatch; +the agent runs without the body it was supposed to carry, and the only trace is a debug-log +warning. A preload that does resolve carries the whole skill body, not only its description. +*Pointer:* for both, see +[subagents: preload skills into subagents](https://code.claude.com/docs/en/sub-agents#preload-skills-into-subagents). +*Why the plugin cares.* A run whose discipline never loaded is indistinguishable from a good one by +every other signal, which is what the liveness token exists to catch. ### `AskUserQuestion` is removed from every non-fork subagent -*Claim.* A dispatched agent cannot ask the user a question directly; open questions reach a human -only through its return payload and the parent. *Basis.* the same page's tool-filter list, which -names `AskUserQuestion` among the tools the first filter "removes these tools, even when listed in -the `tools` field", and states that forks "skip both filters and receive the main conversation's -exact tool pool". +*What we rely on.* A dispatched agent cannot ask the user a question directly; open questions reach +a human only through its return payload and the parent. A fork keeps the main session's tools. +*Pointer:* for the tools every non-fork subagent loses, see +[subagents: available tools](https://code.claude.com/docs/en/sub-agents#available-tools); for what +a fork keeps, see +[subagents: how forks differ from other subagents](https://code.claude.com/docs/en/sub-agents#how-forks-differ-from-other-subagents). ### Plan-mode tools are removed from every non-fork subagent -*Claim.* A dispatched run cannot enter plan mode, so a read-only posture there is the agent's own -instruction rather than a harness boundary. *Basis.* the same tool-filter list: `EnterPlanMode` -unconditionally, and `ExitPlanMode` "unless the subagent's `permissionMode` is `plan`". +*What we rely on.* A dispatched run cannot enter plan mode, so a read-only posture there is the +agent's own instruction rather than a harness boundary. The one exception, a subagent whose +`permissionMode` is `plan` keeping the exit tool, does not apply to any agent here. *Pointer:* the +same available-tools section. ### The `Workflow` tool is absent from every non-fork subagent -*Claim.* Only the main conversation, or a fork of it, can dispatch a workflow engine, which is why -the deep-research tier ladder runs from main context. *Basis.* the same tool-filter list, which -names `Workflow`. +*What we rely on.* Only the main conversation, or a fork of it, can dispatch a workflow engine, +which is why the deep-research tier ladder runs from main context. *Pointer:* the same +available-tools section. ### Background is the default execution mode, and it narrows the tool set again -*Claim.* A dispatched agent runs in the background unless one of the documented foreground cases -applies, and a background subagent keeps only a named subset of built-in tools plus every MCP -tool. *Basis.* the same page: the second filter "reduces the built-in tool set for subagents that -run in the background, which is the default", and the fork-mode section states that Claude Code -"runs the subagents Claude spawns in the background, forks and non-fork subagents alike, apart -from the cases that stay in the foreground". *Do not restate the tool subset.* It is a list the -harness owns and revises; a site that needs it names the page rather than copying the members. +*What we rely on.* A dispatched agent runs in the background unless one of the documented +foreground cases applies, and a background subagent keeps a smaller set of built-in tools plus +every MCP tool. *Pointer:* for the foreground cases, see +[subagents: run subagents in foreground or background](https://code.claude.com/docs/en/sub-agents#run-subagents-in-foreground-or-background); +for the background tool set, the available-tools section. *Do not restate the tool subset.* It is a +list the harness owns and revises; a site that needs it names the page rather than copying the +members. ### A spawn is permission-checked before it launches, and the depth limit is a different mechanism -*Claim.* A denied spawn is not evidence about nesting depth. The two failures have different +*What we rely on.* A denied spawn is not evidence about nesting depth. A deny rule refuses a spawn +before it launches, while the depth limit takes the `Agent` tool away from a non-fork subagent, so +a subagent at the limit has no tool to call rather than a call that comes back denied; a fork at +the limit keeps the tool and gets an error when it calls it. The two failures have different causes and different error text, so read the error rather than inferring a depth ceiling from it. -*Basis.* the same page. A deny rule refuses the spawn: subagents are blocked with an -`Agent(subagent-name)` entry in the settings `deny` array, and denying the `Agent` tool itself -prevents delegation entirely. The depth limit works the other way (re-fetched 2026-09-27, quoted -with link markup removed): at the limit "Claude Code withholds the `Agent` tool from every -subagent except a fork", so a subagent at the limit has no tool to call rather than a call that -comes back denied, while "A fork at the limit keeps `Agent` in -its inherited tool list, but the tool returns an error instead of spawning." *One bound worth -carrying:* in a subagent definition, listing `Agent` permits nesting while the depth limit allows -it, but "any type list inside the parentheses is ignored". +One bound worth carrying: in a subagent definition, listing `Agent` permits nesting while the depth +limit allows it, and a type list in parentheses there restricts nothing. *Pointer:* for deny rules, +see +[subagents: restrict which subagents can be spawned](https://code.claude.com/docs/en/sub-agents#restrict-which-subagents-can-be-spawned); +for the depth limit, see +[subagents: let subagents spawn their own subagents](https://code.claude.com/docs/en/sub-agents#let-subagents-spawn-their-own-subagents). ### `${CLAUDE_EFFORT}` is the loading context's level -*Claim.* `${CLAUDE_EFFORT}` substitutes the effort level of the context that loaded the skill -(`low`, `medium`, `high`, `xhigh`, or `max`; Ultracode reports as `xhigh`). A skill or -subagent frontmatter `effort` pin overrides the session level while that lane is active, so a -skill preloaded into a pinned worker expands the pin, not the parent's session level. A body -Read from disk is unsubstituted: the placeholder remains the literal characters. *Basis.* -[Skills: available string substitutions](https://code.claude.com/docs/en/skills#available-string-substitutions): -"`${CLAUDE_EFFORT}` | The current effort level: `low`, `medium`, `high`, `xhigh`, or `max`. -Ultracode is not a distinct level and reports as `xhigh`." [Skills: frontmatter -reference](https://code.claude.com/docs/en/skills#frontmatter-reference): `effort` "Overrides -the session effort level." [Create custom subagents](https://code.claude.com/docs/en/sub-agents): -the agent-frontmatter `effort` field "Overrides the session effort level. Default: inherits -from session." *Why the plugin cares.* `/discovery:research` scales source breadth by caller -effort, and `discovery:researcher` is pinned `high` so reasoning does not degrade inside a -session tuned down for cost. The worker's substituted value is therefore the pin. The parent -writes `Source breadth:` from its own load so the table still follows the caller. +*What we rely on.* `${CLAUDE_EFFORT}` substitutes the effort level of the context that loaded the +skill. A skill or subagent frontmatter `effort` pin overrides the session level while that lane is +active, so a skill preloaded into a pinned worker expands the pin, not the parent's session level. +A body Read from disk is unsubstituted: the placeholder remains the literal characters. +*Pointer:* for the placeholder, see +[skills: available string substitutions](https://code.claude.com/docs/en/skills#available-string-substitutions); +for the pin, see +[skills: frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference) and +[subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields). +*Why the plugin cares.* `/discovery:research` scales source breadth by caller effort, and +`discovery:researcher` is pinned `high` so reasoning does not degrade inside a session tuned down +for cost. The worker's substituted value is therefore the pin. The parent writes +`Source breadth:` from its own load so the table still follows the caller. ### The built-in Explore agent cannot hold this plugin's contract -*Claim.* Built-in Explore is a read-only locator: `Write` and `Edit` are denied, it preloads no +*What we rely on.* Built-in Explore is a read-only locator: it cannot write or edit, it preloads no skill, it skips the CLAUDE.md hierarchy and the parent's git status, and it is one-shot with no -agent ID to resume. *Basis.* -[Subagents](https://code.claude.com/docs/en/subagents), built-in subagents: "Tools: read-only -tools; Write and Edit are denied"; "Explore and Plan skip your CLAUDE.md files and the parent -session's git status to keep research fast and inexpensive. Every other built-in and custom -subagent loads both, unless its definition sets the `omitClaudeMd` field"; the what-loads-at-startup -list, "Preloaded skills: full content of any skill named in the agent's `skills` field. Built-in -agents don't preload skills"; and "The built-in Explore and Plan agents are one-shot and return no -agent ID, so Claude can't resume them. Use `general-purpose` or a custom subagent when you need to -continue the work." The same section gives the thoroughness knob a caller passes: "quick for -targeted lookups, medium for balanced exploration, or very thorough for comprehensive analysis." +agent ID to resume. A caller passes it a thoroughness level (`quick`, `medium`, or +`very thorough`). *Pointer:* for its tools, the skip, and the thoroughness level, see +[subagents: built-in subagents](https://code.claude.com/docs/en/sub-agents#built-in-subagents); +for preloading and resume, see +[subagents: what loads at startup](https://code.claude.com/docs/en/sub-agents#what-loads-at-startup) +and [subagents: resume subagents](https://code.claude.com/docs/en/sub-agents#resume-subagents). *Why the plugin cares.* Each denial removes one load-bearing piece of the dispatch contract, which is why built-in Explore is a scout under a worker and never the worker: no `Write` means no artifact set for the acceptance gate to grade, no preload means no discipline to fire the liveness token against, no CLAUDE.md means the project's own conventions never reach it, and no agent ID means a truncated run cannot be resumed. Its read depth is a *judgment* this plugin adds rather than -a documented fact: "Built-in agents have predefined prompts", so how much of a file one read is -neither stated by the page nor recoverable from its report, and a worker therefore treats every -scout hit as a pointer backing `verified: grep`, never `verified: read`. *Not verified:* whether a -user- or project-scope subagent *named* `Explore` inherits the CLAUDE.md and git-status skip. The -page attributes the skip to "the built-in Explore and Plan agents" while stating every other custom -subagent loads both, and says elsewhere "Only Explore and Plan skip it" by name. Setting -`omitClaudeMd: true` on such an override makes the question moot. +a documented fact: a built-in agent runs a prompt we cannot inspect, so how much of a file one read +is neither documented nor recoverable from its report, and a worker therefore treats every scout +hit as a pointer backing `verified: grep`, never `verified: read`. *Not verified:* whether a user- +or project-scope subagent *named* `Explore` inherits the CLAUDE.md and git-status skip; the page +ties the skip to the built-in agents by name. Setting `omitClaudeMd: true` on such an override +makes the question moot. ### A per-invocation `model` outranks a subagent's frontmatter -*Claim.* Claude Code resolves a subagent's model as per-invocation parameter, then the definition's -`model` frontmatter (`inherit` selecting the main conversation's model), then -`CLAUDE_CODE_SUBAGENT_MODEL`, then the main conversation's model. `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` -collapses all of it. *Basis.* -[Subagents](https://code.claude.com/docs/en/subagents): "When Claude invokes a subagent, it can -also pass a `model` parameter for that specific invocation", with that four-step order stated -verbatim; "Before v2.1.251, `CLAUDE_CODE_SUBAGENT_MODEL` came first in this order and overrode both -the per-invocation parameter and the frontmatter, including `model: inherit`"; "While -`CLAUDE_CODE_SUBAGENT_MODEL_FORCE` is on, Claude Code ignores the `model` field of every subagent -definition, including the built-in Explore and Plan subagents, and Claude can't pass a model when -it starts a subagent."; "When you omit it, Claude Code picks the model in the subagent model -order". *Why the plugin cares.* Each worker's frontmatter pin is its default, and the dispatching +*What we rely on.* A subagent's model resolves from the per-invocation parameter first, then the +definition's `model` frontmatter (`inherit` selecting the main conversation's model), then +`CLAUDE_CODE_SUBAGENT_MODEL`, then the main conversation's model; an older harness ranked the +environment variable first; and `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` overrides all of it. *Pointer:* +for the order, the version where it changed, and the force switch, see +[subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model) and +[subagents: run every subagent on one model](https://code.claude.com/docs/en/sub-agents#run-every-subagent-on-one-model). +*Why the plugin cares.* Each worker's frontmatter pin is its default, and the dispatching session overrides it per run with the per-call `model`, which replaces the pin in either direction. An omitted `model` is not a neutral default: it falls to `CLAUDE_CODE_SUBAGENT_MODEL` and then to the main conversation's model, so on a machine without the variable an unpinned worker runs on the @@ -551,48 +541,47 @@ environment variable, so it is a cost defect in a worker definition. ### The verdict lane pins `opus` at `effort: high` -*Claim.* `research-verifier` grades outcome-gate rows 4, 7 and 12, the rows the producer may not -grade, so it is a verdict lane and pins `model: opus` and `effort: high`; `explorer`, mechanical -preparation, stays on `sonnet`. *Basis.* [docs/plugin-philosophy.md](../../../docs/plugin-philosophy.md) -line 1042, "a consequential verdict runs at the session-model tier or above, never below", and line -1187, "Consequential-output lanes with a frontmatter surface pin `high`", both read 2026-09-29. -*Recheck:* an edit to the philosophy's tier or lane rule. +*Decision.* `research-verifier` grades outcome-gate rows 4, 7 and 12, the rows the producer may +not grade, so it is a verdict lane and pins `model: opus` and `effort: high`; `explorer`, +mechanical preparation, stays on `sonnet` at `effort: medium`. *Pointer:* +[docs/plugin-philosophy.md](../../../docs/plugin-philosophy.md), "Model tiers" (a consequential +verdict runs at the session-model tier or above) and "Effort tiers" (consequential-output lanes +pin `high`; the pinned-agents record). *As of:* 2026-10-01. *Recheck trigger:* an edit to the +philosophy's tier rule, lane rule, or pinned-agents record. ### A turn-limit stop returns partial output, and the parent can resume the agent -*Claim.* A subagent that reaches `maxTurns` returns its output marked as partial, and the parent -can resume it with `SendMessage` addressed by agent ID; the resumed run keeps its full history and -continues where it stopped. The marking needs Claude Code v2.1.246 or later, and an older harness -may return nothing at all. *Basis.* [Create custom subagents](https://code.claude.com/docs/en/sub-agents), -quoted with link markup removed: the `maxTurns` field row, "When the subagent reaches the limit, -Claude Code returns its output marked as partial, and Claude can resume it to continue. The -partial marking requires Claude Code v2.1.246 or later"; the resume section, "When a subagent -stops at its `maxTurns` limit, Claude Code marks the returned output as partial. For subagents -that return an agent ID, Claude Code also notes in the result that Claude can message the subagent -to continue from where it stopped.", "Claude uses the `SendMessage` tool with the agent's ID or -name as the `to` field to resume it.", "Resumed subagents retain their full conversation history, -including all previous tool calls, results, and reasoning.", and "The subagent picks up exactly -where it stopped rather than starting fresh." *Why the plugin cares.* It is what makes +*What we rely on.* A subagent that reaches its `maxTurns` limit +returns its output marked as partial, and the parent can resume it with `SendMessage` addressed +by agent ID; the resumed run keeps its full history and continues where it stopped. An older harness, below the version the +`maxTurns` field row names, may return nothing at all. *Pointer:* for the marking and its version +floor, see the `maxTurns` row of +[subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields); +for resume, see +[subagents: resume subagents](https://code.claude.com/docs/en/sub-agents#resume-subagents). +*Why the plugin cares.* It is what makes [Resume first, then decide about the slice](#resume-first-then-decide-about-the-slice) the first -rung rather than a hope. *Not verified:* which text the partial output carries. The page says the -output is "marked as partial" and does not say whether a payload block the agent emitted mid-run -is part of it, which is why the agents keep the disk marker as the primary stop signal. +rung rather than a hope. *Not verified:* which text the partial output carries. We found no page +that says whether a payload block the agent emitted mid-run is part of it, which is why the agents +keep the disk marker as the primary stop signal. ### `maxTurns` is set per definition, so 40 is a checkpoint, not a completion budget -*Claim.* A subagent's `maxTurns` comes from its definition, and the parent cannot change it for -one dispatch. Every producing worker definition here (`explorer`, `researcher`, `intent-tracer`) sets `maxTurns: 40`. The read-only `research-verifier` sets `maxTurns: 30` and stops gathering at turn 24. The number is a checkpoint and a +*What we rely on.* A subagent's `maxTurns` comes from its definition, and the parent cannot change +it for one dispatch: the Agent tool takes no per-call `maxTurns`, an `--agents` JSON definition +sets it for the whole session rather than one dispatch, and the CLI's `--max-turns` is a different +setting for print mode. *Pointer:* for the field, see +[subagents: supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields); +for the Agent tool's per-call parameters, see +[subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model) and +[subagents: subagent names](https://code.claude.com/docs/en/sub-agents#subagent-names); +for `--agents` and `--max-turns`, see their rows in +[CLI reference: CLI flags](https://code.claude.com/docs/en/cli-reference#cli-flags). + +*Decision.* Every producing worker definition here (`explorer`, `researcher`, `intent-tracer`) sets `maxTurns: 40`. The read-only `research-verifier` sets `maxTurns: 30` and stops gathering at turn 24. The number is a checkpoint and a runaway guard for unattended fan-out, not a budget sized to finish the work: each agent stops gathering at its own stop turn to write before the limit, and a run that still reaches the limit -completes through the resume in the record above. *Basis.* -[Create custom subagents](https://code.claude.com/docs/en/sub-agents) lists `maxTurns` as a -frontmatter field, "Maximum number of agentic turns before the subagent stops". The Agent tool -call parameters it documents, such as `model` and `name`, include no `maxTurns`. An `--agents` -JSON definition does accept `maxTurns`, but that route is "Current session", "Pass JSON when -launching Claude Code", so it defines an agent for the whole session rather than widening one -dispatch. The CLI's `--max-turns` is a different setting: -[CLI reference](https://code.claude.com/docs/en/cli-reference), "Limit the number of agentic -turns (print mode only). Exits with an error when the limit is reached. No limit by default." +completes through the resume in the record above. *Why the plugin cares.* No documented or measured basis exists for a different number. Raising it would size the budget to a guess, and removing it would drop the guard on unattended runs, while a limit stop now returns partial output the parent resumes. `contract.test.sh` holds the @@ -639,11 +628,13 @@ the agent to say whether a "Preload liveness" section was in its instructions. E `[Agent: discovery:explorer] Preloaded skill 'discovery:explore'` and no skip warning; the probe run confirmed the definition body was in the agent's context; no run Read `SKILL.md`; every first return carried the YAML block with the skill's `preload_token` and `preload: fired`. Harness 2.1.286 also -delivers a child agent's report as a message to its parent, not as a tool result. The -[sub-agents page](https://code.claude.com/docs/en/sub-agents) states that "a subagent that launches -background subagents waits for their results before it finishes" and that background results "reach -Claude as a completion notification in a later turn". *What this does not -establish.* The reported runs' dispatch prompts, debug logs and checkouts were not reachable, so an +delivers a child agent's report as a message to its parent, not as a tool result, which matches +our reading of the sub-agents page: a parent that launched background subagents waits for them, +and their results arrive in a later turn. *Pointer:* for both, see +[subagents: run subagents in foreground or background](https://code.claude.com/docs/en/sub-agents#run-subagents-in-foreground-or-background) +and +[subagents: let subagents spawn their own subagents](https://code.claude.com/docs/en/sub-agents#let-subagents-spawn-their-own-subagents). +*What this does not establish.* The reported runs' dispatch prompts, debug logs and checkouts were not reachable, so an intermittent harness fault, or a definition text that differed from this commit, is neither confirmed nor excluded. The dispatch without `model:` and `name:` was not run. *Consequence.* The preload and payload half of this record changes nothing in the dispatch envelope: the definition body loads, and the parent's acceptance gate @@ -736,23 +727,29 @@ Anything that lets the run proceed without a script exit reintroduces the defect ### The gate ships no permission grant, and the un-run case is a halt -Neither skill declares `allowed-tools`, and that is a conclusion rather than an omission. Re-checked -against (raw markdown, fetched 2026-08-14): +Neither skill declares `allowed-tools`, and that is a conclusion rather than an omission. Three +legs: -1. **`${CLAUDE_PLUGIN_ROOT}` now substitutes in plugin-skill `allowed-tools` Bash rules** (same page: - "In a plugin skill, Claude Code substitutes `${CLAUDE_PLUGIN_ROOT}` and `${CLAUDE_PLUGIN_DATA}` in - the same two places" as `${CLAUDE_SKILL_DIR}` / `${CLAUDE_PROJECT_DIR}`). That removes the old - "token cannot name these scripts" leg. It does **not** confirm that a +1. **`${CLAUDE_PLUGIN_ROOT}` substitutes in plugin-skill `allowed-tools` Bash rules.** That removes + the old "token cannot name these scripts" leg. It does **not** confirm that a `${CLAUDE_PLUGIN_ROOT}`-bearing rule matches at runtime on every host. Treat the docs change as necessary but not sufficient, and do not ship a grant on docs alone. 2. **An interpreter-led rule is still an anti-pattern in this repo.** A grant shaped like `bash` wrapping the script path names the interpreter and is dropped under auto mode. See `docs/conventions/permission-rule-hygiene/README.md`, anti-pattern 1. A direct-path rule that names the `.sh` (or `.py`) under the plugin root is the documented shape, but see leg 3. -3. **The grant would not last long enough anyway.** It "grants permission for the listed tools - during the turn that invokes the skill … The grant clears when you send your next message." The - parent runs this gate *after* a dispatch returns, which is a later turn; criterion 11 on a - multi-phase research run is likewise later than the invoking turn. +3. **The grant would not last long enough anyway.** We read an `allowed-tools` grant as lasting only + for the turn that invokes the skill. The parent runs this gate *after* a dispatch returns, which + is a later turn; criterion 11 on a multi-phase research run is likewise later than the invoking + turn. + +- **Pointer**: for leg 1, see + [skills: available string substitutions](https://code.claude.com/docs/en/skills#available-string-substitutions); + for leg 3, and for allow rules as the session-wide alternative, see + [skills: pre-approve tools for a skill](https://code.claude.com/docs/en/skills#pre-approve-tools-for-a-skill). +- **As of**: 2026-10-01 +- **Recheck trigger**: either section changes where plugin variables substitute or how long an + `allowed-tools` grant lasts. So the honest statement is the one the rest of this plugin already makes about un-run checks: @@ -761,10 +758,9 @@ So the honest statement is the one the rest of this plugin already makes about u > reading of the directory or of the coverage ledger. The context most motivated to call the run > finished is the one that would be doing the reading. -**Operator setup, once per installed version, optional.** The documented way to cover a -multi-turn command is settings, not frontmatter: "To pre-approve tools for the whole session rather -than a single turn, add allow rules to those permission settings instead." The plugin cannot ship -them: a plugin's `settings.json` supports only the `agent` and `subagentStatusLine` keys. So the +**Operator setup, once per installed version, optional.** We cover a multi-turn command with allow +rules in settings, not frontmatter (pointer above). The plugin cannot ship them, because a plugin's +settings cannot carry permission rules (see "Credentials stay unread, stated once"). So the operator adds them to their own `~/.claude/settings.json`, and `/discovery:setup check` prints them resolved for this install. The rules, with `` replaced by the absolute path this plugin's skills render for `${CLAUDE_PLUGIN_ROOT}`: @@ -792,19 +788,20 @@ The trailing space-and-`*` covers `--help` and every gate argument. **Why the rules pin the version instead of wildcarding it.** A cache install's plugin root carries the version (`…/discovery//`), so these rules stop matching after an update and the gates prompt again; re-run `/discovery:setup check` and paste its output. Writing `…/discovery/*/scripts/…` -instead would survive the update but is unsafe. Claude Code "matches everything before the first `*` -as written" and a `*` "matches any text, including spaces". Tested on Claude Code 2.1.285 (probe -linked under *Basis*): a `*` in the version segment matched across `/`, and +instead would survive the update but is unsafe: a `*` in the version segment of a Bash allow rule +spans `/` and is not path-normalized, so `..` escapes the plugin cache. Our probe on Claude Code +2.1.285 showed it: a `*` in the version segment matched across `/`, and `/cache/discovery/../../outside/scripts/gate.sh` was allowed with no prompt, so the rule matches the command text without normalizing `..` and runs a script outside the plugin cache. A prompt after an -update is the safe failure; a rule that approves a script outside the cache is not. *Claim:* a `*` in -the version segment of a Bash allow rule spans `/` and is not path-normalized, so `..` escapes the -plugin cache. *Basis:* , "Wildcard patterns", fetched -2026-09-30, and the probe recorded at - -(allowed 3 of 3 runs, Claude Code 2.1.285, Linux). *As of:* 2026-09-29, Claude Code 2.1.285. -*Recheck when:* the permissions page documents path normalization or a `*` that stops at `/`, or a -Claude Code release changes the probe result, which would make a version wildcard safe. +update is the safe failure; a rule that approves a script outside the cache is not. + +- **Pointer**: for that behavior, see our probe at + + (allowed 3 of 3 runs, Claude Code 2.1.285, Linux); for how Bash rule wildcards match, see + [permissions: wildcard patterns](https://code.claude.com/docs/en/permissions#wildcard-patterns). +- **As of**: 2026-09-29 +- **Recheck trigger**: the permissions page documents path normalization or a `*` that stops at + `/`, or a Claude Code release changes the probe result, which would make a version wildcard safe. ### What this gate does not grade diff --git a/plugins/docs-hygiene/.claude-plugin/plugin.json b/plugins/docs-hygiene/.claude-plugin/plugin.json index 29d4b2dedb..b275ad9d1a 100644 --- a/plugins/docs-hygiene/.claude-plugin/plugin.json +++ b/plugins/docs-hygiene/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "docs-hygiene", - "version": "0.24.4", + "version": "0.24.5", "description": "Documentation-hygiene toolkit: compress (flavor-trim markdown with a semantic-diff safety net), audit-noise (classify markdown noise), extract-ssot (deduplicate repeated content into a single source of truth), audit-encapsulation (detect citations into skill-private surfaces), rename-references (sweep stale references after renames), audit-derivability (classify whether a whole document earns its existence: could a fresh agent re-derive it from the code?), audit-progressive-disclosure (grade instruction files against a load-tier model for split opportunities and hub/spoke disclosure defects), write-for-agents (authoring-time doctrine that fires while agent-consumed markdown is being written), and write-for-humans (the same moment for the other reader, covering end-user READMEs, RFCs, release notes and guides, and resolving the consuming project's own style guide first); the setup check verifies the markdownlint-cli2 that compress requires.", "author": { "name": "Melodic Software", diff --git a/plugins/docs-hygiene/CHANGELOG.md b/plugins/docs-hygiene/CHANGELOG.md index 2e45888f30..4dd5d54b52 100644 --- a/plugins/docs-hygiene/CHANGELOG.md +++ b/plugins/docs-hygiene/CHANGELOG.md @@ -1,5 +1,25 @@ # Changelog: docs-hygiene plugin +## [0.24.5] - 2026-10-02 + +### Changed + +- **`audit-progressive-disclosure` holds its thresholds as our settings with pointer records.** + `context/tier-model.md` states each number and rule in our words, with a pointer, an as-of date + and a recheck trigger, and `SKILL.md` lists its sources by topic only. The missing-TOC check + records that two Anthropic sources disagree on the threshold (100 versus 300 lines), so a file + between the two gets awareness only and neither number is presented as the single official rule. +- **`write-for-humans` records its four fallback layers in the links-only shape.** The layers in + `reference/sources.md` are this plugin's selections from published standards, each with our + decision, a pointer, an as-of date and a recheck trigger. +- **`write-for-humans` runs checks instead of a self-check.** After writing it runs the project's + prose linter and `/ai-slop:audit`, or reports that the AI-tell check did not run. +- **`write-for-agents` names which surfaces are system prompt**: a subagent body, an output style, + and the launch flags that replace or append to the system prompt. Everything else reaches the + model as conversation content. +- **`rename-references` sets a binary done criterion** for its stale-path pass: zero orphans and + zero stale-but-functional rows, or each remaining row named. + ## [0.24.4] - 2026-10-02 ### Fixed diff --git a/plugins/docs-hygiene/README.md b/plugins/docs-hygiene/README.md index 011f3a8fb7..b39372f32b 100644 --- a/plugins/docs-hygiene/README.md +++ b/plugins/docs-hygiene/README.md @@ -21,10 +21,10 @@ here sweeps the references after a rename someone already made. | `/docs-hygiene:audit-encapsulation` | Detects external citations reaching into skill-private surfaces inside `.claude/skills//` (private subdirectories, heading anchors, schema files) and routes each violation to a remediation path. Ships its own public-surface contract reference. | | `/docs-hygiene:rename-references` | Sweeps stale references after renames, the forms plain token grep misses: slash-command tokens, relative paths from moved files, frontmatter chains and globs, via a 12-form pattern library with audit, half-rename detection, and apply modes. | | `/docs-hygiene:audit-derivability` | Read-only, document-level worth classifier: could a fresh agent re-derive this whole document from the code, config, and structure? Weighs derivability, re-derivation cost, drift risk, and fact ownership into a verdict (delete, convert-to-pointer, keep-as-derivation-cache, keep-owns-facts), splits it by audience, and confirms load-bearing deletions with a fresh-context spot-test. Where the other five trim *inside* a doc, this decides whether the doc should exist. | -| `/docs-hygiene:audit-progressive-disclosure` | Read-only progressive-disclosure classifier: grades agent-facing instruction markdown against a three-tier load-cost model (always-loaded / invocation-loaded / on-demand) and emits seven finding shapes in two lanes. Split opportunities (oversize, mixed-concerns, tier-mismatch) and hub/spoke structure defects (blind-pointer, orphan-spoke, deep-nesting, missing-toc), with tiered treatment guidance. Thresholds are advisory and Anthropic-prescribed; a deterministic `detect.sh` emits the facts, the judgment layer adjudicates. | -| `/docs-hygiene:write-for-agents` | The write-side complement to the audit skills: authoring-time doctrine that fires while agent-consumed markdown is being written (CLAUDE.md/AGENTS.md content, rules files, agent-loaded reference docs, pointer lines, doc-plus-pointer extractions). Two-loads budgeting, branch-covering pointers, steps-vs-reference separation, observable completion criteria, split-by-sequence, positive-form prompting, with a verified auto-read surface reference and a trigger-reliability eval suite. | +| `/docs-hygiene:audit-progressive-disclosure` | Read-only progressive-disclosure classifier: grades agent-facing instruction markdown against a three-tier load-cost model (always-loaded / invocation-loaded / on-demand) and emits seven finding shapes in two lanes. Split opportunities (oversize, mixed-concerns, tier-mismatch) and hub/spoke structure defects (blind-pointer, orphan-spoke, deep-nesting, missing-toc), with tiered treatment guidance. Thresholds are advisory settings of ours, each pointing at the Anthropic guidance it follows from the skill's own records; a deterministic `detect.sh` emits the facts, the judgment layer adjudicates. | +| `/docs-hygiene:write-for-agents` | The write-side complement to the audit skills: authoring-time doctrine that fires while agent-consumed markdown is being written (CLAUDE.md/AGENTS.md content, rules files, agent-loaded reference docs, pointer lines, doc-plus-pointer extractions). Two-loads budgeting, branch-covering pointers, steps-vs-reference separation, observable completion criteria, split-by-sequence, positive-form prompting, which surfaces are system prompt, with a verified auto-read surface reference and a trigger-reliability eval suite. | | `/docs-hygiene:setup` | Check-only setup for the plugin's one external prerequisite: resolves `markdownlint-cli2` from `PATH` or the repository's `node_modules/.bin`, runs it with `--version`, and reports the install remediation when it is missing or broken. Installs nothing and writes nothing. | -| `/docs-hygiene:write-for-humans` | The other half of the write-side pair: authoring-time doctrine for prose a **person** reads. End-user READMEs, RFCs, design docs, release notes, tutorials, how-to guides, reference pages, explanations. Resolves the consuming project's own declared style guide first and reaches for a bundled default set only as the fallback: Diátaxis document modes, Google developer style, ASD-STE100 instruction rules, and Global English disambiguation. The plugin therefore never silently imposes a house style. Ships the mode picker, a rhythm section against machine-cadence prose, one sentence-rules spoke, drift-stamped source records, and a seven-item self-check. | +| `/docs-hygiene:write-for-humans` | The other half of the write-side pair: authoring-time doctrine for prose a **person** reads. End-user READMEs, RFCs, design docs, release notes, tutorials, how-to guides, reference pages, explanations. Resolves the consuming project's own declared style guide first and reaches for a bundled default set only as the fallback: Diátaxis document modes, Google developer style, ASD-STE100 instruction rules, and Global English disambiguation. The plugin therefore never silently imposes a house style. Ships the mode picker, a rhythm section against machine-cadence prose, one sentence-rules spoke, links-only source records, and after-writing checks that run the project's prose linter and `/ai-slop:audit`. | ## Requirements diff --git a/plugins/docs-hygiene/skills/audit-progressive-disclosure/SKILL.md b/plugins/docs-hygiene/skills/audit-progressive-disclosure/SKILL.md index 958ac3a5fe..a0dc9cf22b 100644 --- a/plugins/docs-hygiene/skills/audit-progressive-disclosure/SKILL.md +++ b/plugins/docs-hygiene/skills/audit-progressive-disclosure/SKILL.md @@ -52,7 +52,7 @@ threshold, routing rule, or citation posture (Anthropic-prescribed vs corroborat | structure | `blind-pointer` | Pointer with no when-to-read clause, unmarked execute-vs-read intent, or a vague target name (`doc2.md`, `utils`) | 2 | Attach the condition and intent; rename the target descriptively. On skill descriptions, a missing when-NOT-to-use clause is advisory color (community-sourced), never a violation | | structure | `orphan-spoke` | Bundled spoke no hub references. Unreachable by pointer | 2 | 3-way: add the missing pointer, merge the content up, or delete the spoke | | structure | `deep-nesting` | Spoke-to-spoke chain, required reading more than one level from the hub (documented partial-read failure) | 2; 1 when the chain is the only path to required content | Re-link the deep target directly from the hub, or flatten | -| structure | `missing-toc` | Reference file >300 lines with no TOC = definite; 100–300 lines with none = awareness only, citing the official 100-vs-300 conflict | 1 (>300) / 3 (100–300) | Add a TOC at top (or a grep recipe for lookup-shaped content) | +| structure | `missing-toc` | Reference file >300 lines with no TOC = definite; 100–300 lines with none = awareness only, citing the source conflict on the TOC threshold | 1 (>300) / 3 (100–300) | Add a TOC at top (or a grep recipe for lookup-shaped content) | Thresholds are advisory and tier-calibrated, never hard gates; consuming repos refine them via their own `CLAUDE.md` / rules. There is deliberately **no** "should have spokes" shape: disclosure @@ -130,7 +130,7 @@ sibling divergences it owns. |------|------|-------|------|----------|-----------| | 1 | split | tier-mismatch | 41 | "## Deploy procedure" (multi-step) in always-loaded CLAUDE.md | Move to a skill; leave a one-line pointer | | 2 | structure | blind-pointer | 12 | "[details](context/tier-model.md)" — no when-clause | Attach the read condition and intent | -| 3 | structure | missing-toc | — | reference file, 180 lines, no TOC (official guidance conflicts: 100 vs 300) | awareness only | +| 3 | structure | missing-toc | — | reference file, 180 lines, no TOC (sources disagree on the TOC threshold: 100 vs 300) | awareness only | | 1 | split | tier-mismatch | 41 | "## Deploy procedure" in always-loaded file listed as synced in `sync/README.md` | upstream: file with the owner, citing the decision | ``` @@ -158,8 +158,8 @@ Total: file(s) audited — T1=, T2=, T3=. Facts: files= pointers - Ownership is usually stated outside the audited targets. A file that looks local may be synced or vendored; grep the repo, not just the target. -- The 500/200 numbers are **ceilings, not targets**; the official split trigger is *approaching* - the cap, and the internal-practice hub figure (~30 lines) is far below it. Size alone under the +- The 500/200 numbers are **ceilings, not targets**; we treat *approaching* the cap as the split + trigger, and a well-split hub sits far below it (about 30 lines is common). Size alone under the cap never fires `oversize`. - Invocation-loaded is **cheap to have, not cheap to use**: once a skill body loads, every line recurs for the session, so mutually-exclusive content inside one body defeats the tier and @@ -168,8 +168,9 @@ Total: file(s) audited — T1=, T2=, T3=. Facts: files= pointers are alternates (not required reading) are legitimate. - An unresolved pointer (`resolved=no`) is upstream breakage worth surfacing, but rename sweeps belong to `/docs-hygiene:rename-references`, not here. -- The TOC bands exist because Anthropic's own surfaces disagree (100 vs 300); never present - either number as the single official rule. +- The TOC bands exist because two Anthropic sources disagree on the TOC threshold (the Agent + Skills best-practices page and skill-creator, both under Sources); never present either source's + number as the single official rule. ## What this skill is NOT @@ -183,13 +184,17 @@ Total: file(s) audited — T1=, T2=, T3=. Facts: files= pointers ## Sources +Each entry names the topic this skill relies on the source for, never what the source says; the +pointer, as-of date and recheck trigger for each threshold live in +[context/tier-model.md](context/tier-model.md). + - [Claude Code skills docs](https://code.claude.com/docs/en/skills). Loading levels, listing cap, compaction budgets, split triggers - [Claude Code memory docs](https://code.claude.com/docs/en/memory). CLAUDE.md/rules loading, 200-line target, fact-vs-procedure routing - [Claude Code large-codebases docs](https://code.claude.com/docs/en/large-codebases). Root-orients / per-directory layering, escalation ladder - [Agent Skills best practices](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices). Hub-and-spoke patterns, pointer rules, observed-navigation diagnostics, >100-line TOC guidance - [Agent Skills spec](https://agentskills.io/specification). Frontmatter limits, one-level-deep rule, ~100-token metadata -- [Anthropic engineering: Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills). Mutual-exclusivity split rule -- [Anthropic engineering: context engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents). Just-in-time retrieval, pointer doctrine +- [Progressive disclosure patterns](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#progressive-disclosure-patterns). Mutual-exclusivity split rule (correlate with [Anthropic engineering: Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills)) +- [How Skills work](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview#how-skills-work). Just-in-time retrieval, pointer doctrine (correlate with [Anthropic engineering: context engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents)) - [skill-creator](https://github.com/anthropics/skills/tree/main/skills/skill-creator). Approaching-the-limit split trigger, >300-line TOC guidance - [UC Davis disclosure study (arXiv 2607.17598)](https://arxiv.org/abs/2607.17598). Scale boundary, depth>1 harm (academic corroboration) - [happyskills: listing eviction](https://happyskills.ai). Community source for eviction scoring and the when-NOT-to-use description clause (advisory color only) diff --git a/plugins/docs-hygiene/skills/audit-progressive-disclosure/context/tier-model.md b/plugins/docs-hygiene/skills/audit-progressive-disclosure/context/tier-model.md index 2220a13329..161b1de5f6 100644 --- a/plugins/docs-hygiene/skills/audit-progressive-disclosure/context/tier-model.md +++ b/plugins/docs-hygiene/skills/audit-progressive-disclosure/context/tier-model.md @@ -12,116 +12,155 @@ The reference layer behind `/docs-hygiene:audit-progressive-disclosure`. The hub cites these facts; read this file when adjudicating a finding that needs the exact number, the routing rule, or the pointer criteria. -**Citation posture.** Numbers and routing rules below marked *(Anthropic-prescribed)* are -vendor-defined facts from official Anthropic surfaces. Cite them as Anthropic's prescription, -not as independently verified consensus. Items marked *(corroborated)* carry independent -first-hand corroboration (practitioner measurement, independent implementations, cross-vendor -convergence). Items marked *(community)* come from a single non-official source and are advisory -color only. +**Citation posture.** Every number and rule below is this audit's own setting, in our words; the +pages behind them are pointed at, never quoted. Items marked *(Anthropic-prescribed)* are settings +we take from official Anthropic pages: cite them as Anthropic's prescription, not as independently +verified consensus, and read the wording at the pointer. Items marked *(corroborated)* carry +independent first-hand corroboration (practitioner measurement, independent implementations, +cross-vendor convergence). Items marked *(community)* come from a single non-official source and +are advisory color only. ## The three tiers -| Tier | Surfaces | Cost mechanics | +| Tier | Surfaces | Cost model this audit grades with | |---|---|---| -| **always-loaded** | `CLAUDE.md` / `AGENTS.md` (working dir + ancestors, loaded in full, never truncated), `@path` imports (do NOT reduce cost vs inline), `.claude/rules/*.md` without `paths:` frontmatter, the skill listing (~100 tokens/skill metadata), auto-memory `MEMORY.md` head (first 200 lines / 25KB) | Paid every session, held every turn; adherence degrades with size *(Anthropic-prescribed; tier framing corroborated)* | -| **invocation-loaded** | Skill bodies (`SKILL.md`, on description match or `/name`), path-scoped rules (`paths:` frontmatter), subtree `CLAUDE.md`, agent/command bodies | Cheap to have, **not cheap to use**: once loaded, every line is a recurring token cost for the rest of the session; compaction re-attaches the first 5k tokens per skill, 25k combined *(Anthropic-prescribed)* | -| **on-demand** | Bundled `context/` / `reference/` files, scripts (only output enters context), docs read via pointer | Zero cost until read; "no practical limit" on bundled content *(Anthropic-prescribed)* | +| **always-loaded** | `CLAUDE.md` / `AGENTS.md` (working dir + ancestors, loaded in full), `@path` imports (no saving over inline), `.claude/rules/*.md` without `paths:` frontmatter, the skill listing (~100 tokens/skill metadata), auto-memory `MEMORY.md` head (first 200 lines / 25KB) | Paid every session, held every turn; we treat adherence as degrading with size *(Anthropic-prescribed; tier framing corroborated)* | +| **invocation-loaded** | Skill bodies (`SKILL.md`, on description match or `/name`), path-scoped rules (`paths:` frontmatter), subtree `CLAUDE.md`, agent/command bodies | Cheap to have, **not cheap to use**: once loaded, every line is a recurring token cost for the rest of the session; we price the compaction re-attach at the first 5k tokens per skill, 25k combined *(Anthropic-prescribed)* | +| **on-demand** | Bundled `context/` / `reference/` files, scripts (only output enters context), docs read via pointer | Zero cost until read; we set no size limit on bundled content *(Anthropic-prescribed)* | **Grading rule per tier**: always-loaded content must apply broadly, in every session. The -per-line test: "Would removing this cause Claude to make a mistake?" Invocation-loaded content carries the -same conciseness bar as CLAUDE.md once triggered. On-demand content is free until pulled, so -depth belongs there. A recorded reason to stay always-loaded overrides the test, and upstream -ownership overrides the local treatment (the finding stays): see [Boundaries this audit honors](#boundaries-this-audit-honors). +per-line test: flag a line whose removal would not lead Claude into a mistake. Invocation-loaded +content carries the same conciseness bar as CLAUDE.md once triggered. On-demand content is free +until pulled, so depth belongs there. A recorded reason to stay always-loaded overrides the test, +and upstream ownership overrides the local treatment (the finding stays): see +[Boundaries this audit honors](#boundaries-this-audit-honors). ## Size guidance (all advisory: targets and tips, not validation errors) | Number | Bounds | Status | |---|---|---| -| 500 lines | SKILL.md body cap; the split trigger is **approaching** the limit, not exceeding it ("if you're approaching this limit, add an additional layer of hierarchy along with clear pointers") | Anthropic-prescribed | -| <5k tokens | Recommended SKILL.md body size | Anthropic-prescribed | -| 200 lines | Per-CLAUDE.md target ("longer files consume more context and reduce adherence") | Anthropic-prescribed; stricter 80–150 community practice exists but is not official | +| 500 lines | SKILL.md body cap; the split trigger fires when a file is **approaching** the cap, not only past it | Anthropic-prescribed | +| <5k tokens | Target SKILL.md body size | Anthropic-prescribed | +| 200 lines | Per-CLAUDE.md target; we read a longer file as costing context and adherence | Anthropic-prescribed; stricter 80–150 community practice exists but is not official | | ~100 tokens | Per-skill always-loaded metadata cost | Anthropic-prescribed; corroborated (~80 median measured) | | 1,024 chars | `description` frontmatter validation cap | Anthropic-prescribed (enforced) | -| 1,536 chars | Claude Code listing cap for description + when_to_use per skill; truncation is tail-first, so key use case goes first | Anthropic-prescribed | -| 1% of context window | Skill-listing budget; on overflow descriptions drop lowest-priority-first (usage-frequency/recency scored, a *(community)* detail; official phrasing: least-invoked-first) while names always remain | Anthropic-prescribed; corroborated | -| 200 lines / 25KB | MEMORY.md load limit (excess silently not loaded) | Anthropic-prescribed | -| 1 level | Max reference nesting depth from the hub ("keep references one level deep") | Anthropic-prescribed; corroborated (depth >1 "never helps and sometimes hurts", per the academic source) | -| 100 vs 300 lines | Reference-file length above which a TOC is expected, and **officially inconsistent** (platform best-practices says >100; skill-creator says >300) | Anthropic-prescribed, conflicting, hence the two-band treatment | - -**Provenance of the vendor numbers in both tables** (four-part record): the Claude Code-side -values (1,536 listing cap for `description` + `when_to_use` with tail-first truncation, the -1%-of-context listing budget with least-invoked-first dropping, the 5k-per-skill / 25k-combined -compaction re-attach, the `MEMORY.md` head limits) are owned by - and that page's content-lifecycle -and visibility sections plus and -; the authoring-side values (500-line body cap, <5k-token -body, 200-line CLAUDE.md target, ~100 tokens/skill metadata, one-level nesting, the >100-line TOC -band) are owned by -; the 1,024-char -`description` validation cap is the Agent Skills spec's (), enforced on -upload paths and not by Claude Code locally. Verified 2026-08-31. Recheck trigger: a fetch of the -owning page no longer carrying a table row's value re-derives that row; each fleet audit re-runs -the whole table. - -**Two-band TOC treatment** (this skill's resolution of the official conflict): a reference file -**>300 lines with no TOC** is a definite finding (both official sources agree by then); one at +| 1,536 chars | Claude Code listing cap for description + when_to_use per skill; we treat truncation as tail-first, so the key use case goes first | Anthropic-prescribed | +| 1% of context window | Skill-listing budget; we treat an overflow as dropping descriptions, never names (a usage-frequency/recency scoring of the drop order is a *(community)* detail; read the official order at the pointer) | Anthropic-prescribed; corroborated | +| 200 lines / 25KB | MEMORY.md load limit; we treat the excess as not loaded | Anthropic-prescribed | +| 1 level | Max reference nesting depth from the hub | Anthropic-prescribed; corroborated (an academic source finds deeper nesting no help and sometimes harmful) | +| 100 vs 300 lines | Reference-file length above which a TOC is expected; two Anthropic sources disagree on it (record below), hence the two-band treatment | Anthropic-prescribed, conflicting | + +**Provenance of the vendor numbers in both tables.** + +- **Pointer** (Claude Code-side values: the 1,536 listing cap and its truncation, the 1% listing + budget, the 5k-per-skill / 25k-combined compaction re-attach, the `MEMORY.md` head limits, the + 200-line CLAUDE.md target): + [Frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference), + [Skill content lifecycle](https://code.claude.com/docs/en/skills#skill-content-lifecycle), + [Skill descriptions are cut short](https://code.claude.com/docs/en/skills#skill-descriptions-are-cut-short), + [`skillListingBudgetFraction`](https://code.claude.com/docs/en/settings-reference#skilllistingbudgetfraction), + auto memory's [How it works](https://code.claude.com/docs/en/memory#how-it-works) and + [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions). +- **Pointer** (authoring-side values: the 500-line body cap, the <5k-token body, the ~100 + tokens/skill metadata, one-level nesting, the >100-line TOC band): + [Token budgets](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#token-budgets), + [How Skills work](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview#how-skills-work), + [Avoid deeply nested references](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#avoid-deeply-nested-references) + and [Structure longer reference files with table of contents](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#structure-longer-reference-files-with-table-of-contents). +- **Pointer** (the 1,024-character `description` cap): the Agent Skills specification's + [`description` field](https://agentskills.io/specification#description-field). We treat it as + enforced on upload paths, not by Claude Code locally. +- **As of**: 2026-10-01 +- **Recheck trigger**: a fetch of the owning page no longer carrying a table row's value + re-derives that row; each fleet audit re-runs the whole table. + +**TOC threshold conflict.** The platform best-practices section +[Structure longer reference files with table of contents](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#structure-longer-reference-files-with-table-of-contents) +and skill-creator's +[Progressive Disclosure](https://github.com/anthropics/skills/blob/main/skills/skill-creator/SKILL.md#progressive-disclosure) +disagree on the reference-file length above which a TOC is expected. + +- **As of**: 2026-10-01 +- **Recheck trigger**: either source changes its threshold, or the two come to agree. + +**Two-band TOC treatment** (this skill's resolution of that conflict): a reference file +**>300 lines with no TOC** is a definite finding (both sources agree by then); one at **100–300 lines with no TOC** is awareness-tier only, and the finding text cites the conflict. ## Split triggers (when a file earns a split) 1. **Size**: approaching the tier's guidance number *(Anthropic-prescribed)*. -2. **Mutual exclusivity**: "if certain contexts are mutually exclusive or rarely used together, - keeping the paths separate will reduce the token usage" *(Anthropic-prescribed; the strongest - mixed-concern signal: co-resident content that never co-executes)*. -3. **Kind mismatch**: a CLAUDE.md section "has grown into a procedure rather than a fact" → - skill; multi-step or part-of-codebase entries → skill or path-scoped rule - *(Anthropic-prescribed)*. +2. **Mutual exclusivity**: content for situations that never or rarely arise in the same task, + split so each part loads alone *(Anthropic-prescribed; the strongest mixed-concern signal: + co-resident content that never co-executes)*. +3. **Kind mismatch**: a CLAUDE.md section that has turned into a procedure → skill; multi-step or + part-of-codebase entries → skill or path-scoped rule *(Anthropic-prescribed)*. 4. **Scope mismatch**: instructions relevant to only part of the tree → path-scoped rule or per-directory file *(Anthropic-prescribed)*. -5. **Workflow complexity**: workflows "large or complicated with many steps" → separate files - read per task *(Anthropic-prescribed)*. -6. **Adherence symptoms**: a rule repeatedly ignored suggests the file is too long and the rule - is getting lost *(Anthropic-prescribed; a behavioral trigger, visible in use, not in the file)*. - -Mixed-concern signals *(corroborated)*: one category per skill ("straddling several = confused -skill"), one topic per rules file, cross-file contradiction as a smell, -"one topic per file — do not co-mingle" (Microsoft, independent convergence), -and the case-study direction that refactoring -a mixed 600-line instruction file into 50–150-line topic docs measurably improves task success -(single case study; direction corroborated, percentages illustrative). +5. **Workflow complexity**: a long, many-step workflow → its own file, read per task + *(Anthropic-prescribed)*. +6. **Adherence symptoms**: a rule the model keeps skipping, which we read as a sign the file has + grown past what it can hold *(Anthropic-prescribed; a behavioral trigger, visible in use, not in + the file)*. + +- **Pointer**: for separate files loaded only when needed, see + [Progressive disclosure patterns](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#progressive-disclosure-patterns) + and [Use workflows for complex tasks](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#use-workflows-for-complex-tasks); + for what belongs in CLAUDE.md, see + [When to add to CLAUDE.md](https://code.claude.com/docs/en/memory#when-to-add-to-claude-md); for + the ignored-rule symptom, see + [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md). +- **As of**: 2026-10-01 +- **Recheck trigger**: one of those sections stops supporting the trigger that cites it. + +Mixed-concern signals *(corroborated)*: one category per skill (a skill that straddles several +is a confused one), one topic per rules file, cross-file contradiction as a smell, one topic per +file with no co-mingling (Microsoft guidance, independent convergence), and the case-study +direction that refactoring a mixed 600-line instruction file into 50–150-line topic docs +measurably improves task success (single case study; direction corroborated, percentages +illustrative). ## Pointer-quality criteria (what makes a spoke reachable) A pointer is good when *(Anthropic-prescribed unless noted)*: -1. **Direct from the hub, one level deep**: chained pointers trigger partial reads (the - documented `head -100` preview failure). -2. **Condition attached**: the pointer states WHEN to read the target ("For tracked changes: - see REDLINING.md"); a bare link is the documented "missed connection" failure. -3. **Intent marked**: execute vs read ("Run `x.py` to extract" vs "See `x.py` for the - algorithm"). -4. **Self-describing target name**: `form_validation_rules.md`, not `doc2.md` / `helper` / - `utils`; organize by domain. +1. **Direct from the hub, one level deep**: we treat a chained pointer as inviting a partial + preview read of the deeper file instead of a full read. +2. **Condition attached**: the pointer states WHEN to read the target ("For a schema change: + read MIGRATIONS.md"); a bare link is one Claude may never follow. +3. **Intent marked**: execute vs read ("Run `fetch.py` to pull the rows" vs "Read `fetch.py` + for the retry rules"). +4. **Self-describing target name**: `retry_policy.md`, not `notes2.md` / `helper` / `utils`; + organize by domain. 5. **Navigable target**: long references open with a TOC so partial reads still see the scope; a grep recipe beats a full read for lookup-shaped content. 6. **Portable path form**: forward slashes, relative from the skill root; fully-qualified MCP tool names in the form Claude Code resolves, `mcp____` for a configured server - and `mcp__plugin____` for a server a plugin bundles - (basis: [permissions, "MCP"](https://code.claude.com/docs/en/permissions#mcp) and - [MCP, "Plugin-provided MCP servers"](https://code.claude.com/docs/en/mcp#plugin-provided-mcp-servers), - verified 2026-09-10; recheck when either page changes the form). The platform page's - `ServerName:tool_name` form - ([MCP tool references](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#mcp-tool-references)) - applies to other surfaces and is not the harness form. - -Description-as-trigger (the always-loaded pointer to a skill body): state what the skill does AND -when to use it, third person, key use case first (tail-first truncation at 1,536 chars strips + and `mcp__plugin____` for a server a plugin bundles. + Pointer: [permissions, "MCP"](https://code.claude.com/docs/en/permissions#mcp) and + [MCP, "Plugin-provided MCP servers"](https://code.claude.com/docs/en/mcp#plugin-provided-mcp-servers). + As of: 2026-09-10. Recheck trigger: either page changes the form. The platform page's + [MCP tool references](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#mcp-tool-references) + form applies to other surfaces; we grade against the harness form. + +- **Pointer** (criteria 1-5): [Avoid deeply nested references](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#avoid-deeply-nested-references), + [Progressive disclosure patterns](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#progressive-disclosure-patterns), + [Provide utility scripts](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#provide-utility-scripts), + [Runtime environment](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#runtime-environment) + and [Structure longer reference files with table of contents](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#structure-longer-reference-files-with-table-of-contents). +- **As of**: 2026-10-01 +- **Recheck trigger**: one of those sections stops supporting the criterion that cites it. + +Description-as-trigger (the always-loaded pointer to a skill body): name the skill's job AND the +situations that call for it, third person, key use case first (tail-first truncation at 1,536 chars strips trailing keywords). A when-NOT-to-use clause in descriptions is *(community)* guidance, so surface it as advisory color only, never as an official requirement. **Observed-navigation diagnostics** *(Anthropic-prescribed method)*: repeatedly re-read spoke → promote its content to the hub; never-read spoke → demote, re-signal, or delete; failed -reference-follow → make the link more explicit. +reference-follow → make the link more explicit. Pointer: +[Observe how Claude navigates Skills](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices#observe-how-claude-navigates-skills). +As of: 2026-10-01. Recheck trigger: that section is removed or changes the method. ## Boundaries this audit honors diff --git a/plugins/docs-hygiene/skills/rename-references/context/audit-modes.md b/plugins/docs-hygiene/skills/rename-references/context/audit-modes.md index 83467346f2..a5f2563d1c 100644 --- a/plugins/docs-hygiene/skills/rename-references/context/audit-modes.md +++ b/plugins/docs-hygiene/skills/rename-references/context/audit-modes.md @@ -101,9 +101,9 @@ Files containing BOTH (incomplete rename state): **When to invoke:** -- After `/docs-hygiene:rename-references to ` apply phase, double-check no file paths went stale +- After `/docs-hygiene:rename-references to ` apply phase, run it to list file paths the apply left stale - After `git mv `, sweep for `[text]()` markdown links and similar -- Before declaring rename done: orphan check is final safety net beyond apply.md Phase 6 re-sweep +- Before reporting a rename done, after apply.md Phase 6 re-sweep: the rename is done when this mode reports `0 orphans, 0 stale-but-functional`, or each remaining row is named in the report with its reason **Inputs:** diff --git a/plugins/docs-hygiene/skills/write-for-agents/SKILL.md b/plugins/docs-hygiene/skills/write-for-agents/SKILL.md index 94e40652a0..a1b57d8fa8 100644 --- a/plugins/docs-hygiene/skills/write-for-agents/SKILL.md +++ b/plugins/docs-hygiene/skills/write-for-agents/SKILL.md @@ -42,6 +42,24 @@ not transfer. A bottom-line-first opening, headings written to be skimmed, bulle scanning, and bounded bold all serve a person moving down a page fast, and an agent reading a rule needs the rule stated where it applies, not staged for a skim. +## Know which surfaces are system prompt + +Three vehicles shape the system prompt: a subagent definition's body, an output style, and the +launch flags that replace or append to the system prompt (`--system-prompt`, `--append-system-prompt`, and their `-file` +forms). CLAUDE.md files, rules, and skill bodies reach the model as conversation content in the +user turn. Write those as instructions read in the conversation: state the rule, its scope, and +its reason. A rule that must hold at system-prompt level goes in a subagent body, an output style, +or a launch flag. + +- **Pointer**: for how CLAUDE.md content is delivered, see + ; for a subagent body, see + ; for output styles, see + ; for the flags, see + ; for skill content, see + . +- **As of**: 2026-10-01 +- **Recheck trigger**: one of those sections changes how its content reaches the model. + ## Write pointers that cover their branches A pointer is a routing instruction; the reader decides whether to follow it from the pointer diff --git a/plugins/docs-hygiene/skills/write-for-humans/SKILL.md b/plugins/docs-hygiene/skills/write-for-humans/SKILL.md index 4527eb6deb..482ff76852 100644 --- a/plugins/docs-hygiene/skills/write-for-humans/SKILL.md +++ b/plugins/docs-hygiene/skills/write-for-humans/SKILL.md @@ -142,9 +142,13 @@ ambiguity). "If exceeded" gets a subject: the request (ambiguity). ## After writing -- **Check for AI-writing tells.** Invoke `/ai-slop:audit` via the Skill tool when it is available in - the session; when it is not, re-read for the obvious tells yourself: filler, stacked hedging, - negative parallelism, and promotional tone. Then say that you did the lighter pass. +- **Run the project's prose linter.** When the repository configures one (a Vale, markdownlint, or + textlint config), run it on the files you wrote and fix what it reports in them. Report the + command and its result. +- **Check for AI-writing tells.** Invoke `/ai-slop:audit` on the files you wrote via the Skill tool + when it is available in the session; its detector reports filler, stacked hedging, negative + parallelism, and promotional tone by line. When it is not available, report that the AI-tell + check did not run. - **Repeated the same prose in another file. Even a second occurrence, or a recap of an SSOT that already exists?** Invoke `/docs-hygiene:extract-ssot` via the Skill tool. Creating a new shared home still waits for the third occurrence; below that it remedies the repetition in place. @@ -152,26 +156,6 @@ ambiguity). "If exceeded" gets a subject: the request (ambiguity). tool; it trims flavor behind a semantic-diff guard rather than rewriting. - **Writing markdown an agent will load instead?** That is `/docs-hygiene:write-for-agents`. -## Self-check before handing back - -**Check the draft against the standard you resolved.** Questions 4, 6 and 7 restate the three rules -above, so they apply whichever standard that was. The other four come from the bundled layers: when -the project declared its own guide, that guide supplies their equivalents and these four do not -apply. Reaching for them would impose the bundled set on a project that already chose. Answer -whichever apply against the text you just wrote, not from memory of writing it. - -1. Is each document one mode, with links where modes meet? Confirm it by naming the mode. -2. Is every instruction a command, with its condition in front? -3. Does any sentence carry two instructions, or two thoughts? Split it until each carries one. -4. Can any word be cut without losing meaning? Cut it. -5. Is "only" next to the word it changes? Does every "it" point at one obvious thing? Does every - clause keep its verb? -6. Does each thing have exactly one name throughout? -7. Would a developer say these words out loud? Replace invented metaphors and fancy synonyms with - the plain word or the real name. - -The draft is done when every answer is yes. A "no" is a rewrite now, not a note for later. - ## What this skill does NOT do - **Does not impose a style guide.** The consuming project's declared standard wins; the bundled set @@ -192,7 +176,7 @@ The draft is done when every answer is yes. A "no" is a rewrite now, not a note copy, not documentation; they follow your product's own copy guidelines. - **Does not author skills**, a SKILL.md is `playbooks:skill-authoring` and `skill-quality:check` territory. -- **Does not claim to be the standards it names.** Each layer is a paraphrase of a published +- **Does not claim to be the standards it names.** Each layer is our selection from a published standard; see the source records below. ## Gotchas @@ -210,7 +194,7 @@ The draft is done when every answer is yes. A "no" is a rewrite now, not a note ## Source records -The four bundled layers are distilled paraphrases of published standards, each with a four-part -drift stamp, claim, basis, as-of date, recheck trigger, in +The four bundled layers are our selections from published standards, each with a record (our +decision, a pointer, an as-of date, a recheck trigger) in [`reference/sources.md`](reference/sources.md). Read it before citing a layer as the standard: the STE layer is a principles subset, and a document written to it is not thereby STE-conformant. diff --git a/plugins/docs-hygiene/skills/write-for-humans/reference/sources.md b/plugins/docs-hygiene/skills/write-for-humans/reference/sources.md index f0b36923e9..d4af4d7f4a 100644 --- a/plugins/docs-hygiene/skills/write-for-humans/reference/sources.md +++ b/plugins/docs-hygiene/skills/write-for-humans/reference/sources.md @@ -1,39 +1,43 @@ # Source records for the default layer set -The four layers this skill falls back to are **distilled paraphrases** of published standards, never -copies of them, and none of them is this plugin's own invention. Each carries a four-part drift -stamp so a later reader can tell what was claimed, on what basis, when it was last true, and what -event should send someone back to check. +The four layers this skill falls back to are this plugin's selections from published standards, +never copies of them, and none of them is this plugin's own invention. Each carries a record in +the links-only shape (our decision, a pointer, an as-of date, a recheck trigger) so a later reader +can tell what we took, where to read the standard live, when the selection was last checked, and +what event should send someone back to check. Read this when you need to know how faithful a layer is, cite a layer to someone, or decide whether -a standard has moved since the port. +a standard has moved since the selection was made. ## Diátaxis: the mode layer -- **Claim.** The four modes, and the doing/understanding × learning/work compass that selects - between them, are as the framework defines them. -- **Basis.** [diataxis.fr](https://diataxis.fr). -- **As of.** 2026-07-18. -- **Recheck trigger.** The framework publishes a revision that renames a mode or changes either +We use the framework's four modes, and the compass that selects between them, as the framework +defines them; the framework settles any mode question this skill does not cover. + +- **Pointer**: for the modes and the compass, see [diataxis.fr](https://diataxis.fr). +- **As of**: 2026-07-18 +- **Recheck trigger**: the framework publishes a revision that renames a mode or changes either compass axis. ## Google developer documentation style: the address layer -- **Claim.** The address rules paraphrase the guide's own highlights; they are a selection, not the - guide, and the guide settles anything this file does not cover. -- **Basis.** [developers.google.com/style](https://developers.google.com/style). -- **As of.** 2026-07-18. -- **Recheck trigger.** The guide's Highlights page changes a rule stated in `sentence-rules.md`. +The address rules are our selection from the guide's highlights, not the guide; the guide settles +anything this file does not cover. + +- **Pointer**: for the guide and its highlights, see + [developers.google.com/style](https://developers.google.com/style). +- **As of**: 2026-07-18 +- **Recheck trigger**: the guide's Highlights page changes a rule stated in `sentence-rules.md`. ## ASD-STE100 Simplified Technical English: the load layer -- **Claim.** The load rules are the transferable core of the specification's writing rules. The - numbered rules and the controlled dictionary live in the specification itself and are **not** - reproduced here. This layer is a set of principles derived from the standard, and a document - written to it is not thereby STE-conformant. -- **Basis.** [asd-ste100.org](https://asd-ste100.org), Issue 9 (2025). -- **As of.** 2026-07-18. -- **Recheck trigger.** A new Issue of the specification is published. +The load rules are principles we derive from the specification's writing rules. The numbered rules +and the controlled dictionary live in the specification and are **not** reproduced here, and a +document written to this layer is not thereby STE-conformant. + +- **Pointer**: for the specification, Issue 9 (2025), see [asd-ste100.org](https://asd-ste100.org). +- **As of**: 2026-07-18 +- **Recheck trigger**: a new Issue of the specification is published. This caveat is a real constraint, not boilerplate. Anyone claiming STE conformance for a document needs the specification; anyone wanting sentences that load one idea at a time can use the @@ -41,11 +45,12 @@ principles alone. ## Global English: the ambiguity layer -- **Claim.** The ambiguity rules paraphrase Kohl's guidelines for writing prose that survives - non-native readers, translators, and machine parsers. -- **Basis.** Kohl, *The Global English Style Guide* (SAS Press). -- **As of.** 2026-07-18. -- **Recheck trigger.** A new edition is published. +The ambiguity rules are our selection from Kohl's guidelines for prose that survives non-native +readers, translators, and machine parsers. + +- **Pointer**: for the guidelines, see Kohl, *The Global English Style Guide* (SAS Press). +- **As of**: 2026-07-18 +- **Recheck trigger**: a new edition is published. ## Why these four and not one diff --git a/plugins/guardrails/.claude-plugin/plugin.json b/plugins/guardrails/.claude-plugin/plugin.json index 74aee10a08..712db42542 100644 --- a/plugins/guardrails/.claude-plugin/plugin.json +++ b/plugins/guardrails/.claude-plugin/plugin.json @@ -171,5 +171,5 @@ "min": 1 } }, - "version": "0.46.6" + "version": "0.46.7" } diff --git a/plugins/guardrails/CHANGELOG.md b/plugins/guardrails/CHANGELOG.md index bfcb2dd37c..8458282183 100644 --- a/plugins/guardrails/CHANGELOG.md +++ b/plugins/guardrails/CHANGELOG.md @@ -3,6 +3,16 @@ All notable changes to the `guardrails` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.46.7] - 2026-10-02 + +### Changed + +- **The README's hook-behavior notes state our decisions and point at the hooks page.** The guard + timeout tail, the plugin-bundled GitHub server matcher, the decision not to set `async: true` on + the three advisory verify guards, and the `node` and `jq` prerequisites are each a decision with + a pointer to the anchored hooks section, an as-of date and a recheck trigger, and none of the + page's wording is stored. No guard behavior changes. + ## [0.46.6] - 2026-10-02 ### Fixed diff --git a/plugins/guardrails/README.md b/plugins/guardrails/README.md index f20a9b9c21..e6d4a43728 100644 --- a/plugins/guardrails/README.md +++ b/plugins/guardrails/README.md @@ -312,8 +312,10 @@ out of scope until such a signal exists. freeze the session, with the dual-channel notice so the allow is not silent. Stdin timeout and a NUL payload still fail closed. The 60s `hooks.json` `timeout` on this handler is a harness-level fail-open the plugin does not - override: if the process is killed at that bound, the tool call proceeds - ([hooks: Timeouts](https://code.claude.com/docs/en/hooks#timeouts)). What the + override: we treat a guard killed at that bound as letting the tool call + proceed. Pointer: . As of: + 2026-10-01. Recheck trigger: that section changes what a timed-out + `PreToolUse` command hook does to the tool call. What the plugin does instead is keep the row far from that bound: the command tokenizer is linear in the command's length, one parse serves every guard, and a command over `MAX_COMMAND_LEN` is refused by the first guard that @@ -330,10 +332,11 @@ out of scope until such a signal exists. tools.** `secret-pattern-detection` and `hardcoded-path-check` inspect `mcp__github__push_files` (every entry of its `files` array, not just the first) and `mcp__github__create_or_update_file`, and, since **0.37.1**, the - same two tools from a plugin-bundled GitHub server, which Claude Code names - `mcp__plugin__github__` - ([hooks reference](https://code.claude.com/docs/en/hooks.md), "plugin-bundled - MCP server", checked 2026-09-27). This closes a real hole: a + same two tools from a plugin-bundled GitHub server, matched as + `mcp__plugin__github__`. Pointer: + . As of: 2026-09-27. + Recheck trigger: that section changes how a plugin-bundled server's tools are + named. This closes a real hole: a `Write|Edit` matcher does not see an MCP write, so a session could be cleared by these guards and still push the same secret to a repository by another route, where there is no local file to fix afterwards and no `pre-commit` @@ -588,9 +591,9 @@ detected from `.git/shallow`. The three report-only rows stay synchronous. -- **Decision**: do not set `async: true` on `cli-flag-verify`, `skill-reference-verify`, or `stale-path-verify`. -- **Basis**: [hooks reference](https://code.claude.com/docs/en/hooks), "Run hooks in the background", re-fetched 2026-09-28. An async hook's `additionalContext` and `systemMessage` are delivered on the next conversation turn and are not shown to the user. In an idle session the response waits for the next user message. Under `claude -p`, a hook still running at teardown is killed. `timeout` is not enforced on an async hook. These findings are advisory context for the edit that just landed; a next-turn delivery misses that edit. Blocking guards stay synchronous and fail closed. -- **As of**: 2026-09-28. +- **Decision**: do not set `async: true` on `cli-flag-verify`, `skill-reference-verify`, or `stale-path-verify`. These findings are advisory context for the edit that just landed, and an async row would deliver them after that edit, could lose them at the end of a `claude -p` run, and would not be bounded by the row's `timeout`. Blocking guards stay synchronous and fail closed. +- **Pointer**: for async delivery, `-p` teardown and `timeout` on an async hook, see and . +- **As of**: 2026-09-28 - **Recheck trigger**: that section changes when async output is delivered beside the tool result, when `-p` waits for a running async hook, or when `timeout` applies to one. **0.38.0, a long command (#4528).** 2026-09-27, Linux 6.12, bash 5.2.21, @@ -1372,20 +1375,22 @@ as before. fails **open** (disabled) and prints a one-line stderr notice, never a silent disable. - **Node.js** on `PATH`. Every guard row starts through `hooks/exec-bash.mjs`, which finds bash - and runs the guard; the script declares no minimum Node version. Claude Code resolves an - exec-form `command` on `PATH` ([exec form](https://code.claude.com/docs/en/hooks#exec-form-and-shell-form)). - Its hooks reference documents a hook that cannot start as a - [non-blocking error](https://code.claude.com/docs/en/hooks#other-exit-codes) for most events, - with a missing script as the example, and does not document a `command` absent from `PATH`. - `/guardrails:check` reports a missing `node` or `jq`. A `SessionStart` row in shell form + and runs the guard; the script declares no minimum Node version. The exec-form row needs `node` + resolvable on `PATH`, and we treat a row that cannot start because `node` is missing as a guard + that enforced nothing, with no documented error to rely on. `/guardrails:check` reports a missing + `node` or `jq`. A `SessionStart` row in shell form (`"shell": "bash"`, no `args`) runs `command -v node` and needs no node itself. When node is absent it exits 0 with JSON: `systemMessage` shows the user a warning and `additionalContext` tells the model that the guards cannot launch and enforce nothing. It prints nothing when node is present. It does not read the per-guard toggles, because an unset toggle exports no environment variable and the row would need every guard's key listed by hand; a host that turns - every guard off should disable the plugin instead. Basis: https://code.claude.com/docs/en/hooks, "SessionStart" (plain stdout reaches - Claude only, and exit-2 stderr reaches the user only) and "JSON output" (`systemMessage` is a - warning shown to the user). + every guard off should disable the plugin instead. Pointer: for how an exec-form `command` + resolves, see ; for a hook that + cannot start, see ; for where + `SessionStart` stdout goes, see ; for + `systemMessage`, see . As of: 2026-10-01. + Recheck trigger: the hooks page documents a `command` absent from `PATH`, or changes who sees + `SessionStart` stdout or `systemMessage`. - On Windows, **Git Bash** (the hooks run via Git Bash's bash). - `cli-flag-verify` runs ` --help` for the binaries it scans; findings require those binaries on PATH (missing binaries are skipped, never flagged). diff --git a/plugins/harness-config/.claude-plugin/plugin.json b/plugins/harness-config/.claude-plugin/plugin.json index fefc830b34..1e94301869 100644 --- a/plugins/harness-config/.claude-plugin/plugin.json +++ b/plugins/harness-config/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "harness-config", - "version": "1.2.1", + "version": "1.3.0", "description": "Nine configuration-health skills (plus setup) for a repo's Claude Code configuration: audit (settings.json / .mcp.json / hooks / plugins / permissions drift), audit-automation-gaps (evidence-gated verdicts on automation gaps), audit-permission-grants (allow-rule / allowed-tools grants for auto-mode durability and portability), audit-permission-state (the permission rules actually in effect: every settings scope merged with per-rule provenance, what auto mode drops on entry, config written where nothing reads it, and which managed intents are enforced versus loosenable), draft-auto-mode-rules (interview and draft a paste-ready autoMode classifier block; prints only, never writes), audit-instructions (locally-owned instruction surfaces vs current model capability, proposing removals/rewrites of instructions the model no longer needs, and detecting cross-surface instruction conflicts), audit-prompting-postures (the additive lane: posture guidance the prompting guide says a component's purpose needs but the component does not carry), audit-pass (one coordinated, ordered, resumable pass over a named target: three-scope inventory, run-time-derived exclusion set, stable finding identity, suppression memory, resume, one human gate, delegating every check to the plugin that owns it), and unhobble (the empirical bare-baseline experiment: reversibly strip a repo's standing instructions, log real stumbles against the current model, re-add only what evidence earns). Boundary: harness-memory owns the health of CLAUDE.md, AGENTS.md, CLAUDE.local.md, .claude/rules/ and auto-memory (structure, size, placement, index integrity); harness-config audit-instructions judges whether instruction text across those files and skills, agents and hooks still fits the current model, and runs no memory-file hygiene checks.", "author": { "name": "Melodic Software", diff --git a/plugins/harness-config/CHANGELOG.md b/plugins/harness-config/CHANGELOG.md index e0ae7e522b..386628c6dc 100644 --- a/plugins/harness-config/CHANGELOG.md +++ b/plugins/harness-config/CHANGELOG.md @@ -5,6 +5,60 @@ All notable changes to the `harness-config` plugin are documented here. Format f Versions 0.51.8 to 0.51.9 and 0.51.11 to 0.51.14 were reserved by parallel branches and never released. +## [1.3.0] - 2026-10-02 + +### Added + +- **`audit-instructions` rows I36 and I37** (`criteria.md` 1.26.0), both scoped to `sonnet-5-5`: + I36 flags an instruction that limits tool or search use as a general policy on a component that + has a search or retrieval tool; I37 flags model-visible text added after every tool result in an + interactive session without a condition. +- **`audit-prompting-postures` P12, P13 and P14** and an `ideating` purpose: no self-started steps + after the run's end at `xhigh` or `max`, a runnable check behind a done claim, and ideas first on + an open-ended request, each stated as a check over the component's own text. P12 alone carries + the Sonnet 5.5 model condition; the catalog no longer says every row citing that subpage is + model-neutral. +- **The `audit` checklist flags a code-changing or verifying component pinned below `medium`**, + the marketplace effort floor. + +### Changed + +- **The `audit-instructions` criteria catalog keeps its firing rules in our words and names its + sources.** A row's Source line names the page and section that documents its mechanic and quotes + nothing; the firing rule above it works without a live fetch. A row that acts on a volatile + upstream literal keeps the literal as its own setting and carries the links-only record: the + decision, a pointer to the exact section, an as-of date and an observable recheck trigger. The + vendor blog posts the catalog cites are marked correlate-only beside their docs pointer. +- **`audit-instructions` records a source conflict.** The Fable 5 guide and the Opus 5 and Opus 4.8 + guides disagree on throttling subagent dispatch, so the catalog records the disagreement with the + guide sections and keeps its per-target treatment of a throttle. +- **The `audit-permission-state`, `audit-prompting-postures` and `agents-md-liveness` records use the + same shape.** The conditions under which `AGENTS.md` support is unavailable, the four **Project + instructions** values and where the setting is honored are stated as what our checks read, each + with a pointer to its anchored section of the memory page, an as-of date and a recheck trigger, + and none of the page's wording is stored. +- **`audit-instructions` widens to Sonnet 5.5 and current models.** I10 fires on `sonnet-5-5`; I17 + and I17-b cover `between_tools` in API code; the model sets of I17, I17-a, I17-c and I25 cover + Opus 5.5, Sonnet 5.5 and Fable 5.1; I17's ultracode count is narrowed to the effort that reaches + the request, and the row says so. I8-c records a declined `sonnet-5-5` widening and splits + `between_tools` ownership with I17; I17-a's Detect and Must-NOT agree; I17-b leaves the cache cost + of an effort change to Claude Code, with a pointer; rows that keep tokens of older models are + re-justified, and Fable 5 is treated as current. +- **`audit-prompting-postures` repoints P2, P5, P6 and P8** to the current model subpages (P5 to the + one section that still covers its check), with a Sonnet 5.5 pointer block. +- **The `audit` checklist's records are links-only**, and its exec-form, plugin-manifest, + fallback-model, effort-default and managed-settings rows are reduced to firing rules plus + pointers. +- The `AGENTS.md` liveness reference lists three availability conditions with their version + floors, re-read against the memory page. audit-permission-state reads which sessions start in + auto mode at the permission-modes page instead of keeping a copy of its table. audit- + instructions points I25 and its Sources at the Sonnet 5 model page, since the old what's-new URL + now redirects, and names the surface pages behind I17-c's promotion gate. Postures P2 and P6 + state their present-when conditions in our own terms. +- audit-instructions I14 no longer says only the built-in Explore and Plan agents skip `CLAUDE.md`: + a custom agent with `omitClaudeMd: true` does too, so a rule restated for it is not flagged as + redundant. The record names the sub-agents page's internal disagreement on that field. + ## [1.2.1] - 2026-10-02 ### Changed diff --git a/plugins/harness-config/reference/agents-md-liveness.md b/plugins/harness-config/reference/agents-md-liveness.md index 8b55d27728..ac76a0fac1 100644 --- a/plugins/harness-config/reference/agents-md-liveness.md +++ b/plugins/harness-config/reference/agents-md-liveness.md @@ -2,10 +2,10 @@ Every check in this plugin that decides whether an `AGENTS.md` is a live instruction surface turns on two questions, in this order: is `AGENTS.md` support **available** in the session at all, and if -so, which files does the **Project instructions** setting load. Both are recorded here as the -four-part record the -[upstream-drift convention](../../../docs/conventions/upstream-drift/README.md) defines: claim, -basis, as-of date, recheck trigger. +so, which files does the **Project instructions** setting load. Both are recorded here in the +shape the +[upstream-drift convention](../../../docs/conventions/upstream-drift/README.md#required-parts) +defines: our decision, a pointer to the upstream section, the as-of date, the recheck trigger. **Why the record lives here.** A plugin never imports files from a sibling plugin ([plugin philosophy](../../../docs/plugin-philosophy.md), "It never imports files from a sibling @@ -15,73 +15,75 @@ into another plugin's private reference would leave these conditions unresolvabl standalone install. Each record below cites the upstream page directly. Where a sibling plugin keeps its own record of the same upstream fact, that is a parallel record, not this one's source. -## Availability: three documented conditions, any one of which ends the question - -- **Claim**: in these sessions "Claude reads `CLAUDE.md` files only, and **Project instructions** - doesn't appear in the `/config` settings panel", so no `AGENTS.md` is read natively at any path - under any mode: - 1. "You're on a Claude Code version before v2.1.277"; - 2. "You disabled the built-in `agents-md` plugin in `/plugin`"; - 3. "In some cases, it's your first session after you upgrade from v2.1.276 or earlier. Claude - reads `AGENTS.md` from your next session on". - - The page adds a version-bounded fourth: "Before v2.1.281, some sessions, such as those on Amazon - Bedrock or with telemetry disabled, read `CLAUDE.md` files only. On those versions, update Claude - Code." Its remedy for every one of these sessions is the shim: "To give Claude your `AGENTS.md` in - any of these sessions, import it from a `CLAUDE.md`." -- **Basis**: , "When AGENTS.md support is unavailable", - fetched as raw markdown 2026-10-01 (50,074 bytes). The 2.1.281 entry of the - [changelog](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) reads "Changed - AGENTS.md support to also work on Amazon Bedrock, Google Vertex AI, Microsoft Foundry, LLM - gateways, and sessions with telemetry disabled". -- **As of**: 2026-10-01, Claude Code 2.1.287. -- **Recheck trigger**: a condition is added to or removed from that list, the list's opening claim - changes, the pre-2.1.281 sentence changes, or the remedy sentence changes. - -**No hooks setting is an availability condition.** "The settings and flags that stop installed mods, -such as `disableAllHooks`, `--bare`, and `--safe-mode`, don't stop built-in mods", and -`agents-md@builtin` is one ([mods overview](https://code.claude.com/docs/en/plugins/mods/overview), -"Mods built into Claude Code", fetched 2026-10-01). Under `allowManagedHooksOnly`, "Mods built into -Claude Code keep running" ([settings reference](https://code.claude.com/docs/en/settings-reference), -"What runs under `allowManagedHooksOnly`", fetched 2026-10-01). The -[plugin's README](https://github.com/anthropics/claude-code/blob/main/mods/agents-md/README.md), -commit `2282079d6ac8`, agrees: "No hooks setting or CLI mode turns it off". Where the engine -loads no instruction files at all (`--bare` without `--add-dir`, `--safe-mode`, -`CLAUDE_CODE_DISABLE_CLAUDE_MDS`), it finds no `AGENTS.md` either, per that README. Recheck when the -overview's built-in sentence or the README's paragraph changes. - -**All three are resolvable or bounded.** Condition 1 is a version comparison, and so is the -pre-2.1.281 provider and telemetry gap. Condition 2 is whether `agents-md@builtin` is enabled. -Condition 3 is resolvable only as "this may be that session", so it is the one that commonly stays -unresolved. Resolve what is resolvable first, because a condition known TRUE settles it with no +## Availability: three conditions, any one of which ends the question + +Our checks treat `AGENTS.md` support as unavailable, so that no `AGENTS.md` is read natively at +any path under any mode, when any one of these holds for the session: + +1. the Claude Code version is below v2.1.277, or below v2.1.281 in a session that fetches no + feature flags from Anthropic (a third-party provider such as Amazon Bedrock, or telemetry + disabled); +2. the built-in `agents-md` plugin is disabled in `/plugin`; +3. it may be the first session after an upgrade from a version without `AGENTS.md` support. + +For such a session the remedy we recommend is a `CLAUDE.md` that imports the `AGENTS.md`. + +- **Pointer**: for when `AGENTS.md` support is unavailable and the import remedy, see + . +- **As of**: 2026-10-01 +- **Recheck trigger**: a condition is added to or removed from that section's list, a version + floor in it moves, or its remedy changes. + +**No hooks setting is an availability condition.** `agents-md@builtin` is a built-in mod, so our +checks never count `disableAllHooks`, `allowManagedHooksOnly`, `--bare` or `--safe-mode` against +it. A session that loads no instruction files at all (`--bare` without `--add-dir`, `--safe-mode`, +`CLAUDE_CODE_DISABLE_CLAUDE_MDS`) reads no `AGENTS.md` either. + +- **Pointer**: for which settings stop built-in mods, see the "Mods built into Claude Code" + section of and the "What runs under + `allowManagedHooksOnly`" section of ; for the + plugin's own statement, its + [README](https://github.com/anthropics/claude-code/blob/main/mods/agents-md/README.md) at commit + `2282079d6ac8`. +- **As of**: 2026-10-01 +- **Recheck trigger**: the overview's built-in-mods section or the README's settings paragraph + changes. + +**Two of the three resolve from what this plugin already reads.** Condition 1 is a version +comparison, plus the provider and telemetry configuration on a CLI between the two floors. +Condition 2 is whether the built-in `agents-md` plugin is enabled, which the permission-and-settings +lanes already inventory. Condition 3 resolves only as "this may be that session", so it is the one +that commonly stays unresolved. Do not treat the whole question as unresolvable because one +condition is: resolve what is resolvable first, because a condition known TRUE settles it with no further work. ## The mode: four values, and where the value lives -- **Claim**: the **Project instructions** setting takes one of four values. - `claude-md-or-agents-md` reads "Your `CLAUDE.md` files, or your `AGENTS.md` files when you have no - `CLAUDE.md` or `CLAUDE.local.md` in your working directory or above it. **This is the default**". - `claude-md-and-agents-md` reads both, "each directory's `CLAUDE.md` files first and its - `AGENTS.md` after them", and "Claude Code skips an `AGENTS.md` it has already loaded, so one that - your `CLAUDE.md` imports or symlinks to isn't read twice". `claude-md` reads "Your `CLAUDE.md` - files only". `managed-only` reads "Only your organization's managed `CLAUDE.md` and auto memory at - launch", and under it "every `AGENTS.md`" is left out. - **So two of the four values make an `AGENTS.md` unread regardless of displacement**, and - displacement is a condition of the default value alone. -- **Basis**: , "Choose which instruction files load", value - table, fetched 2026-10-01. -- **As of**: 2026-10-01. -- **Recheck trigger**: a value is added, removed or renamed, the default moves, or a value's - description changes which files it loads. +Our checks read the **Project instructions** setting as one of four values: -**Where the value lives**, which is what a check reads rather than the `/config` panel: +- `claude-md-or-agents-md`, the default: an `AGENTS.md` loads only where no `CLAUDE.md` or + `CLAUDE.local.md` sits in the working directory or above it (displacement). +- `claude-md-and-agents-md`: both load, `CLAUDE.md` first per directory, and an `AGENTS.md` that a + `CLAUDE.md` already imports or symlinks counts once. +- `claude-md` and `managed-only`: no `AGENTS.md` loads. -- **Claim**: "Add it under the built-in `agents-md` plugin's ID in `pluginConfigs`, in - `~/.claude/settings.json`, a `--settings` file, or managed settings. **Claude Code ignores it in - project and local settings files.**" The documented shape is the `instructionFiles` option under - the `agents-md@builtin` key. -- **Basis**: the same section, its settings paragraph and JSON example. -- **As of**: 2026-10-01. +**So two of the four values make an `AGENTS.md` unread regardless of displacement**, and +displacement is a condition of the default value alone. + +- **Pointer**: for the values and what each loads, see + . +- **As of**: 2026-09-21 +- **Recheck trigger**: a value is added, removed or renamed, the default moves, or a value changes + which files it loads. + +**Where the value lives**, which is what a check reads rather than the `/config` panel: our checks +read the `instructionFiles` option under the `agents-md@builtin` key of `pluginConfigs`, from user +settings (`~/.claude/settings.json`), a `--settings` file or managed settings, and ignore the key +in project and local settings files. + +- **Pointer**: for where the setting is honored and its JSON shape, see + . +- **As of**: 2026-09-21 - **Recheck trigger**: the option key or plugin id changes, the honored scope set changes, or the setting becomes readable from project or local settings. diff --git a/plugins/harness-config/skills/audit-instructions/SKILL.md b/plugins/harness-config/skills/audit-instructions/SKILL.md index 9e09900c1d..7781f6051e 100644 --- a/plugins/harness-config/skills/audit-instructions/SKILL.md +++ b/plugins/harness-config/skills/audit-instructions/SKILL.md @@ -19,7 +19,7 @@ locally-owned instruction surfaces, cites each finding to current official promp it by how confident the evidence can be, and packages proposed removals or rewrites as a human-gated diff, so instruction surfaces shrink as models get better instead of only ever growing. -The check catalog, covering the checks I1–I35, their evidence tier, authority tag, severity, +The check catalog, covering the checks I1–I37, their evidence tier, authority tag, severity, per-surface applicability, and the `OPINION`-tier enablement policy, lives in [reference/criteria.md](reference/criteria.md); the deterministic pre-scan is `${CLAUDE_PLUGIN_ROOT}/skills/audit-instructions/scripts/instruction-scan.sh`. @@ -36,10 +36,11 @@ Diffs are proposed artifacts. A clean audit is a valid outcome. mechanical. `Write` stays for the Phase D persist and for lane reports under `runs///lanes/`, and `Bash` for the pre-scans, `lane-runs.sh`, and the `run-state.sh` lease writes under `runs/`. Either can mutate a file this skill has already read, so -this is an instruction-held contract with a narrowed accident surface, not an enforced one. Never describe it to an operator as a guarantee. The restriction clears -on their next message (, frontmatter reference, fetched -2026-08-12), so whoever accepts a diff can apply it. `audit-prompting-postures` carries the identical -declaration and the identical caveat, because the two state the same contract. +this is an instruction-held contract with a narrowed accident surface, not an enforced one. Never +describe it to an operator as a guarantee. We treat the restriction as clearing on the operator's +next message, so whoever accepts a diff can apply it (Pointer: the +[skills frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference). +As of: 2026-08-12. Recheck trigger: that section changes how long the restriction lasts). `audit-prompting-postures` carries the identical declaration and caveat: both state one contract. ## Scope boundary (route out) @@ -61,7 +62,7 @@ concerns its siblings already cover, so route rather than re-answer: On **memory-layer surfaces** (CLAUDE.md, a natively read AGENTS.md, CLAUDE.local.md, `.claude/rules/`, and `rules/` under the user root Phase A resolves), this skill runs only the -model-era checks I6–I35. It never runs or reports the hygiene checks I1–I5 (line-necessity, length, +model-era checks I6–I37. It never runs or reports the hygiene checks I1–I5 (line-necessity, length, placement, inferable content, rule-to-hook) on these surfaces; that layer belongs to the `harness-memory` plugin. When it is installed, route memory-layer hygiene to its `audit` skill; when it is not, emit a single one-line pointer to the official CLAUDE.md include/exclude guidance @@ -247,23 +248,22 @@ One partition rule sizes lanes, a token budget: no lane cap and no line-count co dispatch first**: run `lane-runs.sh partition` over the inventoried files, then dispatch those lanes. A lane's budget is **0.25 of the lane model's own context window**, leaving the rest for the -catalog, the lane brief, the lane's reasoning, and its report, at **3.5 bytes per token** -(Anthropic's glossary: "a token approximately represents 3.5 English characters", fetched -2026-09-28 from ; recheck when that entry -changes or a lane overflows its window on a supported model). Non-ASCII text runs more bytes per -character, so the estimate errs toward smaller lanes. Never state the budget as a line count: the -line figure is derived per run from the bytes per line measured over the in-scope files. +catalog, the lane brief, the lane's reasoning, and its report, at **3.5 bytes per token**, our +setting (Pointer: for the characters-per-token estimate, see the [glossary entry for +tokens](https://platform.claude.com/docs/en/about-claude/glossary#tokens). As of: 2026-09-28. +Recheck trigger: that entry changes, or a lane overflows its window on a supported model). +Non-ASCII text runs more bytes per character, so the estimate errs toward smaller lanes. Never state +the budget as a line count: the line figure is derived per run from the bytes per line measured over the in-scope files. -**Claim:** a subagent's context window is sized by its own model, not the parent's. **Basis:** - (model field section). **As of:** 2026-09-29. -**Recheck:** when that page's model or context-window wording changes. +We size a lane by the context window of the model the lane runs on, never the parent's (Pointer: +[subagents: choose a model](https://code.claude.com/docs/en/sub-agents#choose-a-model). As of: +2026-09-29. Recheck trigger: that section changes how a subagent's model or window is set). `` is the `--window-tokens` value: the context window in tokens of the model the -lane runs on, per (fetched 2026-09-29). When the -lane's model resolves to no documented window, pass 200000, the smaller standard window: a smaller -budget only adds lanes. **Claim:** 200000 is the smallest documented window. **Basis:** the -model-config page above. **As of:** 2026-09-29. **Recheck:** when that page documents a smaller -window for a supported model. +lane runs on. When the lane's model resolves to no documented window, pass 200000, our setting for +the smallest standard window: a smaller budget only adds lanes (Pointer: +[model configuration: extended context](https://code.claude.com/docs/en/model-config#extended-context). +As of: 2026-09-29. Recheck trigger: that page documents a smaller window for a supported model). Partition deterministically, feeding every in-scope file as `\t\t`, where the group is its plugin (or the memory layer) and the unit is its skill (or the file itself): @@ -317,9 +317,9 @@ plugin's `block-hook-bypass` guard exempts by design, never through a shell redi carried in a variable or through inline Python, which that guard blocks because it cannot resolve the target. A redirect to a literal absolute path under the host temp tree is exempt only when `CLAUDE_PROJECT_DIR` names a project root not itself under a temp tree, so in a temp-rooted checkout -(a CI clone, a test fixture) the Write tool is the only route. Verified 2026-09-12 against `plugins/guardrails/hooks/block-hook-bypass.sh` -(`_bbh_temp_default_applies` and the scope note in `block_bypass`) and `plugins/guardrails/README.md` -("`block-hook-bypass` ships two scratch roots exempt"); recheck when the guardrails plugin changes +(a CI clone, a test fixture) the Write tool is the only route. Pointer: +`plugins/guardrails/hooks/block-hook-bypass.sh` (`_bbh_temp_default_applies` and the scope note in +`block_bypass`) and `plugins/guardrails/README.md`. Recheck trigger: the guardrails plugin changes that guard's exemption set or its block message. ## Phase B2: Cross-surface conflict pass diff --git a/plugins/harness-config/skills/audit-instructions/evals/evals.json b/plugins/harness-config/skills/audit-instructions/evals/evals.json index 18c459b909..ac7d3e6219 100644 --- a/plugins/harness-config/skills/audit-instructions/evals/evals.json +++ b/plugins/harness-config/skills/audit-instructions/evals/evals.json @@ -17,11 +17,11 @@ "id": 2, "name": "scope-boundary-routes-out", "prompt": "/harness-config:audit-instructions is my CLAUDE.md too long, and can you prune the always-loaded hygiene lines out of it?", - "expected_output": "Recognizes memory-layer hygiene — line budget, whether CLAUDE.md is too long, pruning always-loaded content — as the harness-memory plugin's audit skill's concern. On memory-layer surfaces (CLAUDE.md, a natively read AGENTS.md, CLAUDE.local.md, .claude/rules) this skill runs only the model-era checks I6-I35 and routes the I1-I5 hygiene findings to harness-memory:audit when that plugin is installed, falling back to the official CLAUDE.md include/exclude guidance pointer when it is not, rather than answering the hygiene question itself.", + "expected_output": "Recognizes memory-layer hygiene (line budget, whether CLAUDE.md is too long, pruning always-loaded content) as the harness-memory plugin's audit skill's concern. On memory-layer surfaces (CLAUDE.md, a natively read AGENTS.md, CLAUDE.local.md, .claude/rules) this skill runs only the model-era checks I6-I37 and routes the I1-I5 hygiene findings to harness-memory:audit when that plugin is installed, falling back to the official CLAUDE.md include/exclude guidance pointer when it is not, rather than answering the hygiene question itself.", "files": [], "expectations": [ "Routes memory-layer hygiene (line budget, 'CLAUDE.md too long', pruning) to the harness-memory plugin's audit skill when installed", - "On memory-layer surfaces runs only the model-era checks I6-I35 (I16, I19, and I22 only under --opinion), not the I1-I5 hygiene checks", + "On memory-layer surfaces runs only the model-era checks I6-I37 (I16, I19, and I22 only under --opinion), not the I1-I5 hygiene checks", "Falls back to the official CLAUDE.md include/exclude guidance pointer when harness-memory is not installed instead of silently skipping" ] }, @@ -368,6 +368,32 @@ }, { "id": 29, + "name": "i36-tool-discouraging-language-fires-on-sonnet-5-5-only", + "prompt": "/harness-config:audit-instructions skills --target-model sonnet-5-5. One of my research skills tells the model: \"Keep tool calls to a minimum and answer from what you already know unless a tool is unavoidable.\" The skill has a web search tool. Then run the same audit with --target-model opus-5-5.", + "expected_output": "On the sonnet-5-5 run, reports an I36 tool-discouraging-language finding on that line, behavioral tier, with a proposed change that replaces the line with the Sonnet 5.5 guide's targeted search steer, read live at the row's pointer. The finding is a proposal verified in Phase C, not a confident removal, and it is not reported as I28. On the opus-5-5 run, I36 is inert and the report lists it as skipped-for-target, because the row is scoped to sonnet-5-5 by exact token match.", + "files": [], + "expectations": [ + "Reports I36 on the sonnet-5-5 run against the tool-discouraging line", + "Proposes replacing the line with the guide's targeted search steer from the pointer, not a blanket 'always use the tool' default", + "Does not report the line under I28", + "Lists I36 as skipped-for-target on the opus-5-5 run instead of reporting it" + ] + }, + { + "id": 30, + "name": "i37-per-tool-result-countdown-versus-narrow-hook", + "prompt": "/harness-config:audit-instructions hooks --target-model sonnet-5-5. My interactive setup has two PostToolUse hooks. The first matches every tool and returns additionalContext with the remaining token budget after each call. The second matches only Bash and returns additionalContext only when the command exits non-zero.", + "expected_output": "Reports an I37 finding on the first hook: it puts a countdown into the model's context after every tool result in an interactive session, which raises the chance that a real mid-turn user message is taken for an injected one. The proposed change makes the text rarer, for example a narrower matcher or a threshold, or drops the countdown if nothing acts on it. The second hook is not a finding: it fires on a narrow condition. The run does not report the first hook under I23, which is scoped to Fable targets.", + "files": [], + "expectations": [ + "Reports I37 on the hook that returns a countdown after every tool call", + "Proposes making that text rarer or dropping the countdown, rather than deleting the hook outright", + "Does not flag the Bash hook that fires only on a non-zero exit", + "Does not report I23 on a sonnet-5-5 target" + ] + }, + { + "id": 31, "name": "content-in-claude-md-routes-to-migrate", "prompt": "/harness-config:audit-instructions claude-md, treating evals/fixtures/content-in-claude.md relative to the skill directory as the repository's tracked root CLAUDE.md, and the same text as a tracked .claude/CLAUDE.md; the repository has no AGENTS.md.", "expected_output": "Besides any model-era findings on the file, the report's Routing subsection carries one AGENTS.md content-home advisory for the root CLAUDE.md, because its content is not the single line `@AGENTS.md`, and none for the .claude/CLAUDE.md, which migrate's plan does not cover. The advisory points to `/instruction-placement:migrate plan` when that plugin is installed, and otherwise names the move as a repository-wide change for the operator. It says the CLAUDE.md stays as the one-line `@AGENTS.md` shim and that migrate's `cutover-check` decides when shims come out, notes that Claude-specific text goes to `.claude/rules/.md` with a `paths:` glob, and routes progressive-disclosure questions to `/docs-hygiene:audit-progressive-disclosure`. It has no Finding ID, severity, or diff and does not appear in the findings table.", diff --git a/plugins/harness-config/skills/audit-instructions/reference/criteria.md b/plugins/harness-config/skills/audit-instructions/reference/criteria.md index 3ca30be1ac..ddca14a013 100644 --- a/plugins/harness-config/skills/audit-instructions/reference/criteria.md +++ b/plugins/harness-config/skills/audit-instructions/reference/criteria.md @@ -1,5 +1,5 @@ --- -version: 1.25.0 +version: 1.26.0 last-updated: 2026-10-01 --- @@ -46,6 +46,8 @@ Look up a specific check by ID: run `grep -n '^### I:'` over this file. - [I33: Sibling-file meta-commentary](#i33-sibling-file-meta-commentary) - [I34: Maintainer rationale inside model-facing YAML comments](#i34-maintainer-rationale-inside-model-facing-yaml-comments) - [I35: Settled-answers instruction where later steps revise earlier ones](#i35-settled-answers-instruction-where-later-steps-revise-earlier-ones) + - [I36: Tool-discouraging language](#i36-tool-discouraging-language) + - [I37: Harness text after every tool result](#i37-harness-text-after-every-tool-result) - [Stopping condition](#stopping-condition) - [Out-of-catalog defects](#out-of-catalog-defects) - [AGENTS.md content-home advisory](#agentsmd-content-home-advisory) @@ -67,12 +69,14 @@ which the catalog-wide trigger already covers. One staleness event fires the who check that noticed it. Model-specific pages, the per-model prompting guides under Sources, are superseded on each model generation. -**Per-row verification stamps.** A row that restates a volatile upstream *literal*, such as a level -name, a model range, or a type predicate, additionally carries the four-part record that claim needs: the -claim, its basis, an as-of date, and a recheck trigger naming an observable event (the shape is -`docs/conventions/upstream-drift/README.md` in this monorepo; in a standalone install the four -parts, not the path, are the requirement). A row that restates nothing and only points at its page -carries no stamp, because a pointer cannot go stale. +**Per-row verification stamps.** A row whose firing rule acts on a volatile upstream *literal*, +such as a level name, a model range, or a type predicate, keeps that literal as its own setting and +additionally carries the record that decision needs: the decision in our words, a pointer to the +exact upstream section, an as-of date, and a recheck trigger naming an observable event, with no +upstream text (the shape is `docs/conventions/upstream-drift/README.md` in this monorepo; in a +standalone install those parts, not the path, are the requirement). A row whose Source line only +names its page and section, and whose firing rule acts on no such literal, carries no stamp: the +catalog-wide trigger already covers it. A per-row stamp **supplements** the catalog-wide trigger above; it never replaces or narrows it. The catalog trigger already fires every row on any Sources change, so a per-row trigger adds no @@ -84,9 +88,13 @@ because it is the wider one and a staleness signal is not something to resolve b narrower authority. The requirement **binds on touch**, per the convention above: rows predating this rule keep their -citations as they are and adopt the four parts the next time they change. A missing stamp on an +citations as they are and adopt the record parts the next time they change. A missing stamp on an older row is therefore not itself a defect in this catalog. +**Source lines name, never quote.** A row's Source line names the page and the section that +documents its mechanic; the firing rule above it is this catalog's decision, in its own words, and +works without a live fetch. Read the section's wording at the page. + **Admission.** A row's observable must be **anchored to text that is present**. A check detects a passage a surface actually contains: either what it says, or an attribute it lacks while saying it. I6 (a prohibition carrying no rationale marker) and I7 (a request stating no motivation) are the @@ -121,6 +129,21 @@ model name in prose. Promotion to fleet-wide (unscoped) happens only through the authoritative model-agnostic upstream doc states the claim, OR multiple model guides converge on it. Unannotated checks are model-agnostic and always fire. +**Tokens of models that are no longer current.** A scope token stays while Claude Code can still +put a session on its model by fallback. On 2026-10-01 that held for `opus-5`, `opus-4-8` and +`sonnet-5`, so no row drops one. `fable-5` is not in this set: we treat Fable 5 as a current model, +since Claude Code still offers it for selection. Each row scoped to a token in this set, or to +`fable-5`, carries a "Re-justified" line saying why its scope neither drops nor widens to another +current model. + +- **Pointer**: for the fallback targets, see + [model configuration: automatic model fallback](https://code.claude.com/docs/en/model-config#automatic-model-fallback); + for how Fable 5 is selected, see + [model configuration: work with Fable](https://code.claude.com/docs/en/model-config#work-with-fable). +- **As of**: 2026-10-01 +- **Recheck trigger**: a model leaves Claude Code's model page, which retires its token in every + row that names it, or Fable 5 stops being selectable there. + **`OPINION` enablement.** Enablement attaches to *detection*, never to advice, and splits on what a rule does: @@ -144,10 +167,10 @@ non-memory surfaces (skill bodies, agent definitions, hook instruction text, out memory-layer surfaces (CLAUDE.md, a natively read AGENTS.md, CLAUDE.local.md, `.claude/rules/`, `~/.claude/rules/`) their findings route to the `harness-memory` plugin's `audit` skill when it is installed, and fall back to the official include/exclude guidance (I1–I5 source below) when it is -not. Checks I6–I12, I16–I28, I30, and I35 apply to all surfaces. I15 also applies to all surfaces, -but its unit is a pair, so Phase B2 answers it rather than a per-surface lane. I13, I14, I29, I31, -I32, I33, and I34 name narrower surface sets in their own rows, and a lane runs each only on the -surfaces its row names. +not. Checks I6–I12, I16–I28, I30, and I35–I37 apply to all surfaces. I15 also applies to all +surfaces, but its unit is a pair, so Phase B2 answers it rather than a per-surface lane. I13, I14, +I29, I31, I32, I33, and I34 name narrower surface sets in their own rows, and a lane runs each only +on the surfaces its row names. ## Sources @@ -158,35 +181,30 @@ surfaces its row names. - Prompting Claude Opus 5: -- The bundled `claude-api` skill's model-migration reference. The `fable-5-1` widenings below were - taken from Claude Code 2.1.258 (sections Migrating to Claude Fable 5.1 and Migrating to Claude - Fable 5.1 from Claude Fable 5). Re-read 2026-09-28 from the skill inside Claude Code 2.1.282: - those sections are still present, the guide adds `## Ground the migration with an eval` (the - 2.1.260 refresh, which also moved the Go, Java, and C# samples onto current-generation model - ids), and the bundled `prompt-audit` guide still runs Steps 0–7 over Groups 1–4. Changelog - 2.1.283 is the next release that names `prompt-audit`, so this stamp does not claim the guide is - byte-identical past 2.1.282. **Recheck trigger:** publication of a Fable 5.1 prompting guide, - which replaces this basis and joins this list in its place, or a release note that changes - `prompt-audit` or the model-migration sections this catalog cites. +- Prompting Claude Fable 5.1: + . + The `fable-5-1` widenings below rest on it, in place of the bundled `claude-api` skill's + model-migration reference they first cited; that skill's record lives in + [bundled-claude-api.md](bundled-claude-api.md). +- What's new in Claude Fable 5.1 (its refusal categories, for I10): + +- Prompting Claude Sonnet 5.5 (the `sonnet-5-5` rows and widenings below): + - Prompting Claude Opus 5.5: -- Getting the most out of Opus 5.5 in Claude and Claude Code (vendor blog, published 2026-09-22, - which carries the Opus 5.5 guide's chat-scoped claims to saved Claude Code instructions; a dated - post, cited where it adds that reach and otherwise corroborating, so the citing rows keep the - `ANTHROPIC-DOCS` Authority of the guide above): - -- Prompting Claude Sonnet 5.5: - +- Getting the most out of Opus 5.5 in Claude and Claude Code (vendor blog, published 2026-09-22). + Not a pointer (correlate only): the docs pointer is Prompting Claude Opus 5.5 above, and the + citing rows keep that guide's `ANTHROPIC-DOCS` Authority; + correlate with - Prompting Claude Sonnet 5: - Prompting Claude Opus 4.8: - The new rules of context engineering for Claude 5 generation models (vendor blog, published - 2026-07-24, which corroborates I6 from the model-delta side and I15 from the reasoning-cost side; - a dated post, static once published, so a recheck is expected to find it unchanged; it - corroborates rather than defines, so the rows citing it keep the `ANTHROPIC-DOCS` Authority of - their primary documentation sources and the closed four-value Authority set above is unchanged): - + 2026-07-24). Not a pointer (correlate only): the docs pointer is Prompting best practices above, + for I6 and I15, whose rows keep the `ANTHROPIC-DOCS` Authority of their documentation sources, so + the closed four-value Authority set above is unchanged; + correlate with - Memory (CLAUDE.md, a natively read AGENTS.md, rules, auto memory): - The `.claude` directory: @@ -202,8 +220,9 @@ surfaces its row names. - Introducing Claude Fable 5 and Claude Mythos 5 (which models carry the safety classifiers): - Thinking (the sanctioned reasoning-visibility path, the `display` field, the thinking-block - round-trip protocol, the models that reject a thinking-disable outright, and what a thinking or - effort change does to the cache prefix): + round-trip protocol, the per-model table of accepted `thinking` values including `between_tools`, + the models that reject a thinking-disable outright or non-default sampling parameters, and what a + thinking or effort change does to the cache prefix): - Steering thinking (the turn-validation relaxation, and the models that still enforce a leading thinking block): @@ -211,12 +230,14 @@ surfaces its row names. - Troubleshooting thinking (the per-request 400s, the models the effort restriction covers, and the internal-tag leakage a don't-think directive worsens): -- Model migration guide (the model ranges over which manual extended thinking is rejected, and - the ranges over which non-default sampling parameters are rejected): - -- What's new in Claude Sonnet 5 (the sampling-parameter constraint's arrival on the Sonnet class, - the new tokenizer, and the launch behavior changes): - +- Model migration guides. The migration-guide URL is an index of per-model guides + (); I25 cites the Fable + and Mythos guide's section on migrating from Claude Opus 5: + . + The model ranges that reject manual extended thinking (I17-c) and non-default sampling + parameters (I25) are read from Thinking, above. +- Claude Sonnet 5 model page, "Good to know" (the sampling-parameter constraint on the Sonnet class): + - Effort (the levels, `high`'s equivalence to omitting the parameter, the carry-over sweep advice, and where thinking may not be disabled): @@ -227,10 +248,12 @@ surfaces its row names. - Settings (the `effortLevel` value set): - Environment variables (`CLAUDE_CODE_EFFORT_LEVEL`, `MAX_THINKING_TOKENS`, and `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` with the models it reaches): - ; read it verbatim per the + ; read it whole per the [fetch route](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/upstream-drift/README.md#reading-the-basis-the-fetch-route), because a summarizing fetch truncates this page well before these rows - Prompt caching (what belongs to the cache key): +- Errors (what Claude Code sends, or reports, when thinking is off at an effort level the model + refuses with it): - CLI reference (`claude doctor` and the other terminal forms): - Subagents (what loads into a subagent at startup): @@ -267,8 +290,7 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface on I5's terms. This bar asks whether removal would change behavior *today*; a protected rail's removal changes behavior only on the occasion it was written for, which this criterion cannot observe. -- **Source:** best-practices, "For each line, ask: *Would removing this cause Claude to make - mistakes?* If not, cut it." +- **Source:** best-practices, "Write an effective CLAUDE.md" (the per-line removal test). ### I2: Length and skimmability @@ -277,8 +299,7 @@ Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface - **Detect:** a surface long or dense enough that its own rules start getting ignored; the tell is the model breaking a rule the file contains. - **Remediate:** prune, split into path-scoped rules or skills, tighten structure. -- **Source:** best-practices, "Bloated CLAUDE.md files cause Claude to ignore your actual - instructions." +- **Source:** best-practices, "Write an effective CLAUDE.md" (file length and ignored rules). ### I3: Broad-applicability placement @@ -316,7 +337,7 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface trades guaranteed presence for a deferral the agent cannot rely on. Name a destination the agent itself reaches, meaning a skill the agent's definition **invokes at runtime** or text kept in the definition, and never a `paths:`-scoped rule. **A `skills:` preload is not such a destination**: - the full content of each listed skill is injected into every dispatch of that agent, so the + we treat every preloaded skill's whole body as loaded into each dispatch of that agent, so the content is resident for every unrelated use exactly as it was in the definition, and the move defers nothing. That is the same disqualification `@path` imports carry above. When the agent has no conditional runtime invocation to move the content to, report that no safe deferral is @@ -325,19 +346,18 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface - **Adjacent axis:** this check is load *timing*. Definition-site *locality*, an instruction sitting away from the thing it governs, is I16, and an instruction can be correctly deferred here and still misplaced there. -- **Source:** best-practices, "only include things that apply broadly. For domain knowledge or - workflows that are only relevant sometimes, use skills instead."; memory, "splitting into `@path` - imports helps organization but doesn't reduce context, since imported files load at launch"; - context-window, "What survives compaction", for the per-destination cost; skills, "skill - descriptions are loaded into context so Claude knows what's available, but full skill content only - loads when invoked", the combined `description` and `when_to_use` text "is truncated at 1,536 - characters in the skill listing to reduce context usage", and "Plugin skills are not affected by - `skillOverrides`."; subagents, on a skill named in an agent's `skills:` field, "The full content - of each listed skill is injected into the subagent's context at startup." All quoted spans - verified 2026-08-31 against (the - 1,536 cap and invocation-control quotes) and (the - preload quote); recheck trigger: a fetch of either page no longer carrying its quoted span - re-derives this check's Remediate mechanics. +- **Source:** best-practices, "Write an effective CLAUDE.md" (broad applicability, skills for + sometimes-relevant content); memory, "Import additional files" (imports load at launch); + context-window, "What survives compaction", for the per-destination cost; skills, "Frontmatter + reference" (description loading, the listing truncation this check keeps as its own 1,536 + setting, and the invocation-control fields) and "Override skill visibility from settings" + (`skillOverrides` and plugin skills); subagents, the `skills:` field (preloaded skill content). + Pointer: , + and + . As of: 2026-08-31. Recheck trigger: either page + changes the listing cap, which field keeps a description out of context, whether + `skillOverrides` reaches plugin skills, or what a `skills:` preload injects; that re-derives + this check's Remediate mechanics. ### I4: Inferable or redundant content @@ -354,8 +374,8 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface the hold and its class; propose compression in place instead. The register is non-exhaustive, so a candidate absent from it is judged on this criterion's normal terms, never deleted *because* it is absent. -- **Source:** best-practices include/exclude table, which excludes "Anything Claude can figure out by - reading code" and "Standard language conventions Claude already knows." +- **Source:** best-practices, "Write an effective CLAUDE.md", the include/exclude table (its exclude + column). ### I5: Rule-to-hook or delete @@ -372,8 +392,8 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `info` · Surfaces: "The model already does this" is the weakest possible evidence against a rail whose absence is unrecoverable, and the hook conversion stays available: converting a protected rule to a deterministic mechanism is a remediation, deleting it is not. -- **Source:** best-practices, "If Claude already does something correctly without the instruction, - delete it or convert it to a hook." +- **Source:** best-practices, "Avoid common failure patterns" (already-followed rules) and "Set up + hooks". ### I6: Bare prohibition to positive reframing @@ -388,18 +408,10 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface - **Remediate:** reframe positively, stating what to do instead, as the primary fix. Where a genuine hard "never" survives, keep it but add its rationale (see I7) as the fallback. - **Bounded by:** the **Stopping condition** below, which is enabled by default. -- **Source:** prompting best-practices, "Tell Claude what to do instead of what not to do." - Corroborated from the model-delta side at the context-engineering blog, under "Then and now" in - the paired "Then: Give Claude rules" / "Now: Let Claude use judgment" headings. The bare - prohibition quoted below was a guardrail for older models, since "newer models have better - judgment and can handle these decisions well without explicit rules", and its shipped - replacement is an instance of this row's remediation shape: "Write code that reads like the - surrounding code: match its comment density, naming, and idiom." - - - "In code: default to writing no comments. Never write multi-paragraph docstrings or - multi-line comment blocks — one short line max." - +- **Source:** prompting best-practices, "Control the format of responses" (positive framing). + Corroborated from the model-delta side (correlate with the context-engineering blog under + Sources, its "Then and now" section, whose worked example is an instance of this row's + remediation shape). ### I7: Reason with the request @@ -410,12 +422,8 @@ so this fires for every target model. - **Detect:** an instruction that states a request with no intent or motivation attached. - **Remediate:** add the why: the model connects the task to relevant context instead of inferring intent on its own. -- **Source:** Fable 5 guide, "Give the reason, not only the request": "Claude Fable 5 tends to - perform better when it understands the intent behind a request." Convergent model-agnostic - source (the gate-meeting one): Prompting best practices, "Add context to improve performance": - "Providing context or motivation behind your instructions, such as explaining to Claude why - such behavior is important, can help Claude better understand your goals and deliver more - targeted responses." +- **Source:** Fable 5 guide, "Give the reason, not only the request". Convergent model-agnostic + source (the gate-meeting one): Prompting best practices, "Add context to improve performance". ### I8: Model-era re-audit @@ -427,12 +435,10 @@ convergent model guides meet the gate for each (see the rows); the base row's de worked instance keeps a `fable-5` scope of its own. **Base row** · Unscoped. Promotion gate MET on 2026-08-08: the model-agnostic best-practices page -states the claim under its all-current-models framing, "Prefer general instructions over -prescriptive steps. A prompt like 'think thoroughly' often produces better reasoning than a -hand-written step-by-step plan. Claude's reasoning frequently exceeds what a human would -prescribe." **The worked instance +states the claim under its all-current-models framing (section "Leverage thinking & interleaved +thinking capabilities", on general instructions over prescriptive steps). **The worked instance below keeps a `fable-5` scope of its own**, because its basis is Fable-specific and the Opus guides -run the other way. +disagree with it. - **Detect:** prior-model workarounds and over-prescriptive step lists: instructions enumerating behaviors a current model handles from a brief instruction, or scaffolding that pins an approach. @@ -440,11 +446,11 @@ run the other way. only on a `fable-5` or `fable-5-1` resolved target: a delegation throttle**, meaning a cap on concurrent workers, a one-at-a-time rule, or an instruction to block until each subagent returns before dispatching the next, where the surface's own ground for it is that subagent handling is - unreliable. The Fable 5 guide runs the other way, asking for readier dispatch and asynchronous - orchestrator-to-worker communication, so a throttle resting on that premise is the generic case - with a name on it. On `opus-5` and `opus-4-8` targets this instance is inert, not merely - unattested: those guides recommend delegation caps and note fewer spawns by default, so a - throttle there is the recommended shape rather than a workaround. **A cap carrying its own + unreliable. On a Fable target we treat a throttle resting on that premise as the generic case + with a name on it. The Fable 5 guide ("Parallel subagents") and the Opus 5 and Opus 4.8 guides + ("Controlling subagent spawning") disagree on throttling subagent dispatch, so on `opus-5` and + `opus-4-8` targets this instance is inert, not merely unattested: there we treat a throttle as + the recommended shape rather than a workaround. **A cap carrying its own non-model rationale is not this instance.** Reviewability of returns, rate limits, cost, or shared mutable state each justify a bound on their own terms, and that justification is the surface's to make, not this row's to override. @@ -452,32 +458,24 @@ run the other way. consequential removal needs a closed watch, an editorial one does not) that default performance holds or improves. - **Bounded by:** the **Stopping condition** below, which is enabled by default. -- **Source:** prompting best practices, "Leverage thinking & interleaved thinking capabilities", - the prefer-general-instructions statement quoted above (the gate-meeting, model-agnostic one). - Convergent model guide: Fable 5, "Skills developed for prior models are often too prescriptive - for Claude Fable 5 and can degrade output quality." The worked instance's basis is the same - guide, "Parallel subagents": "Claude Fable 5 dispatches parallel subagents more readily than - prior models. Use subagents frequently … and prefer asynchronous communication between - orchestrator and subagents over blocking until each subagent returns"; its Opus counter-basis is - the Opus 5 guide's "Controlling subagent spawning" ("set deterministic caps … keep spawn counts - low") and the Opus 4.8 guide's "Controlling subagent spawning" ("tends to spawn fewer subagents - by default"). -- **The general principle, and why it is cited separately.** The migration-framed sentences above +- **Source:** prompting best practices, "Leverage thinking & interleaved thinking capabilities" + (the gate-meeting, model-agnostic one). Convergent model guide: Fable 5, "Recommended + scaffolding changes" (prior-model skills as too prescriptive). The worked instance's basis is + the same guide, "Parallel subagents"; its Opus counter-basis is the Opus 5 guide's and the Opus + 4.8 guide's "Controlling subagent spawning". +- **The general principle, and why it is cited separately.** The migration-framed sources above point a reader at what looks like leftover prior-model scaffolding, walking straight past freshly authored over-enumeration, which is the same defect with no legacy provenance to - recognize it by. The principle is also stated on its own in the Fable 5 guide's "Strong - instruction following": "Instruction-following is - improved enough that you can steer most behaviors with a brief instruction rather than enumerating - each behavior by name," and that guide's own worked case is a *newly written* brevity - instruction replacing a list of patterns, not a migration. **Age is not an element of this row.** - Detect over-enumeration wherever it was written and whenever. + recognize it by. The Fable 5 guide's "Strong instruction following" states the principle on its + own (a brief instruction in place of an enumeration), and that section's worked case is a *newly + written* instruction, not a migration. **Age is not an element of this row.** Detect + over-enumeration wherever it was written and whenever. **Row I8-a: instructed self-check removal** · Tier `behavioral` · Model scope: `opus-5`. -- **Detect:** instructions telling the model to re-check work it already checks: "double-check - your answer," "re-verify before responding," "include a final verification step for any - non-trivial task," "use a subagent to verify". This includes legacy harness scaffolding that adds - separate verification steps. +- **Detect:** instructions telling the model to re-check work it already checks. The pre-scan's + stems are `double-check`, `re-verify`, `final verification step`, and `subagent to verify`. Older + harness text that bolts an extra checking pass onto every task counts too. - **Classify by reviewer INDEPENDENCE, not invocation source:** architected independent review, meaning a fresh-context reviewer blind to the producing rationale or a different-vendor verifier, is NOT a finding; the anti-pattern is the instructed self-check. **Carve-out lanes (never @@ -486,20 +484,30 @@ run the other way. - **Remediate:** propose removal; verify per Deletion tiers (a consequential removal needs a closed watch, an editorial one does not). - **Bounded by:** the **Stopping condition** below. -- **Source:** Opus 5 guide, "Task scope and over-verification", which says to remove explicit - verification instructions: they "cause over-verification on Claude Opus 5, and removing them - reduces wasted tokens with no loss in quality"; "Self-correction", which says to avoid instructing - re-checks it already performs. +- **Source:** Opus 5 guide, "Task scope and over-verification" and "Self-correction". - **The independence carve-out is corroborated by a second guide, and the scope does not move.** The - Fable 5 guide reaches the same line from the opposite direction: it asks for self-verification to - be made explicit on long runs, and states that "separate, fresh-context verifier subagents tend to - outperform self-critique" ("Recommended scaffolding changes"). Read without that sentence, the two - guides look contradictory, remove verification instructions versus add them, and a reader has to - resolve it alone. They are not: the anti-pattern is the instructed **self**-check, and an - architected independent verifier is the thing the Fable 5 guide is asking for. **This does not meet - the promotion gate**, because the gate wants a second guide stating this row's *detection* claim, - that verification instructions cause over-verification, and the Fable 5 guide states no such - thing. The scope annotation stands; only the carve-out gains a second source. + Fable 5 guide's "Recommended scaffolding changes" section reaches the same line from the opposite + direction, on making verification explicit on long runs with fresh-context verifiers. Read + without that section, the two guides look contradictory, remove verification instructions versus + add them, and a reader has to resolve it alone. We read them as consistent: the anti-pattern is + the instructed **self**-check, and an architected independent verifier is what the Fable 5 + section covers. **This does not meet the promotion gate**, because the gate wants a second guide + stating this row's *detection* claim, that verification instructions cause over-verification, + and the Fable 5 guide states no such thing. The scope annotation stands; only the carve-out gains + a second source. +- **Re-justified 2026-10-01 against the current models:** the scope stays `opus-5` (see "Tokens of + models that are no longer current"). Our probe of the Fable 5.1, Opus 5.5 and Sonnet 5.5 guides + that day (each read whole as raw markdown; no artifact stored) found no statement of the + detection claim. We do not widen to `sonnet-5-5`: a finding there would remove the check the + posture catalog's P13 asks a code-changing component to carry, and P13 cites this section. + Pointer: [Sonnet 5.5 guide, verification on coding + tasks](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#verification-on-coding-tasks); + the probe covered the whole of [Prompting Claude Fable + 5.1](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1) + and [Prompting Claude Opus + 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5), + so it names no section. As of: 2026-10-01. Recheck trigger: a current model's guide stating that + verification instructions cause over-verification. **Row I8-b: conservative-reporting detection** · Tier `behavioral`. Unscoped. Promotion gate MET on its second arm: a second model guide, the Sonnet 5 one, states the same claim about the shared @@ -527,21 +535,20 @@ target model. self-filter is genuinely wanted, keep it but **state the bar concretely**, as an enumerable test the reader can decide a novel finding against, rather than a qualitative term. - **Bounded by:** the **Stopping condition** below, which is enabled by default. -- **Source:** Opus 5 guide, "Code review and bug-finding": if the prompt says "only report - high-severity issues" or "be conservative," the model "may follow that instruction literally and - report less; ask it to report everything and filter in a separate pass instead." Convergent - second model guide (the gate-meeting one): Sonnet 5 guide, "Code review harnesses", on the same - three phrases: "Claude Sonnet 5 may follow that instruction more faithfully than earlier models - did: it may investigate the code just as thoroughly, identify the bugs, and then not report - findings it judges to be below your stated bar." The third trigger phrase, **"don't nitpick", - which appears nowhere in the Opus 5 guide**, is stated in the Sonnet 5 guide and again in the - Opus 4.8 guide ("Code review harnesses"), which repeats the claim, the coverage prompt, and the - concrete-bar half near-verbatim for its own model; the Sonnet 5 guide states that half as: "be - concrete about where the bar is rather than using qualitative terms like 'important'" (the - upstream page double-quotes the word). (Opus 4.8 corroboration verified 2026-08-08 against that - guide's raw `.md`, 15,905 bytes, MD5 `6b9db5b784ad6a7b2e6307c1481b8be9`; the gate was already met - without it. The "nowhere in the Opus 5 guide" negative re-verified 2026-08-08 against the Opus 5 - guide's raw `.md`: zero occurrences of "nitpick".) +- **Source:** the gate rests on two model guides that name the trigger phrases: the Opus 5 guide + for the first two and the report-everything-then-filter remediation, and the Sonnet 5 guide for + all three and the Remediate line's concrete-bar half. The Opus 4.8 guide corroborates the third + phrase, "don't nitpick"; the gate was met without it. The Opus 5 guide does not use that phrase. + Pointer: [Opus 5 guide, capability + improvements](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5#capability-improvements) + (its code review and bug-finding item), [Sonnet 5 guide, code review + harnesses](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5#code-review-harnesses) + and [Opus 4.8 guide, code review + harnesses](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-4-8#code-review-harnesses). + As of: 2026-08-08 (our probes, no artifact stored: the Opus 4.8 guide's raw `.md`, 15,905 bytes, + MD5 `6b9db5b784ad6a7b2e6307c1481b8be9`; the Opus 5 guide's raw `.md`, zero occurrences of + "nitpick"). Recheck trigger: either gate guide drops the trigger phrases from its section, or + the Opus 5 guide starts using "nitpick". **Row I8-c: don't-think / don't-reason directive** · Tier `behavioral` · Model scope: `opus-5`, `opus-5-5`, `sonnet-5-5`. @@ -550,25 +557,22 @@ claim (see Source), and it is a model-agnostic feature page, the surface where a appear, yet it names Claude Opus 5 anyway. The promotion gate stays unmet by upstream's own choice, on the same reasoning I10 applies to a declined widening. -- **Detect:** instructions telling the model not to think or not to reason. With thinking - disabled these increase internal-tag leakage. Also flag tag-hygiene rules that name thinking - tags specifically (less effective than the general form). -- **Where it shows, and why it outlives the turn.** The leakage is "most commonly on tool-heavy - workloads such as search", so a surface governing a tool-driven lane is where to look, and the - damage is not confined to the response that leaks: "A leaked tool call never runs, and in agentic - loops the leaked text stays in the conversation history, so later turns are affected as well." - The page states the history effect, not this consequence. Read here, that means an autonomous - lane carries the poisoned turn forward as context. +- **Detect:** a directive forbidding the model to think or reason. We treat such a directive as + making internal-tag leakage worse when thinking is off. Also flag tag-hygiene rules that name + thinking tags specifically (less effective than the general form). +- **Where it shows, and why it outlives the turn.** We look first at a surface governing a + tool-driven lane, such as search, since the troubleshooting page places the leakage on tool-heavy + workloads, and we treat the damage as not confined to the response that leaks: a tool call + emitted as text is never executed, and that text remains in the loop's history afterwards. The + page states the history effect, not this consequence. Read here, that means an autonomous lane + carries the poisoned turn forward as context. - **Remediate:** remove the directive; where output-tag hygiene is genuinely needed, use the general "internal or system XML tags" phrasing. - **Bounded by:** the **Stopping condition** below, which is enabled by default. -- **Source:** Opus 5 guide, "Running with thinking disabled": "If your system prompt contains a - rule instructing the model not to think or not to reason, remove it; that kind of instruction - increases tag leakage"; naming thinking tags is "less effective than the general form." - Corroborated at troubleshooting thinking, "Tool calls or XML tags appear in the text output", - which reaches the same claim from the symptom side, "System-prompt rules instructing the model - not to think or not to reason increase the tag leakage", and is the source of the condition and - consequence above. **Verified 2026-08-04** against that page, fetched as raw markdown. +- **Source:** Opus 5 guide, "Running with thinking disabled" (the removal, and the general form over + naming thinking tags). Corroborated at troubleshooting thinking, "Tool calls or XML tags appear in + the text output", which reaches the same claim from the symptom side and is the source of the + condition and consequence above. **As of 2026-08-04** (that page read as raw markdown). **Recheck trigger:** a second model name appearing beside Claude Opus 5 in either section that states the claim: the Opus 5 guide's "Running with thinking disabled", or this page's "Tool calls or XML tags appear in the text output". A new name re-opens the scoping question, not the @@ -578,27 +582,44 @@ choice, on the same reasoning I10 applies to a declined widening. enumerates the models that do *not* leak, so those two sections are the whole of what there is to re-read. - **Widened to `opus-5-5` on 2026-09-23:** the Opus 5.5 guide, "Prompts written for thinking - disabled", says to re-test the thinking-disabled mitigations and to "remove the no-thinking rule - either way", since thinking is always on for that model. On `opus-5-5` the Detect clause's - leakage premise does not apply; the finding stands on that removal instruction alone. **Verified - 2026-09-23** against the guide's raw `.md` (28,311 bytes, MD5 + disabled", prescribes removing the no-thinking rule on that model, where thinking is always on. + On `opus-5-5` the Detect clause's leakage premise does not apply; the finding stands on that + removal alone. **As of 2026-09-23** (our probe: the guide's raw `.md`, 28,311 bytes, MD5 `fb3bff7f41e20fbbb71be78770edb8cb`). **Recheck trigger:** that section ceasing to prescribe the removal. -- **Widened to `sonnet-5-5` on 2026-10-01:** the Sonnet 5.5 guide, "Running without up-front - thinking", says that when a request sends `between_tools` any instruction not to think should go, - because it makes internal XML tags in the visible output more likely; "Calibrate effort" adds - that asking in the system prompt for less thinking "doesn't reliably reduce its thinking", so - lowering effort is the control. The leakage premise applies where `between_tools` is in use; - elsewhere on `sonnet-5-5` the finding stands on the second statement. **Verified 2026-10-01** - against the guide's raw `.md` (27,412 bytes, MD5 `2bcb67cc9f72b68e8823f197034c06d6`). **Recheck - trigger:** either section dropping its statement. +- **Considered for `sonnet-5-5` on 2026-10-01 and declined.** We read the Sonnet 5.5 guide as + covering this directive only under the `between_tools` thinking setting (pointer below), and no + Claude Code surface exposes that setting, so a session reading a Claude Code surface never runs + under it. On a `sonnet-5-5` target the row stays inert. Ownership of `between_tools` text is + split once, the same way in this row and in I17's third arm: a prompt in application source that + sends `between_tools` is the bundled `claude-api` skill's (SKILL.md, "Boundary"); instruction + text that prescribes a `between_tools` request is in this catalog only through I17's third arm, + which flags the refused pairings and not this directive. + Pointer: [Sonnet 5.5 guide, running without up-front + thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#running-without-up-front-thinking). + For the Claude Code negative, our probe read [model + configuration](https://code.claude.com/docs/en/model-config), + [settings](https://code.claude.com/docs/en/settings), [settings + reference](https://code.claude.com/docs/en/settings-reference) and [environment + variables](https://code.claude.com/docs/en/env-vars#variables) whole as raw markdown and found + no control for that setting (no artifact stored). As of: 2026-10-01. Recheck trigger: Claude + Code gains a `between_tools` control, or the Sonnet 5.5 guide states the claim outside + `between_tools`. +- **Re-justified 2026-10-01 against the current models:** `opus-5` stays (see "Tokens of models + that are no longer current"), and both sections the trigger above watches still name Claude Opus + 5 alone. Pointer: [Opus 5 guide, running with thinking + disabled](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5#running-with-thinking-disabled) + and [troubleshooting thinking: tool calls or XML tags appear in the text + output](https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting#tool-calls-or-xml-in-text) + (the rendered page gives this heading the id `tool-calls-or-xml-in-text`). As of: 2026-10-01. + Recheck trigger: the same as this row's trigger above. **Row I8-d: short-turn assumptions** · Tier `behavioral` · Model scope: `fable-5, fable-5-1`. - **Detect:** instruction text resting on the premise that a turn is short: a directive to answer quickly or keep turns brief, or any required progress rhythm pinned to a turn rather than to the - work. Individual requests now run for many minutes at higher effort and autonomous runs for hours, - so a rhythm calibrated to the old turn length fires as noise on work that has not reached a + work. We treat a single high-effort request as lasting minutes and an autonomous run as lasting + hours, so a rhythm calibrated to the old turn length fires as noise on work that has not reached a reportable boundary, and it interrupts precisely the long uninterrupted runs the model is being used for. - **The forced interim-status cadence shape is owned fleet-wide by I8-e**, which is unscoped since @@ -625,14 +646,27 @@ choice, on the same reasoning I10 applies to a declined widening. client configuration rather than instruction content, so it is not audited here and no row claims it; a surface whose *instruction text* prescribes a short client timeout is the shape that would reach this catalog, and none is attested. -- **Source:** Fable 5 guide, "Longer turns by default": "Individual requests on hard tasks can run - for many minutes at higher effort settings … and autonomous runs can extend for hours. This is one - of the largest shifts teams encounter when adjusting to Claude Fable 5." -- **Widened to `fable-5-1` on 2026-09-03:** the bundled `claude-api` skill's model-migration - reference (Claude Code 2.1.258), sections Migrating to Claude Fable 5.1 and Migrating to Claude - Fable 5.1 from Claude Fable 5, restates this behavior for Claude Fable 5.1 and states that Fable 5 - prompt guidance carries over. **Recheck trigger:** publication of a Fable 5.1 prompting guide, - whose statement of this claim replaces this basis and joins `## Sources`. +- **Source:** Fable 5 guide, "Longer turns by default". +- **Widened to `fable-5-1` on 2026-09-03**, first on the bundled `claude-api` skill's + model-migration reference. That basis's trigger fired when the Fable 5.1 guide was published. The + widening now rests on our rule that a Fable 5 row applies to Fable 5.1 unless the Fable 5.1 guide + names a difference on the row's subject, and on our reading of that guide on 2026-10-01 (read + whole as raw markdown; no artifact stored): no section names a turn-length difference. Pointer: + [Prompting Claude Fable + 5.1](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1), + the text above its first heading, which carries no heading id on the rendered page, so the link + is to the page. As of: 2026-10-01. Recheck trigger: the Fable 5.1 guide naming a turn-length + difference from Fable 5, or its opening changing what it says about Fable 5 prompts. +- **Re-justified 2026-10-01 against the current models:** `fable-5` stays, since Fable 5 is a + current model (see "Tokens of models that are no longer current"). Our probe of the Opus 5.5 and + Sonnet 5.5 guides that day (each + read whole as raw markdown; no artifact stored) found no statement of the short-turn claim, so + the row does not widen to them. Pointer: [Prompting Claude Opus + 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5) + and [Prompting Claude Sonnet + 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5), + whole pages, since the probe is a negative over every section. As of: 2026-10-01. Recheck + trigger: either guide gains a section on turn length. **Row I8-e: forced interim-status cadence** · Tier `behavioral`. Unscoped. Promotion gate MET: two model guides state the claim (see Source). @@ -642,17 +676,17 @@ Unscoped: two model guides state the claim (see Source), which meets the promoti **This row owns the cadence shape on every target**; I8-d cedes it (see that row) so the two report one finding per line rather than two. -- **Detect:** an instruction requiring interim status output on a fixed mechanical interval. The - guide's own example is "After every 3 tool calls, summarize progress"; equivalents this row also - reaches, the catalog's rather than the guide's, are "check in after each file" and "post an update - every N minutes". The subject is the *forced rhythm*, not the reporting: an instruction to report +- **Detect:** an instruction requiring interim status output on a fixed mechanical interval, such + as a summary after every few tool calls, "check in after each file", or "post an update every N + minutes". The subject is the *forced rhythm*, not the reporting: an instruction to report at a genuine work boundary (a phase completing, a gate failing) pins to the work and is not a finding. - **Remediate:** name the guarantee the cadence was protecting, that the user can see progress or that a long run stays interruptible, and either state that outcome and let the model meet it, or move it to a mechanism rather than an instructed rhythm. Where the *content* of native updates is - miscalibrated rather than absent, describe what a good update contains and give examples; that is - the upstream remediation and it does not reintroduce a cadence. Verify per Deletion tiers (a + miscalibrated rather than absent, describe what a good update contains and give examples; that + remediation comes from the Sonnet 5 guide and does not reintroduce a cadence. Verify per + Deletion tiers (a consequential removal needs a closed watch, an editorial one does not). - **Bounded by:** the **Stopping condition** below, which is enabled by default. - **Must NOT flag: a cadence carrying its own explicit observability or interruptibility @@ -667,33 +701,27 @@ report one finding per line rather than two. chapter counter-steering it, or a verification record quoting it, on the same audience test I8-b applies. This catalog's own detect text is the canonical instance; the deterministic pre-scan seeds no pattern for this row, so it carries no fixtures of its own. -- **Source:** Sonnet 5 guide, "User-facing progress updates": "Claude Sonnet 5 provides regular, - higher-quality updates to the user throughout long agentic traces. If you've added scaffolding to - force interim status messages ("After every 3 tool calls, summarize progress"), try removing it." - That guide also supplies the Remediate line's second half: where updates are miscalibrated, - "explicitly describe what these updates should look like in the prompt and provide examples." - Convergent second model guide (the gate-meeting one): Opus 4.8 guide, "User-facing progress - updates": "Claude Opus 4.8 provides more regular, higher-quality updates to the user throughout - long agentic traces. If you've added scaffolding to force interim status messages ("After every 3 - tool calls, summarize progress"), try removing it." -- **Verified 2026-08-08** against both gate sources, fetched as raw markdown: the Sonnet 5 guide - (15,864 bytes, MD5 `6d23959f0ed226feb06bf20c314029e3`, byte-identical to 2026-07-29 and - 2026-08-04 captures) and the Opus 4.8 guide (15,905 bytes, MD5 - `6b9db5b784ad6a7b2e6307c1481b8be9`). The 2026-08-04 **verified negative** on the Fable 5 guide - was re-verified 2026-08-08 against that guide's raw `.md` and is retained as a reading of that - guide on which the scope does not rest. That negative: "Longer turns by default" prescribes only - client-side adjustments, no section prescribes removing instructed status cadence, and "Create a - send-to-user tool" runs the other way. **Recheck trigger:** either gate source ceasing to - prescribe removal of forced status scaffolding, which re-opens the scoping question. +- **Source:** Sonnet 5 guide, "User-facing progress updates" (removing forced status scaffolding, + and the Remediate line's second half). Convergent second model guide (the gate-meeting one): Opus + 4.8 guide, "User-facing progress updates". +- **As of 2026-08-08** (our probe of both gate sources as raw markdown: the Sonnet 5 guide, 15,864 + bytes, MD5 `6d23959f0ed226feb06bf20c314029e3`, byte-identical to 2026-07-29 and 2026-08-04 + captures; the Opus 4.8 guide, 15,905 bytes, MD5 `6b9db5b784ad6a7b2e6307c1481b8be9`). The + 2026-08-04 **verified negative** on the Fable 5 guide was re-checked 2026-08-08 against that + guide's raw `.md` and is retained as our reading of that guide, on which the scope does not rest: + no section of it prescribes removing instructed status cadence ("Longer turns by default" covers + client-side adjustments only, and "Create a send-to-user tool" points the other way). **Recheck + trigger:** either gate source ceasing to prescribe removal of forced status scaffolding, which + re-opens the scoping question. **Row I8-f: think-carefully steer** · Tier `behavioral` · Model scope: `opus-5-5`. **The gate is -unmet by contradiction, not only by absence:** the base row's model-agnostic source recommends -"think thoroughly" over a hand-written plan, so an unscoped row would contradict it. +unmet by contradiction, not only by absence:** the base row's model-agnostic source favors a +general think-thoroughly prompt over a hand-written plan, so an unscoped row would contradict it. - **Detect:** a standing instruction telling the model to think carefully, hard, deeply, or step by step before answering, or an `ultrathink`-style keyword written into a saved instruction rather - than typed for one piece of work. The Opus 5.5 model always thinks and decides how much itself, - so the line adds latency without a clear quality gain. + than typed for one piece of work. We treat the line as adding latency without a clear quality + gain on Opus 5.5, where thinking is always on and the model sets its own depth. - **Remediate:** delete the line. Where the intent was more or less depth, change effort, the documented control; where a fast answer to simple questions was the intent, lower effort first, and add an "Answer directly." line only after measuring quality with it, since less thinking can @@ -702,19 +730,21 @@ unmet by contradiction, not only by absence:** the base row's model-agnostic sou human reader; a per-invocation keyword the human types (I21 owns effort pinning); a document *about* the pattern, on the audience test I8-b applies. - **Bounded by:** the **Stopping condition** below, which is enabled by default. -- **Source:** Opus 5.5 guide, "Thinking instructions in chat system prompts": for instructions - "that tell Claude to think carefully before answering, consider removing them"; removing such a - line "made replies start sooner, with no clear decline in the quality of the reply"; "Calibrate - effort" for the lower-effort-first remediation. The guide scopes the claim to chat system prompts; - the vendor usage guide (Sources) extends it to "your prompts and your saved instructions" and - names effort as the Claude Code control. **Verified 2026-09-23** against the guide's raw `.md` - (hash as in I8-c). **Recheck trigger:** a second model guide stating the claim, which re-opens - the scoping question, or the section dropping it. -- **Not widened to `sonnet-5-5`:** the Sonnet 5.5 guide does not state this claim. For JSON answers - that need a few steps of working out it prescribes a think-first line, "Think the problem through - before you answer.", so the removal premise does not carry ("Reasoning tasks with JSON output"; - verified 2026-10-01, hash as in I8-c's `sonnet-5-5` widening). **Recheck trigger:** that section dropping the line, or a later guide on the model - stating the removal claim. +- **Source:** Opus 5.5 guide, "Thinking instructions in chat system prompts" (the removal and its + effect); "Calibrate effort" for the lower-effort-first remediation. The guide scopes the claim to + chat system prompts; we apply it to saved Claude Code instructions as well, with effort as the + Claude Code control (correlate with the vendor usage guide under Sources). **As of 2026-09-23** + (our probe: the guide's raw `.md`, hash as in I8-c). **Recheck trigger:** a second model + guide stating the claim, which re-opens the scoping question, or the section dropping it. +- **Cross-reference, Sonnet 5.5.** The row stays scoped to `opus-5-5` and is inert on a + `sonnet-5-5` target. Its firing set on an `opus-5-5` target is unchanged: a surface that also + serves Sonnet 5.5 is still flagged. When the flagged line asks for reasoning on a task answered in + JSON, the finding adds that the line may be serving Sonnet 5.5, and proposes splitting it per + model rather than deleting it; for that model's position, follow the pointer. + Pointer: [Sonnet 5.5 + guide, reasoning tasks with JSON + output](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#reasoning-tasks-with-json-output). + As of: 2026-10-01. Recheck trigger: either guide changes its position on thinking instructions. ### I9: Example hygiene @@ -729,61 +759,77 @@ Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `info` · Surfaces: `argument-hint`. That destination clause is **`OPINION`-derived**, since no official page states it, so it rides this check's enablement and severity per the `OPINION` policy above, is labeled as `OPINION` in the finding, and is never fix-applied. -- **Source:** prompting best-practices, "Use examples effectively": examples are "one of the most - reliable ways to steer Claude's output format, tone, and structure"; keep them diverse enough - "that Claude doesn't pick up unintended patterns." +- **Source:** prompting best-practices, "Use examples effectively" (examples as format, tone and + structure steering, and diversity against unintended patterns). ### I10: Reasoning-echo directives Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `error` · Surfaces: all · Model scope: -`fable-5, fable-5-1, opus-5-5, sonnet-5-5` (the cited refusal category is documented per model; promotion gate -unmet). +`fable-5, fable-5-1, opus-5-5, sonnet-5-5` (the cited refusal category is documented per model; +promotion gate unmet, see "Unscoping considered" below). -- **Detect:** instructions telling the model to show, echo, transcribe, or explain its internal - reasoning as response text. The deterministic pre-scan marks show-your-thinking phrasing. +- **Detect:** instructions asking the model to put its internal reasoning into the reply itself, + whether by showing, repeating, writing out, or narrating it. The deterministic pre-scan marks + show-your-thinking phrasing. - **Remediate:** remove them; where reasoning visibility is genuinely needed, read structured `thinking` blocks through the surface that already exposes them: in Claude Code, `Ctrl+O` verbose mode and the `showThinkingSummaries: true` setting (model configuration); on the API, `display: "summarized"` (Thinking). A send-to-user tool remains the path when the reasoning has to reach the user as ordinary response text. -- **Source:** Fable 5 guide: such instructions "can trigger the `reasoning_extraction` refusal - category on Claude Fable 5, causing elevated fallbacks." Corroborated by the Thinking page, which - states the same refusal for the same model: "On Claude Fable 5, a request that attempts to elicit - the model's internal reasoning as part of the response text can be refused with - `stop_details.category: "reasoning_extraction"`." That second citation does **not** move the - promotion gate: its own section names both Claude Fable 5 and Claude Mythos 5 for the adjacent - raw-chain-of-thought property, then names Fable 5 alone for the refusal. That is a - sentence-adjacent chance to widen, declined, so the narrower scope is deliberate. +- **Source:** Fable 5 guide, "Recommended scaffolding changes" (the `reasoning_extraction` refusal + category on Claude Fable 5). Corroborated by the Thinking page, which documents the same refusal + category for the same model. That second citation does **not** move the promotion gate: its own + section names both Claude Fable 5 and Claude Mythos 5 for the adjacent raw-chain-of-thought + property, then names Fable 5 alone for the refusal. That is a sentence-adjacent chance to widen, + declined, so the narrower scope is deliberate. **`Model scope: fable-5` is positively sourced**, - in two statements each taken from the page that owns its half. The page that owns Mythos 5 states - the exclusion at the level of the whole classifier set: "Claude Fable 5 includes safety - classifiers that can decline certain requests. Claude Mythos 5 does not include these classifiers, - so this section applies to Claude Fable 5 only" ([Introducing Claude Fable 5 and Claude Mythos + in two statements each taken from the page that owns its half. The page that owns Mythos 5 scopes + the whole classifier set to Claude Fable 5 and excludes Claude Mythos 5 ([Introducing Claude + Fable 5 and Claude Mythos 5](https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5), - fetched 2026-08-03). Refusals and fallback places this row's category inside that set, listing + read 2026-08-03). Refusals and fallback places this row's category inside that set, listing `reasoning_extraction` among the classifier categories a refusal reports. Which models carry the classifier set is a per-model fact and moves, so the introducing page joins `## Sources`: the catalog-wide trigger then fires this row whenever that page changes, and no narrower per-row trigger is owed. -- **Widened to `fable-5-1` on 2026-09-03:** the bundled `claude-api` skill's model-migration - reference (Claude Code 2.1.258), sections Migrating to Claude Fable 5.1 and Migrating to Claude - Fable 5.1 from Claude Fable 5, restates this behavior for Claude Fable 5.1 and states that Fable 5 - prompt guidance carries over. **Recheck trigger:** publication of a Fable 5.1 prompting guide, - whose statement of this claim replaces this basis and joins `## Sources`. -- **Widened to `opus-5-5` on 2026-09-23:** the Opus 5.5 guide, "Safeguard refusals": "Requests - that push the model to reproduce its internal reasoning in the response text can be declined with - the `reasoning_extraction` category, which is new if you're coming from Claude Opus 5." The same - section notes that server-side fallback returns these declines to the caller instead of - retrying them on a fallback model. Remediate there as above, or ask for what the reader needs instead, such - as the rationale in a few sentences. **Verified 2026-09-23** against the guide's raw `.md` (hash - as in I8-c). **Recheck trigger:** that section dropping the category. -- **Widened to `sonnet-5-5` on 2026-10-01:** the Sonnet 5.5 guide, "Safeguard refusals", lists - `reasoning_extraction` among the categories a refusal reports: "the request asks the model to - reproduce its internal reasoning in the response text." Remediate there as above. **Verified - 2026-10-01** against the guide's raw `.md` (hash as in I8-c's `sonnet-5-5` widening). **Recheck - trigger:** that section dropping the category. +- **Widened to `fable-5-1` on 2026-09-03**, first on the bundled `claude-api` skill's + model-migration reference. That basis's trigger fired when the Fable 5.1 guide was published. + The widening now rests on the refusals section of Fable 5.1's own model page. Pointer: [What's + new in Claude Fable 5.1, refusals, fallback, and + billing](https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1#refusals-fallback-and-billing). + As of: 2026-10-01. Recheck trigger: that section stops covering the `reasoning_extraction` + category for Fable 5.1. +- **Widened to `opus-5-5` on 2026-09-23:** the Opus 5.5 guide, "Safeguard refusals", documents the + `reasoning_extraction` category for that model, and how server-side fallback handles those + declines. Remediate there as above, or ask for what the reader needs instead, such as the + rationale in a few sentences. **As of 2026-09-23** (our probe: the guide's raw `.md`, hash as in + I8-c). **Recheck trigger:** that section dropping the category. +- **Widened to `sonnet-5-5` on 2026-10-01.** On a `sonnet-5-5` target, fire on the same Detect and + remediate as above. Pointer: [Sonnet 5.5 guide, safeguard + refusals](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#safeguard-refusals). + As of: 2026-10-01. Recheck trigger: that section drops the `reasoning_extraction` category. +- **Unscoping considered on 2026-10-01 and declined.** Four model guides now name the category + (Sonnet 5.5, Opus 5.5, Fable 5.1 and Fable 5), which on its face meets the convergent-guides arm + of the promotion gate. We keep the row scoped because we read the category as a per-model + classifier, not a behavior every model shares, so an unscoped row would flag text on targets + where the line draws no refusal. Pointer: the four places that name the category, [Sonnet 5.5 + guide, safeguard + refusals](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#safeguard-refusals), + [Opus 5.5 guide, safeguard + refusals](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#safeguard-refusals), + [Prompting Claude Fable + 5.1](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1) + (the note above its first heading, which carries no heading id on the rendered page, so the link + is to the page) and [Fable 5 guide, recommended scaffolding + changes](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5#recommended-scaffolding-changes); + for the categories themselves, see [refusals and fallback: refusal + response](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#refusal-response). + As of: 2026-10-01. Recheck trigger: a model-agnostic page stating that every current model + declines reasoning extraction. +- **Re-justified 2026-10-01 against the current models:** `fable-5` stays, since Fable 5 is a + current model (see "Tokens of models that are no longer current"). ### I11: CLI over MCP where equivalent @@ -793,8 +839,7 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `info` · Surfaces: the surface's concern is context cost rather than a capability the MCP server uniquely provides. - **Remediate:** prefer the CLI for the equivalent operation; keep the MCP path where it adds capability. -- **Source:** best-practices, "CLI tools are the most context-efficient way to interact with - external services." +- **Source:** best-practices, "Use CLI tools" (context efficiency). ### I12: Stale or misattributed harness-capability claim @@ -831,9 +876,8 @@ Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface their `settings.json` chains or which hooks they wire. That describes a file, not the harness, so it is not a harness-behavior claim. When the evidence in hand shows it false, report it under [Out-of-catalog defects](#out-of-catalog-defects). -- **Source:** CLI reference, "Print read-only installation and settings diagnostics from the - terminal without starting a session … For the in-session setup checkup that can also apply - fixes, run `/doctor`." +- **Source:** CLI reference, "CLI commands", the `claude doctor` row (terminal diagnostics versus + the in-session `/doctor`). ### I13: Citation form that does not load @@ -861,9 +905,9 @@ Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface package scope (`@anthropic-ai/…`), a decorator, an email address, or a `@username` handle. A backticked `` `@path` ``, which the import parser skips by design and which is the documented way to mention a path without importing it. A path cited without an `@` at all. -- **Source:** memory, "CLAUDE.md files can import additional files using `@path/to/import` - syntax", against skills, where supporting files are instead referenced "so Claude knows what - each file contains and when to load it" and no import syntax is defined. +- **Source:** memory, "Import additional files" (import syntax on the CLAUDE.md family), against + skills, "Add supporting files", where supporting files are referenced for Claude to load when + needed and no import syntax is defined. ### I14: Retrieval of an already-loaded surface @@ -930,21 +974,26 @@ skill bodies. current disk contents. The startup copy is a snapshot taken at launch; another process can have changed the file since, and a pre-edit read cut on the grounds that "it is already in context" produces a patch against stale text. A rule restated in a - delegation prompt for the built-in Explore and Plan agents, which are documented as the only - subagents that skip `CLAUDE.md` and have no per-agent setting to change that. -- **Source:** subagents, "What loads at startup": a non-fork subagent's initial context contains - "every level of the CLAUDE.md hierarchy the main conversation loads, including - `~/.claude/CLAUDE.md`, project rules, `CLAUDE.local.md`, and managed policy files." The qualifier - *the main conversation loads* is what bounds this check: memory documents lazy loading for - "path-specific rules or lazy-loaded files in subdirectories", so those are outside the guarantee. - memory, on `@path` imports, is what puts an imported supporting document inside it: "Imported - files are expanded and loaded into context at launch alongside the CLAUDE.md that references - them", and "Imported files can recursively import other files, with a maximum depth of four hops"; - memory's `AGENTS.md` guidance names exactly such an import as what carries an `AGENTS.md` into a - session that cannot read it directly, and it is the portable form on Windows, where the symlink - alternative needs elevation (code.claude.com/docs/en/memory, "When AGENTS.md support is - unavailable" and "Share one file with other coding tools"; verified 2026-09-19; recheck trigger: - that section stops naming the import, or a release note names `AGENTS.md` loading). + delegation prompt for a subagent that starts without the user, project and local `CLAUDE.md`: + the built-in Explore and Plan agents, and a custom agent whose definition sets + `omitClaudeMd: true`. In such an agent's own body, an instruction to read the root `CLAUDE.md` + is likewise not redundant. Pointer: [What loads at startup](https://code.claude.com/docs/en/sub-agents#what-loads-at-startup) + and the `omitClaudeMd` row of [Supported frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields). + The page's system-prompt section disagrees with both on whether `CLAUDE.md` still loads under + that field. As of: 2026-10-01. Recheck trigger: the page reconciles the two, or the field is renamed. +- **Source:** subagents, "What loads at startup": we treat a non-fork subagent's initial context as + holding every level of the CLAUDE.md hierarchy *the main conversation loads*, except an agent + whose definition sets `omitClaudeMd: true` (see the Must NOT flag above), and that qualifier + is what bounds this check: memory documents lazy loading for path-specific rules and + subdirectory files ("How CLAUDE.md files load"), so those are outside the guarantee. memory, + "Import additional files", is what puts an imported supporting document inside it: imports load + at launch with the file that references them and recurse up to four hops. We treat such an + import as what carries an `AGENTS.md` into a session that cannot read it directly, and as the + portable form on Windows, where the symlink alternative needs elevation. Pointer: + and + . As of: + 2026-09-19. Recheck trigger: that section stops naming the import, or a release note names + `AGENTS.md` loading. ### I15: Cross-surface instruction conflict @@ -991,13 +1040,10 @@ a **pair**, so this row is answered by Phase B2 rather than by a per-surface lan objects. An absolute carrying its own exception beside a directive presupposing that exception. A pair one of whose sides already states which wins. The full set with worked instances is in [conflict-criteria.md](conflict-criteria.md). -- **Source:** memory, "If two rules contradict each other, Claude may pick one arbitrarily", which - is why an unarbitrated pair is a finding rather than a stylistic note. Corroborated at the - context-engineering blog, "Unhobbling Claude", where Anthropic's own system prompt, skills, and - user requests clash, "several conflicting messages in a single request like 'leave documentation - as appropriate,' or 'DO NOT add comments'", with the cost stated even for the resolved case: - "Claude must think more carefully about these overlapping and conflicting messages before - deciding what to do." So a conflict taxes reasoning even when no arbitrary pick occurs. +- **Source:** memory, "Write effective instructions" (contradicting rules and an arbitrary pick), + which is why an unarbitrated pair is a finding rather than a stylistic note. We also treat a + conflict as taxing reasoning even when no arbitrary pick occurs (correlate with the + context-engineering blog under Sources, its "Unhobbling Claude" section). ### I16: Definition-site locality @@ -1028,145 +1074,197 @@ by `--opinion`. Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `error` · Surfaces: all. Unscoped. Promotion gate MET: the claim is stated on a model-agnostic feature page, not in a model guide. -**The model ranges are Detect conditions, not a `Model scope` annotation.** One source says the -restriction "applies to Claude Opus 5 and later models"; the other names Fable 5, Mythos 5 and -Mythos Preview. The annotation's exact-string matching has no range form. Annotating `opus-5` would +**The model ranges are Detect conditions, not a `Model scope` annotation.** The pairing arm is an +Opus 5 range; the outright-disable arm is a set spanning three families (Opus, Sonnet, and Fable +and Mythos), listed in the second arm below. The annotation's exact-string matching has no range +form. Annotating `opus-5` would make the row inert on the next generation while the restriction still holds, and no single annotation spans two disjoint families at once. I20 handles a model range the same way. Each row below carries its own decisive source; they share a subject, not a citation. -**Base row: the configurations the model rejects.** Two arms with different shapes: a pairing that -fails only at the top of the effort ladder, and a disable that fails at every level. Both are -`error`, since both are a rejected request. +**Base row: the configurations the model rejects.** Three arms with different shapes: a pairing +that fails only at the top of the effort ladder, a disable that fails at every level, and an API +thinking setting that refuses the top levels and per-turn effort changes. All three are `error`, +since each prescribes a request the model refuses. - **Detect:** a surface that recommends, documents, or sets a **thinking-disable surface**, meaning `MAX_THINKING_TOKENS=0`, `alwaysThinkingEnabled: false`, the `/config` global toggle, the `Alt+T` / `Option+T` session toggle, or API `thinking: {"type": "disabled"}`, together with `xhigh` or `max` effort, on Claude Opus 5 or a later model. Both operands are configuration - literals, so a surface prescribing both publishes a per-request 400 that nothing recovers. + literals. Fire on the API form and on the harness forms alike: in neither does the prescribed + level reach the model, so the finding and its remediation are the same for both. What Claude + Code sends in place of the prescribed level is behind the pointer. + Pointer: for the harness outcome, see + [model configuration: extended thinking](https://code.claude.com/docs/en/model-config#extended-thinking) + and [errors: effort isn't available with thinking turned + off](https://code.claude.com/docs/en/errors#effort-isnt-available-with-thinking-turned-off). + As of: 2026-10-01. Recheck trigger: either section changes what Claude Code sends when thinking + is off and the prescribed level is `xhigh` or `max`. - **Effort literals do not all reach every surface, and the literal set is not the whole set.** - `max` reaches a session through `CLAUDE_CODE_EFFORT_LEVEL`, `--effort`, `/effort`, or skill and - subagent `effort` frontmatter, the frontmatter case being a surface this skill already - inventories. **The `ultracode` *setting* also trips this** without matching either literal: it is - a Claude Code setting rather than an effort level and "sends `xhigh` to the model", so a surface - pairing it with a thinking-disable surface produces the identical rejection. Only the - effort-setting forms count: instruction text prescribing `/effort ultracode`, `--effort - ultracode`, or `--settings` / Agent SDK `"ultracode": true` or `effortLevel: "ultracode"`. Match - on the effort that reaches the request, not on the spelling. -- **Second arm: the models that reject the disable outright, at every effort level.** Claude - Fable 5, Claude Mythos 5, and Claude Mythos Preview reject `thinking: {type: "disabled"}` - whatever effort is in force, so on that family the disable surface alone is the finding and no - effort operand has to be present for the request to fail. Read the effort operand as a condition - that *narrows* the Opus 5 arm, never as a precondition the whole row inherits. Carried across, it - would pass a surface prescribing thinking-off at `high` on Fable 5 as compliant. **Only the API - form belongs to this arm.** On **Fable 5** the harness thinking-disable surfaces fail differently, - and that failure is I17-a's, not this row's: model configuration states thinking cannot be turned - off there and that the session toggle, `alwaysThinkingEnabled` and `MAX_THINKING_TOKENS=0` "have - no effect there", so they are silent no-ops rather than errors. **For Mythos 5 and Mythos - Preview the harness pages state nothing**, so this row makes no claim about their harness surfaces - in either direction; the API reject is the whole of what is stated for them. + Count `max` wherever text sets it through `CLAUDE_CODE_EFFORT_LEVEL`, `--effort`, `/effort`, or + skill and subagent `effort` frontmatter, the frontmatter case being a surface this skill already + inventories. **Count two `ultracode` forms as `xhigh`:** `--effort ultracode`, and + `effortLevel: "ultracode"` sent through the Agent SDK. Count `/effort ultracode` and the + `ultracode` setting only where the level already in effect is `xhigh` or `max`. Match on the + effort that reaches the request, not on the spelling. **Narrowed on 2026-10-01:** this row + formerly counted all four `ultracode` forms as `xhigh`. It now counts `/effort ultracode` and the + setting only on the condition above, because the row matches the effort that reaches the request + and we treat those two forms as leaving the level in effect unchanged. + Pointer: for which ultracode forms set the level, see + [model configuration: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level). + As of: 2026-10-01. Recheck trigger: that section changes which ultracode forms set the effort + level. +- **Second arm: the models that reject the disable outright, at every effort level.** We treat + `thinking: {type: "disabled"}` as failing whatever effort is in force on this row's set: Claude + Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos + 5 and Claude Mythos Preview. On that set the disable surface alone is the finding and no effort + operand has to be present for the request to fail. Read the effort operand as a condition that + *narrows* the Opus 5 arm, never as a precondition the whole row inherits. Carried across, it + would pass a surface prescribing thinking-off at `high` on Fable 5.1 as compliant. On Sonnet 5.5 + the remediation also points at the guide's thinking-off alternative, which the third arm bounds. + **Only the API form belongs to this arm.** On **the models in I17-a's no-effect + set** the harness thinking-disable surfaces fail differently, and that failure is I17-a's, not + this row's: we treat thinking as unable to be turned off there, and the session toggle, + `alwaysThinkingEnabled` and `MAX_THINKING_TOKENS=0` as silent no-ops rather than errors. **For + the Mythos models the harness pages state nothing**, so this row makes no claim about their + harness surfaces in either direction; the API reject is the whole of what is stated for them. + Pointer: for the per-model accepted values, see [thinking: configuring + thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#configuring-thinking) + and the [troubleshooting + table](https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting#supported-models); + for the harness no-ops, see + [model configuration: extended thinking](https://code.claude.com/docs/en/model-config#extended-thinking). + As of: 2026-10-01. Recheck trigger: a model gains or loses a 400 for `"disabled"` in either + table, or the harness no-op list changes. +- **Third arm, API requests only: `between_tools` at a level or with a change it refuses.** Flag + instruction text that prescribes, or a code sample it tells the reader to send, a request to + Claude Sonnet 5.5 carrying `thinking: {"type": "between_tools"}` together with `xhigh` or `max` + effort, or together with a per-message effort change. No Claude Code surface sets + `between_tools` (see I8-c's declined `sonnet-5-5` widening), so this arm reaches only instruction + surfaces that prescribe Messages API requests, such as a skill or reference doc that tells the + reader what to send. A prompt in application source that sends `between_tools` is not this + catalog's: it belongs to the bundled `claude-api` skill (SKILL.md, "Boundary"), the same split + I8-c records. Pointer: for the levels and changes + `between_tools` refuses, see the [Sonnet 5.5 + guide, running without up-front + thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#running-without-up-front-thinking) + and [thinking: configuring + thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#configuring-thinking). + As of: 2026-10-01. Recheck trigger: either page changes which effort levels `between_tools` + accepts, or per-message effort changes become valid with it. - **Remediate:** on the Opus 5 arm, lower the effort to `high` or below, or leave thinking on, and state which, since the pairing has no third resolution. **On the second arm there is only one - resolution: leave thinking on.** No effort level permits the disable on that family, so a - remediation that offers the reader the choice sends them to a request that still fails. -- **Scope, and where the config check lives:** this row audits **instruction text**. Either arm + resolution: leave thinking on.** No effort level permits the disable on that set, so a + remediation that offers the reader the choice sends them to a request that still fails. The one + addition is Sonnet 5.5 in a prescribed API request: point the reader at the third arm's guide + pointer for the model's thinking-off alternative rather than naming a setting here. **On the + third arm**, drop whichever operand the surface does not need, the top effort level or the + per-turn change; where the surface needs both, replace the request shape with the guide's form at + the third arm's pointer. +- **Scope, and where the config check lives:** this row audits **instruction text**. Any arm expressed as *settings keys* is a config-mechanics finding and belongs to `harness-config:audit`, per this skill's own routing. An instruction-content catalog that also scanned settings files would claim authority a sibling already holds. Instruction text that happens to *live* in a settings file, such as a prompt-type hook's injected text, stays here: the discriminator is whether the content instructs, not which file holds it. -- **Must NOT flag:** `effortLevel: max` as a literal to hunt **in instruction text**. The settings - schema's `enum` accepts `"low"`, `"medium"`, `"high"`, `"xhigh"` only, so a schema-aware editor - flags the value where it is actually written, and an instruction-text auditor sent after the - literal finds nothing and learns nothing. The value is writable, not unreachable: the schema is +- **Must NOT flag:** `effortLevel: max` as a literal to hunt **in instruction text**. We rely on + the settings schema rejecting `max` for that key (pointer: the SchemaStore document + `https://json.schemastore.org/claude-code-settings.json`, its `properties.effortLevel` entry; as + of: 2026-10-01; recheck trigger: the schema accepts `max` there), so a schema-aware editor flags + the value where it is actually written, and an instruction-text auditor sent after the literal + finds nothing and learns nothing. The value is writable, not unreachable: the schema is advisory and the harness reads a file that violates it, which is why the settings-file check is - `harness-config:audit` category H rather than absent. **A document that states either arm to + `harness-config:audit` category H rather than absent. **A document that states any arm to describe or forbid it**, such as this row, a model-adaptation delta chapter, or a verification - record quoting it, on the same audience test I8-b applies: either arm prescribed inside an + record quoting it, on the same audience test I8-b applies: any arm prescribed inside an operative directive is a finding; a document *about* it is not. **The bare `ultracode` prompt keyword.** Instruction text telling a reader to include it in a typed prompt runs one task as a - workflow "without changing the session's effort level", so no effort reaches the request and the + workflow without changing the session's effort level, so no effort reaches the request and the rejected pairing never assembles. **A thinking-disable surface named with no effort level in reach of it, on the Opus 5 arm only**, where the pairing is what fails. On the second arm that is the finding itself, so this fence is scoped to the arm that earns it rather than to the row. -- **Source:** effort, "On Claude Opus 5, thinking cannot be disabled at `xhigh` or `max` effort: - requests that set `thinking: {"type": "disabled"}` at those levels return a 400 error." - Corroborated at thinking-troubleshooting, which supplies the model range and adds that the - restriction "is enforced on each request". The per-surface value sets are read from the surfaces' - own pages: settings for `effortLevel`, environment variables for `CLAUDE_CODE_EFFORT_LEVEL`, - skills and subagents for `effort` frontmatter, and model configuration for `/effort`, the session - and global thinking toggles, and ultracode, the last enumerating the three routes that turn the - *setting* on (`/effort`, `--effort`, `--settings` / Agent SDK). The keyword's separation from the - setting is read from workflows, "Ask for a workflow in your prompt": including `ultracode` in a - prompt runs "a single task as a workflow without changing the session's effort level". The second - arm is thinking's, stated in the paragraph directly after that page's own statement of the Opus 5 - arm: "Claude Fable 5, Claude Mythos 5, and Claude Mythos Preview reject `thinking: {type: - "disabled"}`: thinking cannot be turned off on these models." That sentence carries no effort - qualifier, which is what makes the arm unconditional rather than a wider pairing, and the - adjacency is why the two must be read as separate arms rather than one range. +- **Source:** effort, its Opus 5 section (the `xhigh`/`max` disable rejection). Corroborated at + thinking-troubleshooting, which supplies the model range and the per-request enforcement. The + per-surface value sets are read from the surfaces' own pages: settings for `effortLevel`, + environment variables for `CLAUDE_CODE_EFFORT_LEVEL`, skills and subagents for `effort` + frontmatter, and model configuration for `/effort`, the session and global thinking toggles, and + ultracode, the last enumerating the three routes that turn the *setting* on (`/effort`, + `--effort`, `--settings` / Agent SDK). The keyword's separation from the setting is read from + workflows, "Ask for a workflow in your prompt". The second arm is thinking's per-model table, + where those models refuse `"disabled"` with no effort qualifier while Opus 5 refuses it only at + the top levels: that is what makes the arm unconditional rather than a wider pairing, and why + the two are read as separate arms rather than one range. - **Local coverage of the second arm, measured 2026-08-04: zero operative instances in the repository that authored it.** The disable literal occurs six times across four files: three in this catalog, once in the Opus 5 model-adaptation delta chapter, twice in changelog entries. Every one is a document *about* the restriction, which is the audience-test fence above rather than a passed check. **Re-measure when** a surface here begins prescribing a thinking-disable instead of describing one. -- **Verified 2026-08-04** against those pages, fetched as raw markdown. **Recheck trigger:** the - effort level set gaining or losing a name, the set of `ultracode` forms that reach `xhigh` - changing, the restriction's model range moving, or the set of models that reject the disable - outright changing. +- **As of 2026-08-04** for the Opus 5 pairing as first written (those pages read as raw markdown); + the harness consequence, the `ultracode` forms, the second arm's model set and the third arm + carry their own 2026-10-01 records above. **Recheck trigger:** the effort level set gaining or + losing a name, the set of `ultracode` forms that reach `xhigh` changing, the restriction's model + range moving, or the set of models that reject the disable outright changing. **Row I17-a: `MAX_THINKING_TOKENS=0` presented as a universal off switch** · Tier `mechanical` · Severity `warning`. -- **Detect:** text stating or implying that `MAX_THINKING_TOKENS=0` turns thinking off generally. - It does not. On Fable 5 it has no effect at all, and neither do the session toggle or - `alwaysThinkingEnabled`. On third-party providers it omits the `thinking` parameter instead, - so an adaptive-reasoning model may still think. Also flag text treating - `CLAUDE_CODE_DISABLE_THINKING` as equivalent: that variable omits the parameter on every - provider, which on a model that thinks by default leaves it still thinking. Also flag text - presenting the session thinking toggle or `alwaysThinkingEnabled` as turning thinking off on - Fable 5. Model configuration states they "have no effect there", so the reader is promised a - control that is a silent no-op on that model. +- **Detect:** text stating or implying that `MAX_THINKING_TOKENS=0` turns thinking off generally, + naming neither exception this row keeps: the whole no-effect set (Fable 5.1, Fable 5, Opus 5.5 + and Sonnet 5.5), or third-party providers. Also flag text treating `CLAUDE_CODE_DISABLE_THINKING` + as equivalent to it, and text presenting the session thinking toggle or `alwaysThinkingEnabled` + as turning thinking off on a model in the no-effect set. For what each control does on each + model and provider, follow the pointer. - **Remediate:** carry the exceptions with the claim, or point at the page instead of restating it. - **Adjacent axis:** this is also a harness-capability claim, so **I12 can fire on the same line**. I12 asks whether the claim matches its page; this row asks whether a reader following it gets the behavior they were promised. Report both when both hold. -- **Must NOT flag:** a mention that already carries the Fable 5 or third-party exception. A bare - reference to the variable making no claim about its reach. -- **Source:** environment variables, `MAX_THINKING_TOKENS` "Set to `0` to disable thinking on the - Anthropic API, except on Fable 5, which cannot have thinking turned off; on third-party providers, - `0` omits the `thinking` parameter instead". Model configuration heads the same control "Disable - regardless of effort", so a surface repeating that heading unqualified inherits a claim the - variable's own page contradicts. -- **Verified 2026-08-02** against those two pages, fetched as raw markdown; the session-toggle and - `alwaysThinkingEnabled` arm re-verified 2026-08-04 against model configuration. **Recheck - trigger:** the set of models that cannot disable thinking changing. +- **Must NOT flag:** a mention that already names the whole no-effect set or the third-party + exception; either one is enough. A mention naming only part of the set, such as Fable alone, + names neither and still fires. A bare reference to the variable making no claim about its reach. +- **Source:** environment variables, the `MAX_THINKING_TOKENS` and `CLAUDE_CODE_DISABLE_THINKING` + rows, and model configuration, "Extended thinking" (the no-effect set, the toggles, and the + third-party behavior). Pointer: [environment variables: + variables](https://code.claude.com/docs/en/env-vars#variables), those two rows (read whole per + the fetch route under Sources), and [model configuration: extended + thinking](https://code.claude.com/docs/en/model-config#extended-thinking). +- **As of 2026-10-01** (both pages read as raw markdown that day). **Recheck trigger:** the set + of models that cannot have thinking turned off changing, or the third-party behavior changing. **Row I17-b: mid-session thinking or effort change prescribed without its cost** · Tier `mechanical` · Severity `info`. - **Detect:** an instruction directing a reader to change **effort**, or the **thinking - configuration**, part-way through a session without naming what it costs. Both are rendered into - the request, so either change starts a new cache prefix and the next request re-reads the whole - conversation uncached. The thinking half covers switching among `adaptive`, `enabled` and - `disabled`, and changing `budget_tokens`. -- **Must NOT flag: a Claude Code surface prescribing an *effort* change**, where the harness already - surfaces the cost: it "asks you to confirm before applying the change", and a change resolving to - the level already in effect skips the dialog and keeps the cache. Nor flag a change prescribed - *with* its cost stated, which is the remediation. + configuration**, part-way through a session, with no statement of the prompt-cache cost beside + it. The thinking half covers switching among `adaptive`, `enabled` and `disabled`, and changing + `budget_tokens`. For why such a change costs the cache, follow the Source. +- **Must NOT flag: a Claude Code surface prescribing an *effort* change.** We leave that cost to + Claude Code's own handling of an effort change, so the surface owes no warning of its own. Nor + flag a change prescribed *with* its cost stated, which is the remediation. + Pointer: for how Claude Code handles an effort change, see + [prompt caching: changing effort level](https://code.claude.com/docs/en/prompt-caching#changing-effort-level). + As of: 2026-10-01. Recheck trigger: that section stops covering how Claude Code handles the + cache cost of an effort change, or Claude Code's changelog names a change to that handling. +- **On an API surface, the remediation may offer the per-message effort route, never with + `between_tools`.** Offer it where the model supports it. Never offer it to a surface that sends + `between_tools`: that combination is I17's third arm, not this row. Pointer: for the + per-message route, see [effort: change effort mid-conversation + (beta)](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta); + for the `between_tools` combination, see the [Sonnet 5.5 guide, calibrate + effort](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#calibrate-effort). + As of: 2026-10-01. Recheck trigger: per-message effort leaves beta or changes its cache + behavior, or `between_tools` starts accepting it. - **Reach differs by half, and this is the whole of it.** The effort half reaches every surface, with Claude Code surfaces carved out above. **The thinking half reaches API and Agent SDK surfaces - only.** That is where the page's claim is anchored and where no dialog exists. A Claude Code - surface prescribing a mid-session thinking toggle is **out of reach of this row**, neither excused - by the effort carve-out nor flagged by the thinking half. + only.** A Claude Code surface prescribing a mid-session thinking toggle is **out of reach of + this row**, neither excused by the effort carve-out nor flagged by the thinking half. - **Why the carve-out does not simply extend to thinking, and why the row stops short instead.** - Claude Code's prompt-caching page names exactly two settings that sit outside the prompt text and - are still part of the cache key, model and effort level, and documents the confirmation dialog - for effort alone. So the dialog's protection cannot be assumed for a thinking toggle; but the - harness-side *consequence* of one is equally undocumented, and this catalog does not flag what its - sources do not state. Hence out of reach rather than covered. **Re-scope when** the harness - documents what a mid-session thinking change costs. + The carve-out's pointer covers effort changes only, and on our reading of 2026-10-01 no Claude + Code page covers what a mid-session thinking change costs. So we neither assume Claude Code + handles a thinking toggle the way it handles effort nor flag a cost no source states: out of + reach rather than covered. **Re-scope when** a Claude Code page covers what a mid-session + thinking change costs. - **Why the thinking half is not I17-c, and why both can fire, on accepted changes only.** This row asks what a change *costs*. A switch among the modes, or a change to `budget_tokens`, restarts the cache when the new configuration is accepted and a turn runs under it. I17-c asks @@ -1181,34 +1279,43 @@ Severity `warning`. - **Local coverage of the thinking half, measured 2026-08-04: zero operative instances here.** The session-toggle and `budget_tokens` literals appear only in this catalog, in two model-adaptation delta chapters, and in changelog entries, which are descriptions, not prescriptions. -- **Remediate:** name the re-read cost, and prefer choosing both dials at session start. -- **Source:** prompt caching: "**Effort level**: each effort level has its own cache for the same - model. Changing it mid-session recomputes the entire request, and Claude Code asks you to confirm - before applying the change." The thinking half is thinking's, which puts the thinking - configuration and the resolved effort level in the same position, since both "are rendered into the - prompt itself, so changing any of them starts a new cache prefix". It then enumerates the - changes: "Switching between `adaptive`, `enabled`, and `disabled`, changing `budget_tokens`, - and changing the effort value all invalidate cache breakpoints: message-level breakpoints always - miss, and tool and system-prompt breakpoints can miss too, depending on where the model renders - the configuration." -- **Verified 2026-08-04** against those two pages, fetched as raw markdown. **Recheck trigger:** - effort or the thinking configuration leaving the cache key, the confirmation behavior changing, or - the harness gaining a documented dialog for thinking changes. +- **Remediate:** name the re-read cost, and prefer choosing both dials at session start. On an API + surface, offer the per-message effort change where the model supports it and the request does + not send `between_tools`. +- **Source:** prompt caching, "Changing effort level" (the effort half in Claude Code). The thinking + half, and the API effort half, point at [thinking: thinking and prompt + caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching). +- **As of 2026-08-04** for the thinking half (that page read as raw markdown); the carve-out and the + API per-message route carry their own 2026-10-01 records above. **Recheck trigger:** effort or + the thinking configuration leaving the cache key on the thinking page, or a Claude Code page + starting to cover what a mid-session thinking change costs. **Row I17-c: fixed thinking budget prescribed where adaptive reasoning ignores or rejects it** · Tier `mechanical` · Severity `warning`. Unscoped. Promotion gate MET: the claim is stated on -model-agnostic surface pages and a cross-model migration guide, not in a model guide. **The model +model-agnostic surface pages (thinking, environment variables, model configuration), not in a model guide. **The model ranges below are Detect conditions, not a `Model scope` annotation**, for the reason I17 base states. - **Detect:** instruction text directing a reader to control thinking *depth* with a fixed token - budget on a model that always uses adaptive reasoning. Two arms, with opposite failure modes: + budget on a model that always uses adaptive reasoning: this row's set is Opus 4.7 and later + (Opus 4.7, Opus 4.8, Opus 5, Opus 5.5), Sonnet 5 and later (Sonnet 5, Sonnet 5.5), and the Fable + and Mythos 5-series models (Fable 5.1, Fable 5, Mythos 5.1, Mythos 5). Two arms, with opposite + failure modes: - **Harness arm: silent no-op.** A nonzero `MAX_THINKING_TOKENS`, or - `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1` offered as the way to make one take effect. Nonzero - values are ignored on adaptive-reasoning models, and that variable reaches none of the models - that always use adaptive reasoning, so a reader who follows the instruction sees no error and no - effect, which is the worst of the two failures. + `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1` offered as the way to make one take effect, on a + model in the set. We rank this the worse of the two arms because nothing tells the reader it + failed. - **API arm: hard 400.** `thinking: {type: "enabled", budget_tokens: N}`, or prose presenting a - thinking budget as a tunable number, on Opus 4.7 and later, Sonnet 5, Fable 5, or Mythos 5. + thinking budget as a tunable number, on any model in the set. Claude Mythos Preview is outside + this row's set. + Pointer: for the API arm, see [thinking: configuring + thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#configuring-thinking) + (its `"enabled"` column); for the harness arm, see the `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` + row of [environment variables: + variables](https://code.claude.com/docs/en/env-vars#variables) and [model configuration: + adaptive reasoning and fixed thinking + budgets](https://code.claude.com/docs/en/model-config#adaptive-reasoning-and-fixed-thinking-budgets). + As of: 2026-10-01. Recheck trigger: a model gains or loses a 400 for `"enabled"` in that column, + or the set of always-adaptive models changes. - **Why this is not I17-a.** That row is about `MAX_THINKING_TOKENS=0`, the claim that thinking can be turned *off*, and whether the exceptions travel with it. This row is the claim that thinking depth can be *set to a number*. Different literal, different promise, different failure; both can @@ -1229,26 +1336,22 @@ ranges below are Detect conditions, not a `Model scope` annotation**, for the re config-mechanics or source-code finding on the same discriminator I17 base, I21 and I22 apply, and this catalog audits instruction text. - **Remediate:** point at the effort parameter as the depth control on adaptive-reasoning models, or - carry the model gate with the claim. Upstream's own framing is "It has no direct replacement: - thinking is adaptive, and the `effort` parameter is a separate output-level control, not a thinking - budget". -- **Source:** environment variables: `MAX_THINKING_TOKENS` "Nonzero values are ignored on adaptive - reasoning models unless `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` is set", and that variable, "From - v2.1.111, has no effect on Fable 5, Sonnet 5, or Opus 4.7 and later, which always use adaptive - reasoning". The version qualifier is the second half of the gate fence above. Model - configuration states the same partition from the other side: "Fable 5, Sonnet 5, and Opus 4.7 and - later always use adaptive reasoning. The fixed thinking budget mode and - `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` do not apply to them", while "On Opus 4.6 and Sonnet 4.6, - you can set `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1` to revert", which is the fence above, - stated upstream. The API arm comes from the migration guide: - `thinking: {type: "enabled", budget_tokens: N}` - "is no longer supported on Claude Opus 4.7 or later models and returns a 400 error", with the same - stated for Fable 5 and Mythos 5; corroborated for this model generation by the Sonnet 5 guide, - "Calibrating effort and thinking depth", where manual extended thinking "is not supported on Claude - Sonnet 5 and returns a 400 error. It was deprecated on Claude Sonnet 4.6 and is now removed." -- **Verified 2026-08-04** against those four pages, fetched as raw markdown. **Recheck trigger:** the - set of models that always use adaptive reasoning changing, `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` - regaining or losing reach, or manual extended thinking being reinstated on any model in the range. + carry the model gate with the claim. We do not offer effort as a like-for-like budget + replacement: effort is a separate control, not a thinking budget (effort page). +- **Source:** environment variables, the `MAX_THINKING_TOKENS` and + `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` rows (nonzero values ignored on adaptive-reasoning + models; the latter's loss of reach from v2.1.111 over the always-adaptive models). The version + qualifier is the second half of the gate fence above. Model configuration, "Adaptive reasoning + and fixed thinking budgets", covers the same partition from the other side, including the Opus + 4.6 and Sonnet 4.6 revert that is the fence above. The API arm's model range is read from + thinking's per-model table (the Pointer under Detect); the migration-guide URL that first carried + it is now an index of per-model guides. Corroborated for this model generation by the Sonnet 5 + guide, "Calibrating effort and thinking depth". +- **As of 2026-10-01** for the model set under Detect; **as of 2026-08-04** for the v2.1.111 + version fence, which the environment-variables row no longer states (re-read 2026-10-01), so that + fence rests on its first reading. **Recheck trigger:** the set of models that always use adaptive + reasoning changing, `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` regaining or losing reach, or manual + extended thinking being reinstated on any model in the range. **Row I17-d: tool reliance with thinking disabled and no explicit tool nudge** · Tier `behavioral` · Severity `warning` · Model scope: `sonnet-5`. @@ -1256,32 +1359,40 @@ Severity `warning` · Model scope: `sonnet-5`. - **Detect:** a surface that both (a) prescribes running with thinking off, meaning any thinking-disable surface I17 base enumerates, or a workload the surface states runs thinking-disabled, and (b) depends on the model reaching for tools (search, retrieval, self-verification loops, agentic tool - chains) while stating no explicit instruction about when and how to use those tools. The guide - states the coupling and its remedy in one sentence: "With thinking disabled, the model is less - likely to reach for tools or consider searching; if you rely on tool calls with thinking off, add - an explicit nudge in the system prompt." A brief that turns thinking off and then relies on - default tool reach depends on a disposition that configuration reduced, and the failure is - silent: fewer tool calls, not an error. -- **Remediate:** add the explicit nudge the sentence above prescribes, describing which tools, - when, and why, or leave thinking on. Effort is a second lever: "`high` or `xhigh` effort - settings show substantially more tool usage in agentic search and coding." + chains) while stating no explicit instruction about when and how to use those tools. We treat + thinking-off as reducing the model's reach for tools, so a brief that turns thinking off and then + relies on default tool reach depends on a disposition that configuration reduced, and the + failure is silent: fewer tool calls, not an error. +- **Remediate:** add an explicit tool nudge to the system prompt, describing which tools, when, and + why, or leave thinking on. Effort is a second lever: we treat `high` or `xhigh` as raising tool + use in agentic search and coding. - **Must NOT flag:** a thinking-disable with no tool dependence. A tool-dependent surface that already instructs its tool use explicitly. That is the remediation, present. A surface with no control over and no claim about the thinking configuration, whose tool reliance runs under the default (thinking on). A document *about* the pattern, such as this row, a model-adaptation delta chapter, or a verification record, on the audience test I8-b applies. - **Why scoped:** the coupling claim is stated only in the Sonnet 5 guide. The Opus 4.8 guide's - "Tool use triggering" section states a different default for its model, "a tendency to favor - reasoning over tool calls", with no thinking-off coupling, so it is not a second statement of - this claim; the halves the two guides do share (effort as a tool-usage lever, describe-why-and-how - tool instruction) are general advice, not this row's detect condition. -- **Source:** Sonnet 5 guide, "Tool use triggering": the sentence quoted above, plus the - effort-lever sentence. -- **Verified 2026-08-08** against the Sonnet 5 guide (15,864 bytes, MD5 - `6d23959f0ed226feb06bf20c314029e3`) and, for the scope negative, the Opus 4.8 guide (15,905 - bytes, MD5 `6b9db5b784ad6a7b2e6307c1481b8be9`), both fetched as raw markdown. **Recheck + "Tool use triggering" section covers a different default for its model with no thinking-off + coupling, so it is not a second statement of this claim; the halves the two guides do share + (effort as a tool-usage lever, describe-why-and-how tool instruction) are general advice, not + this row's detect condition. +- **Source:** Sonnet 5 guide, "Tool use triggering" (the coupling, the nudge, and the effort + lever). +- **As of 2026-08-08** (our probe of the Sonnet 5 guide, 15,864 bytes, MD5 + `6d23959f0ed226feb06bf20c314029e3`, and, for the scope negative, the Opus 4.8 guide, 15,905 + bytes, MD5 `6b9db5b784ad6a7b2e6307c1481b8be9`, both read as raw markdown). **Recheck trigger:** a second model guide stating the thinking-off tool-reach coupling, which would meet the promotion gate and unscope this row. +- **Re-justified 2026-10-01 against the current models:** `sonnet-5` stays (see "Tokens of models + that are no longer current"). The row does not widen to `sonnet-5-5`: that model is in I17-a's + no-effect set, so half (a) of Detect cannot hold there, and our probe of its guide that day (read + whole as raw markdown; no artifact stored) found no thinking-off coupling. Tool-discouraging text + on that target is I36's. Pointer: [model configuration: extended + thinking](https://code.claude.com/docs/en/model-config#extended-thinking) for the no-effect set; + [Prompting Claude Sonnet + 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5), + the whole page, for the negative. As of: 2026-10-01. Recheck trigger: Sonnet 5.5 leaves the + no-effect set, or its guide states a thinking-off tool-reach coupling. ### I18: Thinking blocks altered on the way back to the model @@ -1305,8 +1416,8 @@ is what a surface believes about blocks that are *not there*. They share a subje omits `redacted_thinking` blocks, which the protocol requires back unchanged. 3. **Within-turn echo integrity.** An instruction to reorder, edit, truncate, or partially drop the consecutive `thinking` blocks of the latest assistant message, including "keep only the - last one" and "strip thinking before resending" advice. Modified blocks are rejected with a - 400. + last one" and "strip thinking before resending" advice. We treat any altered block as a 400 + rejection. - **Reach: this is wider than Messages API client code.** Any instruction whose output eventually becomes a request body is in scope: Agent SDK callers, harness integrations, and **tooling that parses, excerpts or rewrites a stored transcript that will later be replayed or resumed**. What @@ -1324,19 +1435,14 @@ is what a surface believes about blocks that are *not there*. They share a subje a model-adaptation delta chapter, or a verification record quoting it, on the same audience test I8-b applies: the predicate quoted inside an operative directive is still operative and is a finding; a document *about* the pattern is not. -- **Source:** Thinking, "Preserving thinking blocks": "Pass every `thinking` block back to the API - complete and unmodified, alongside the `tool_use` block it accompanied", and "Within the latest - assistant message, the sequence of consecutive `thinking` blocks must match what the model - generated in the original request: you can't rearrange, edit, or partially drop them." Same page, - "Thinking encryption": "Full thinking content is encrypted and returned in the `signature` field - on each thinking block", and "Redacted thinking blocks": "Filtering on - `block.type == "thinking"` alone silently drops `redacted_thinking` blocks and breaks the - multi-turn protocol…". +- **Source:** Thinking, "Preserving thinking blocks" (echoing blocks back complete and unmodified, + and within-turn sequence integrity); same page, "Thinking encryption" (the `signature` field) + and "Redacted thinking blocks" (what a type-equality filter drops). - **Local coverage, measured 2026-08-02: zero instances of all three shapes in the repository that authored this row**, which ships it consumer-facing and unexercised by its own corpus. Stated so the absence reads as an as-of measurement rather than as a passed check. **Re-measure when** a round-trip or transcript-replay path lands here. -- **Verified 2026-08-02** against the Thinking page, fetched as raw markdown. **Recheck trigger:** a +- **As of 2026-08-02** (the Thinking page read as raw markdown). **Recheck trigger:** a content-block type joining or leaving the set the protocol requires echoed back. **Row I18-a: a leading thinking block treated as required where the model does not require one** · @@ -1345,18 +1451,18 @@ Tier `mechanical` · Severity `warning`. - **Detect:** instruction text asserting, or directing work premised on, a validation rule that assistant turns must begin with a thinking block. Three shapes: 1. **Reinsertion.** An instruction to insert, synthesize, or restore a leading `thinking` block - when assembling history from mixed sources, so that each assistant turn "starts with one". + when assembling history from mixed sources, so that each assistant turn starts with one. 2. **History rewriting on resume.** An instruction to rewrite, normalize, or discard a conversation because it began without thinking or ran under a different thinking configuration. 3. **Presence-assuming logic.** An instruction to read, index, or branch on an assistant turn's - first content block as though it were a `thinking` block. A turn where Claude chose not to - think carries none, and the same conversation can hold turns of both kinds. + first content block as though it were a `thinking` block. We treat an assistant turn produced + without thinking as having no such block, and one conversation as able to mix both kinds. - **Why the belief is a finding and not a harmless one.** The remediation a reader reaches for is fabrication, and a hand-built block carries no valid `signature`, which is the base row's shape 1, and a rejected request. This row is therefore the upstream cause of the base row's violation, not a restatement of it; report both when a surface states the premise *and* acts on it. -- **Remediate:** pass history back in whatever shape you have it, and treat a thinking block as +- **Remediate:** send history back as it is, without reshaping it, and treat a thinking block as optional per assistant turn, in tests too, where a no-thinking turn is the case the assumption hides. - **Reach: the base row's, unchanged, and for all three shapes.** A path back to the model is what @@ -1367,37 +1473,29 @@ Tier `mechanical` · Severity `warning`. logic rather than a rejected request, which is a code-correctness matter this catalog does not audit. **Re-scope when** the stored transcript's content-block shape is documented. - **Must NOT flag: text scoped to a legacy manual thinking budget AND to the final assistant - turn**, where the requirement is real. The page carves it out itself, stating that those models - "enforce that the final assistant turn of a thinking-enabled request begins with one", and the - enforcement is exactly that wide: the final assistant turn of a thinking-enabled request, no other turn. A + turn**, where the requirement is real. The page carves it out itself, and we treat the + enforcement as exactly that wide: only the last assistant turn, and only when the request has + thinking on. A legacy-scoped instruction demanding a leading block on *every* assistant turn over-requires past its own source and still flags. The gate is the model's thinking mode plus the turn it names, not the sentence's confidence, and as in I17-c **the finding is the missing gate, never the mention.** - **The base row's own advice**, which is not this row's inverse: the relaxation "is about - validation, not about what you should send", so an instruction to pass blocks you *have* back - unmodified, particularly during tool use, is correct and stays correct. Reading this row as + **The base row's own advice**, which is not this row's inverse: we read the relaxation as about + validation, not about what to send, so an instruction to return the blocks you *have* untouched, + tool-use turns above all, is correct and stays correct. Reading this row as license to drop blocks inverts both rows at once. A document *about* the assumption, on the audience test I8-b applies. -- **Source:** Steering thinking, "Turn validation": "Assistant turns don't need to start with a - thinking block", with the three consequences stated there, one per shape above: turns where - Claude chose not to think "are valid history as-is"; a conversation begun without thinking, or - under a different thinking configuration, resumes "without rewriting its history"; and history - assembled from mixed sources "doesn't need thinking blocks reinserted at the start of each - assistant turn to pass validation". The legacy carve-out is that same "Turn validation" section's - own parenthetical, quoted in the fence above. The presence half is the same page, "How Claude - decides when to think": "a turn where Claude chose not to think contains no thinking block. Don't - build application logic that assumes every assistant turn starts with one." +- **Source:** Steering thinking, "Turn validation" (the relaxation, with one consequence per shape + above, and the legacy carve-out the fence above relies on); the presence half is the same page, + "How Claude decides when to think". - **Why this page is cited and not the sibling.** The Thinking page carries the same pair, but - compressed into a single sentence inside "Thinking with tool use": in extended (manual) mode the - API "additionally enforces that the final assistant turn of a thinking-enabled request begins with - a thinking block", and "Adaptive mode relaxes this: no assistant turn needs to start with one." - That corroborates this row; it does not carry it. Steering thinking is where the relaxation is - stated operatively, with the three history-shape consequences the detect shapes are drawn from, - plus the presence caution, so it is cited as decisive and the sibling as corroboration. Separate from - both is that page's *strip* claim, that the API "may strip thinking blocks that would create an - invalid turn structure": server-side degradation of a request, not a rule about what history a - caller may send, and it licenses nothing here. + compressed into one passage inside "Thinking with tool use" (manual-mode enforcement and the + adaptive-mode relaxation). That corroborates this row; it does not carry it. Steering thinking + is where the relaxation is stated operatively, with the three history-shape consequences the + detect shapes are drawn from, plus the presence caution, so it is cited as decisive and the + sibling as corroboration. Separate from both is that page's *strip* claim about turn structure: + we read it as server-side degradation of a request, not a rule about what history a caller may + send, and it licenses nothing here. - **Local coverage, measured 2026-08-04: zero operative instances here**, on the same footing as the base row, since nothing in this repository assembles, rewrites, or replays history back to the model. @@ -1405,7 +1503,7 @@ Tier `mechanical` · Severity `warning`. own `type` rather than by position, so it is correct by construction rather than by this rule. Stated as an as-of measurement, not a passed check. **Re-measure when** a history-assembly or replay path lands here. -- **Verified 2026-08-04** against the Steering thinking page, fetched as raw markdown. **Recheck +- **As of 2026-08-04** (the Steering thinking page read as raw markdown). **Recheck trigger:** the turn-validation relaxation narrowing, or the set of models that enforce a leading thinking block changing. @@ -1420,10 +1518,11 @@ by `--opinion`. report against a different harness, and a later release reorders the table, so a figure with no stated re-derivation event silently becomes a claim about the past told in the present tense. - **Remediate:** either point at the vendor's announcement and restate nothing, or keep the figure - and attach the four-part record: the claim, the announcement it came from, the as-of date, and a - trigger naming an observable event (a new frontier-model release, a suite version bump, a decision - that would turn on the figure). Label the figures as launch-day snapshots where that is what they - are; leaving them as history is a valid outcome and usually the right one. + as the surface's own recorded decision and attach the record beside it: a pointer to the exact + source section the figure came from, the as-of date, and a recheck trigger naming an observable + event (a new frontier-model release, a suite version bump, a decision that would turn on the + figure). Label the figures as launch-day snapshots where that is what they are; leaving them as + history is a valid outcome and usually the right one. - **Must NOT flag: a verbatim upstream baseline held for drift detection.** A vendored copy exists to be compared byte-for-byte against its source, so stamping it would corrupt the comparison it exists to serve. This is a genuine suppression, not a routing case, which is what distinguishes @@ -1433,9 +1532,10 @@ by `--opinion`. - **Must NOT flag:** a benchmark named as a pointer with no figure attached. A figure already carrying a trigger, whatever heading that trigger sits under. - **Source:** none. No official page states that a restated benchmark figure needs a re-derivation - event, which is why this check is `OPINION`-tier and off by default. The four-part shape it asks - for is this monorepo's `docs/conventions/upstream-drift/README.md`; in a standalone install the - four parts, not the path, are the requirement. + event, which is why this check is `OPINION`-tier and off by default. The record shape it asks for + (the decision, a pointer, an as-of date, a recheck trigger) is this monorepo's + `docs/conventions/upstream-drift/README.md`; in a standalone install those parts, not the path, + are the requirement. ### I20: Prefilled assistant response @@ -1460,13 +1560,10 @@ condition, not a `Model scope` annotation**, for the reason I17 states. row, a migration guide, or a model-delta chapter, on the same audience test I8-b applies: a prefill prescribed inside an operative directive is a finding; a document *about* prefill is not. Instructions targeting an explicitly pinned earlier model, which still supports it. -- **Source:** prompting best practices, "Migrating away from prefilled responses": "Starting with - Claude 4.6 models and Claude Mythos Preview, prefilled responses (providing a partial assistant - message for Claude to continue from) on the last assistant turn are no longer supported. Requests - with prefilled assistant messages to these models return a 400 error… Earlier models continue to - support prefills, and adding assistant messages elsewhere in the conversation is not affected." -- **Verified 2026-08-02** against that page, fetched as raw markdown; the standalone prefill - technique page now redirects to the prompt-engineering overview. **Recheck trigger:** any change to +- **Source:** prompting best practices, "Migrating away from prefilled responses" (the unsupported + model range, the 400, and what stays unaffected). +- **As of 2026-08-02** (that page read as raw markdown; the standalone prefill technique page + redirected to the prompt-engineering overview). **Recheck trigger:** any change to that page, or the unsupported-model range moving. ### I21: Effort level pinned across a model change with no re-sweep @@ -1480,30 +1577,29 @@ not a `Model scope` annotation**, for the reason I17 states. - **Detect:** a surface prescribing a **durable** effort level that states no re-derivation when the pinned model changes: a fleet-wide or project-wide pin, a "set effort to X and leave it" - instruction, a level tied to a named model lane. The effort scale is calibrated per model, so the same - level name does not carry the same underlying value across models; a level measured against one - model and carried to the next is a pin nobody re-measured. -- **The consequence varies by model, which is why the range sits in Detect.** The first-run hold - sentence for Fable 5, Opus 4.8, and Opus 4.7 is no longer on the model-config page. Opus 5.5 - starts at `medium` unless an explicit choice sets a level, and a top-level `effortLevel` in the - user settings file does not count for Opus 5.5. That key still applies on Opus 5, Fable 5.1, and - earlier models. Opus 5.5 and models released after it start at their own default until `/effort` - or the `/model` picker saves a level for them. A top-level `effortLevel` in project, local, or - managed settings, or one passed with `--settings`, applies to every model. Launching with - `--effort` is an explicit choice and applies to that launch. **Claim, basis, as of, recheck:** - that paragraph, - [model-config: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level), - 2026-09-28, and a re-fetch of that section that no longer matches it. The row fires on the - missing re-derivation regardless of model; the hold is severity context, never a fence. + instruction, a level tied to a named model lane. We treat the effort scale as calibrated per + model, so a level name measured against one model is not the same setting on the next; a level + carried across a model change is a pin nobody re-measured. +- **The consequence varies by model, which is why the range sits in Detect.** We treat Opus 5.5 as + starting at `medium` unless an explicit choice sets a level, with a top-level `effortLevel` in + the user settings file not counting for it (Opus 5, Fable 5.1 and older models still honor that + key), and models released after Opus 5.5 as starting at their own default until `/effort` or the + `/model` picker saves a level. Every model takes the level a top-level `effortLevel` sets in the + project, local or managed file or through `--settings`, and `--effort` sets it for one launch. + The model-config page no longer carries a first-run effort hold for Fable 5, Opus 4.8 or Opus + 4.7. The row fires on the missing re-derivation regardless of model; the hold is + severity context, never a fence. + Pointer: [Adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level). + As of: 2026-09-28. Recheck trigger: that section changes which sources set a model's level, or + the user-settings exemption for Opus 5.5. - **Remediate:** attach the re-derivation to the pin, naming the model the level was measured - against and stating that a model change re-opens it, or run the sweep. Upstream's own wording for the - action: "If you carried effort settings over from an earlier model, run a fresh effort sweep on - your evals rather than reusing them." -- **Must NOT flag: a prescription of `high` where `high` is the resolved target's default.** Setting - the default "produces exactly the same behavior as omitting the `effort` parameter entirely", so - on a model that defaults to `high` such a pin carries no measured calibration that could go stale. - **The exemption keys to the resolved target, never to the wording.** In Claude Code `high` is the - default on every model that supports effort **except Opus 5.5 and Sonnet 5.5, which default to + against and stating that a model change re-opens it, or run a fresh effort sweep on the + surface's evals rather than reusing carried-over levels. +- **Must NOT flag: a prescription of `high` where `high` is the resolved target's default.** We + treat setting the default as equivalent to omitting the `effort` parameter, so on a model that + defaults to `high` such a pin carries no measured calibration that could go stale. + **The exemption keys to the resolved target, never to the wording.** In Claude Code every + effort-capable model defaults to `high` **except Opus 5.5 and Sonnet 5.5, which default to `medium`, and Opus 4.7, which defaults to `xhigh`**, so when the run's resolved target is one of those the exemption lifts and a `high` pin is a finding, **including a broad model-agnostic "always use `high`" that names no model at all**. That broad pin is the sharper case rather than @@ -1534,13 +1630,12 @@ not a `Model scope` annotation**, for the reason I17 states. chapter, or a verification record, on the audience test I8-b applies. Nor a level **reported as a named third party's practice** rather than prescribed to the reader: a practitioner's stated setup is `OPINION`-tier testimony, not a pin the surface owns. -- **Source:** model configuration: "The effort scale is calibrated per model, so the same level name - does not represent the same underlying value across models", stated with no model qualifier, and - the whole basis for the check. The same page supplies the resolution order, with the default carve-out: - "`high` on every model that supports effort, except that Opus 5.5 and Sonnet 5.5 default to - `medium`, Opus 4.7 defaults to `xhigh`". Effort supplies the remediation's wording and the - equivalence of the default to omitting the parameter. -- **Verified 2026-09-28** against both pages, fetched as raw markdown (model configuration 109,848 +- **Source:** model configuration, "Choose an effort level" (the per-model calibration, stated with + no model qualifier, and the whole basis for the check) and "Adjust effort level" (the resolution + order and the per-model defaults this row keeps as its own settings above). Effort, "How effort + works" and its per-model sections, supplies the remediation and the equivalence of the default + to omitting the parameter. +- **As of 2026-09-28** (our probe: both pages read as raw markdown, model configuration 109,848 bytes; effort 39,458 bytes). **Recheck trigger:** the calibration property being restated as cross-model-stable, a first-run effort hold returning to the model-config page, the resolution order or the `effortLevel` user-settings exemption for Opus 5.5 changing, or `high` ceasing to be @@ -1558,8 +1653,9 @@ by `--opinion`. matrices are revised on every release, so lanes derived from one reading and written down without their provenance become a claim about a model generation that has since passed, told in the present tense. -- **Remediate:** name the baseline and the triggers. State which vet or reading the lanes came from - and when, then list the events that re-open it: the pinned model changes, per-model guidance or +- **Remediate:** name the baseline and the triggers, as the record parts beside the doctrine: a + pointer to the vet or reading the lanes came from, its as-of date, and the observable events + that re-open it: the pinned model changes, per-model guidance or its notes change, the selection-matrix rows change, a volatile figure a lane turns on drifts. **The action on a trigger is a targeted delta check against the named baseline, never a re-derivation from scratch.** That is what makes the trigger cheap enough to honor, and a trigger @@ -1582,16 +1678,17 @@ by `--opinion`. instead of restating lanes has nothing to go stale. Nor doctrine already carrying a baseline and triggers, whatever heading they sit under. - **Why this is not I19, and not the catalog trigger.** I19 covers a restated *benchmark figure* and - asks for the four-part record; it says nothing about lane assignments and nothing about how to + asks for the record parts; it says nothing about lane assignments and nothing about how to *act* when a trigger fires. The delta-not-re-run discipline is this row's own contribution. The catalog-wide recheck trigger does not reach it either: that trigger governs **this catalog's** staleness against its Sources, not an audited surface's staleness against the pages its doctrine was read from. - **Source:** none. No official page states that model-routing doctrine must name a baseline and delta triggers, which is why this check is `OPINION`-tier and off by default, the same footing as - I19, and it adds no Sources entry for the same reason. The four-part shape it asks for is this - monorepo's `docs/conventions/upstream-drift/README.md`; in a standalone install the four parts, not - the path, are the requirement. + I19, and it adds no Sources entry for the same reason. The record shape it asks for (a pointer to + the baseline, an as-of date, observable recheck triggers beside the doctrine) is this monorepo's + `docs/conventions/upstream-drift/README.md`; in a standalone install those parts, not the path, + are the requirement. ### I23: Context-budget directive to stop, summarize, or hand off @@ -1602,7 +1699,7 @@ Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface readable, which tempts a `mechanical` tag, but I8-b's Detect is a literal three-phrase match and is even seeded in the pre-scan, and it is `behavioral`. The `mechanical` rows rest on a documented hard consequence: I10 on a refusal category the API returns, I21 on a property its page states outright. -This row rests on a reported model *tendency*, "can occasionally suggest a new session", with no +This row rests on a reported model *tendency* (an occasional new-session suggestion), with no documented hard consequence, which is the behavioral tier's definition. The stake is the Output format rule: behavioral findings ship as proposals verified per Deletion tiers, never as confident removals. @@ -1611,8 +1708,8 @@ confident removals. summarize, hand off, trim its work, or start a new session **on that basis**, and instruction text or injected hook output that surfaces a remaining-context count to the model where the surface could avoid it. The guide names the count as the usual trigger for the behavior, so the disclosure - and the directive are one subject; it also hedges the disclosure arm to "where possible", and this - row tracks that hedge rather than reading it as an absolute. + and the directive are one subject; it also hedges the disclosure arm, and this row tracks that + hedge rather than reading it as an absolute. - **The discriminator is who decides, on what evidence.** A directive tells the model to judge its own window and act; a mechanism resolves the window from an instrumented signal and acts itself. Only the first is this row's subject. @@ -1659,8 +1756,9 @@ confident removals. reasoning with it. - **Residency is a severity input, not an admission test.** A trigger in a `description` is resident whenever the skill listing admits it, which is the default, since `disable-model-invocation: true` - also suppresses the description from context ([Skills](https://code.claude.com/docs/en/skills), - verified 2026-08-08). A body-borne trigger costs context only once the skill loads, or at startup + also suppresses the description from context (pointer: + [Skills](https://code.claude.com/docs/en/skills#frontmatter-reference), as of 2026-08-08). A + body-borne trigger costs context only once the skill loads, or at startup in a subagent with the skill preloaded. Both are findings; the resident one is the more expensive to leave. **Second-source recheck trigger:** that page's invocation-control table changing which fields keep a description in context, which would re-rank the two residencies and is the only fact @@ -1679,24 +1777,37 @@ confident removals. over-production is the same contract I8's families carry. It is deliberately **not** anchored to the bare term "context window", which is ordinary vocabulary in any surface discussing sessions and would return the corpus instead of a candidate set. -- **Source:** Fable 5 guide, "Rare cases of context-budget concern": "In very long sessions, Claude - Fable 5 can occasionally suggest a new session, offer to summarize and hand off, or trim its own - work. This is most often triggered when the harness shows a remaining-token countdown to the model. - Avoid surfacing explicit context-budget counts where possible." -- **Verified 2026-08-08** against that guide, fetched as raw markdown (177 lines). **Verified +- **Source:** Fable 5 guide, "Rare cases of context-budget concern" (the behaviors, the + remaining-token countdown as the usual trigger, and the hedged advice against surfacing counts). +- **As of 2026-08-08** (our probe: that guide read as raw markdown, 177 lines). **Verified negative, which is what holds the scope annotation on:** the Opus 5 guide (11,225 bytes) and the - Sonnet 5 guide (15,864 bytes) were fetched as raw markdown the same day and searched for this - claim. Neither states it. Opus 5's only mention of the context window is a capability statement, - that its instruction following, tool calling, and reasoning "stay consistent throughout the - window", which is the opposite subject: a reason the concern does not arise, not a counter-steer - against it. **Recheck trigger:** a second model guide stating the claim, which would meet the + Sonnet 5 guide (15,864 bytes) were read as raw markdown the same day and searched for this + claim. Neither states it. Opus 5's only mention of the context window is a capability statement + about consistency across the window, which is the opposite subject: a reason the concern does + not arise, not a counter-steer against it. **Recheck trigger:** a second model guide stating + the claim, which would meet the promotion gate and unscope this row, or that section ceasing to name the remaining-token countdown as the trigger, which is what joins the disclosure arm to the directive arm. -- **Widened to `fable-5-1` on 2026-09-03:** the bundled `claude-api` skill's model-migration - reference (Claude Code 2.1.258), sections Migrating to Claude Fable 5.1 and Migrating to Claude - Fable 5.1 from Claude Fable 5, restates this behavior for Claude Fable 5.1 and states that Fable 5 - prompt guidance carries over. **Recheck trigger:** publication of a Fable 5.1 prompting guide, - whose statement of this claim replaces this basis and joins `## Sources`. +- **Widened to `fable-5-1` on 2026-09-03**, first on the bundled `claude-api` skill's + model-migration reference. That basis's trigger fired when the Fable 5.1 guide was published. The + widening now rests on our reading of that guide on 2026-10-01 (read whole as raw markdown; no + artifact stored): no section names a context-budget difference from Fable 5, so we keep the + Fable 5 claim for `fable-5-1`. Pointer: [Prompting Claude Fable + 5.1](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1), + the text above its first heading, which carries no heading id on the rendered page, so the link + is to the page. As of: 2026-10-01. Recheck trigger: the Fable 5.1 guide naming a context-budget + difference from Fable 5, or its opening changing what it says about Fable 5 prompts. +- **Re-justified 2026-10-01 against the current models:** `fable-5` stays, since Fable 5 is a + current model (see "Tokens of models that are no longer current"). The row does not widen to + `sonnet-5-5`: a countdown after tool + results on that target is I37's subject, and our probe of the Opus 5.5 and Sonnet 5.5 guides that + day (each read whole as raw markdown; no artifact stored) found no statement of this row's claim. + Pointer: [Prompting Claude Opus + 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5) + and [Prompting Claude Sonnet + 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5), + whole pages, since the probe is a negative over every section. As of: 2026-10-01. Recheck + trigger: either guide states this row's claim, which meets the promotion gate. ### I24: Instruction relying on silent generalization @@ -1705,10 +1816,9 @@ Promotion gate MET: two model guides state the identical claim (see Source). - **Detect:** instruction text that demonstrates or names ONE instance while the author's evident intent is a whole class, with no explicit scope statement: text a literal-minded executor would - satisfy by doing exactly the one instance and stopping. Current models "interpret prompts - literally and explicitly, particularly at lower effort levels": they do "not silently generalize - an instruction from one item to another", and do "not infer requests you didn't make". Four - shapes: + satisfy by doing exactly the one instance and stopping. We treat current models as reading + instructions literally, more so at lower effort: an instruction about one item is not carried to + its siblings, and unstated requests are not inferred. Four shapes: 1. **A worked example standing in for a rule**, "rename this field like so" meaning every such field, with no "apply to every / all / each" scope line. 2. **An enumeration whose tail the executor must guess**: a list ended with "etc." or "and @@ -1717,9 +1827,9 @@ Promotion gate MET: two model guides state the identical claim (see Source). where the surrounding procedure plainly processes many. 4. **A per-item step whose iteration is implied but never stated**: "check the frontmatter" in a skill that processes N files. -- **Remediate:** state the scope explicitly. The guides' own worked remediation: "Apply this - formatting to every section, not just the first one." Name the class an "etc." tail was standing - in for; attach the iteration to the per-item step. +- **Remediate:** state the scope explicitly, as one line naming every item the instruction covers + rather than only the first. Name the class an "etc." tail was standing in for; attach the + iteration to the per-item step. - **Must NOT flag:** an instruction whose single-instance reading is correct, because the request really is one item. Scope stated anywhere in reach of the instruction (a "for each X below" frame, a table iterated by contract, a stated general rule the example sits inside as a labeled example). @@ -1731,9 +1841,9 @@ Promotion gate MET: two model guides state the identical claim (see Source). proposes scope statements, so the Stopping condition's high-consequence withholding does not bind it: adding explicitness to a safety gate is safe where trimming one is not. - **Source:** Sonnet 5 guide, "More literal instruction following", and Opus 4.8 guide, "More - literal instruction following". The two sections state the Detect sentences verbatim-identically - for their respective models, and both give the same remediation example quoted above. -- **Verified 2026-08-08** against both guides, fetched as raw markdown (Sonnet 5: 15,864 bytes, MD5 + literal instruction following". The two sections state the same claim for their respective + models, and both give the same remediation example. +- **As of 2026-08-08** (our probe of both guides as raw markdown; Sonnet 5: 15,864 bytes, MD5 `6d23959f0ed226feb06bf20c314029e3`; Opus 4.8: 15,905 bytes, MD5 `6b9db5b784ad6a7b2e6307c1481b8be9`). **Recheck trigger:** either guide ceasing to state the literalism claim, or a model guide stating that its model resumes generalizing instructions, @@ -1743,22 +1853,26 @@ Promotion gate MET: two model guides state the identical claim (see Source). Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `error` · Surfaces: all. The severity is `error` because following the instruction produces a rejected request, the same consequence class -as I17, I18 and I20. Unscoped. Promotion gate MET: the claim is stated in the cross-model migration -guide, not only in model guides. **The model range is a Detect condition, not a `Model scope` annotation**, -for the reason I17 base states. +as I17, I18 and I20. Unscoped. Promotion gate MET: the claim is stated on a model-agnostic feature +page (thinking, "Sampling parameters", the pointer under Detect), not only in model guides. **The +model range is a Detect condition, not a `Model scope` annotation**, for the reason I17 base +states. -- **Detect:** instruction text directing a reader to set `temperature`, `top_p`, or `top_k` to a - non-default value, commonly "raise the temperature" for variety, creativity, or design +- **Detect:** instruction text directing a reader to move `temperature`, `top_p`, or `top_k` off + its default, commonly "raise the temperature" for variety, creativity, or design divergence, or "set `temperature = 0`" for determinism, where the run's resolved target model is - Claude Opus 4.7 or later, Claude Sonnet 5, Claude Fable 5, or Claude Mythos 5 (the same range - I17-c's API arm names). On those models a non-default sampling parameter returns a 400 error; the - SDK request types still define the fields for compatibility, so the instruction type-checks and - fails only at the API. -- **Remediate:** remove the parameter and steer the behavior in prompt text. Upstream's framing: - "Remove these parameters when migrating, and use system-prompt instructions to guide tone and - variety instead." For design variety specifically, the propose-options pattern is the documented - replacement (see I26). Where the prescription was `temperature = 0` for determinism, carry - upstream's note that "it never guaranteed identical outputs" on prior models either. + in this row's set: Claude Opus 4.7 or later (Opus 4.7, Opus 4.8, Opus 5, Opus 5.5), Claude + Sonnet 5 or later (Sonnet 5, Sonnet 5.5), Claude Fable 5.1, Claude Fable 5, Claude Mythos 5.1, + Claude Mythos 5, or Claude Mythos Preview. Fire on those models whether or not the surface also + sets thinking, and even where the instruction would type-check against an SDK. A model outside + the set is outside this row, whatever it does with thinking on. + Pointer: for the set, see [thinking: sampling + parameters](https://platform.claude.com/docs/en/build-with-claude/thinking#sampling-parameters). + As of: 2026-10-01. Recheck trigger: a model joins or leaves that section's list. +- **Remediate:** remove the parameter and steer tone and variety with system-prompt instructions + instead. For design variety specifically, the propose-options pattern is the documented + replacement (see I26). Where the prescription was `temperature = 0` for determinism, note that + it did not guarantee identical outputs on prior models either. - **Must NOT flag: a claim carrying its own model gate.** Text scoped to a pinned earlier model where the parameters are live is correct rather than stale. As in I17-c, **the finding is the missing gate, never the mention.** **The parameter expressed as an SDK request field, config @@ -1767,21 +1881,18 @@ for the reason I17 base states. of the word**, such as body temperature, disk or thermal temperature, or color temperature, which share the token and nothing else. A document *about* the pattern, on the audience test I8-b applies. -- **Source:** migration guide, "Migrating to Claude Sonnet 5", where sampling parameters "set to a - non-default value are not accepted and return a 400 error"; same guide for the Opus range: - "Setting `temperature`, `top_p`, or `top_k` to any non-default value on Claude Opus 4.7 or later - models, including Claude Opus 5, returns a 400 error", with the SDK-compatibility and - determinism notes quoted from its Opus 5 section; same guide for the Fable/Mythos arm, "Migrating - to Claude Mythos 5 and Claude Fable 5 from Claude Opus 5": "The prefill and sampling-parameter - restrictions, and the thinking display behavior, carry over from Claude Opus 5 unchanged." - Corroborated at What's new in Claude Sonnet 5 ("This is new for Sonnet-class models; the same - constraint was previously introduced on Claude Opus 4.7") and in the Sonnet 5 guide, "Tone and - writing style", which supplies the Remediate quote. The migration guide and What's new in Claude - Sonnet 5 are Sources entries. -- **Verified 2026-08-08** against those pages, fetched as raw markdown (migration guide 148,590 - bytes, MD5 `bfe459a13cd59d6ac93a6826910d5a28`; whats-new-sonnet-5 11,490 bytes, MD5 - `19acce78670ceb337b99ce8fbac03fc5`). **Recheck trigger:** the rejecting model range moving, or - sampling parameters being reinstated on any model in it. +- **Source:** thinking, "Sampling parameters" (the set and the gate, under Detect). The Fable and + Mythos guide's section on migrating from Claude Opus 5 (under Sources) keeps the Fable/Mythos + carry-over. Corroborated at the Claude Sonnet 5 model page and in the Sonnet 5 guide, "Tone and + writing style", which supplies the Remediate line. Both are Sources entries. + Pointer: [Claude Sonnet 5, good to + know](https://platform.claude.com/docs/en/models/sonnet-5/overview#good-to-know) and + [Sonnet 5 guide, tone and writing + style](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5#tone-and-writing-style). + As of: 2026-10-01 for the set and the corroborating pages; the determinism note under Remediate + rests on our 2026-08-08 reading of a page no current page replaces. Recheck trigger: the + rejecting model set moving, sampling parameters being reinstated on any model in it, or a docs + page starting to cover the determinism note. ### I26: Generic negative steering on open-ended design briefs @@ -1789,21 +1900,20 @@ Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `info` · Surfaces: Promotion gate MET: two model guides converge (see Source). - **Detect:** operative instruction text steering visual design away from a model's default style - with generic negatives or vague qualifiers such as "don't use that color", "make it clean and - minimal", "less corporate", or "avoid a generic AI look", with neither a concrete specification nor a propose-options step. - Both guides - state the failure the same way: such instructions "tend to shift the model to a different fixed - palette rather than producing variety." Also flag text recommending sampling parameters as the - design-variety mechanism, which additionally reaches I25 on an in-range target. -- **Remediate:** either of the two approaches both guides state work reliably: (1) specify a - concrete alternative, since the model "follows explicit specs precisely"; or (2) have the model - propose distinct visual directions first (each as background / accent / typeface plus a one-line - rationale), have the user pick one, and implement only that, on Sonnet 5 "the recommended way to - produce meaningfully different design directions across runs", since `temperature` is not - accepted there. A short anti-generic-aesthetics directive with concrete, enumerable negatives + with generic negatives or vague qualifiers (banning a color without naming its replacement, + asking for "clean" or "minimal", "less corporate", "not generic") with neither a concrete + specification nor a propose-options step. We treat such instructions as trading one default look + for another default look, not as producing variety. Also flag text recommending sampling + parameters as the design-variety mechanism, which additionally reaches I25 on an in-range target. +- **Remediate:** either of the two approaches both guides cover: (1) specify a concrete + alternative, which the model follows precisely; or (2) have the model propose distinct visual + directions first (each naming its colors and type with a short reason), have the + user pick one, and implement only that, which is the recommended route to distinct directions + across runs on Sonnet 5, where `temperature` is not accepted. A short anti-generic-aesthetics + directive with concrete, enumerable negatives (named fonts, named schemes) is the guides' own sanctioned snippet shape, not a finding. Pair the - exclusion list with an iteration step: check which styles the result used instead, and extend - the list when those are unwanted too. + exclusion list with an iteration step: look at what the output fell back to, and add those + styles to the list when they are unwanted as well. - **Must NOT flag:** concrete enumerable negatives. Naming the exact fonts, palettes, or patterns to avoid is the sanctioned shape, distinct from a vague qualifier. Non-design uses of "clean" / "minimal" (a clean audit, a minimal reproduction). A surface that already runs the propose-options @@ -1813,11 +1923,9 @@ Promotion gate MET: two model guides converge (see Source). frontend defaults", convergent on the default-style behavior, the fixed-palette failure of generic instructions, and both remediations; the Sonnet 5 guide adds the temperature-is-gone ground for preferring propose-options. Third convergent guide: Opus 5.5, "Frontend design - defaults", where "avoid a generic AI look" "mostly swaps one default for another", named patterns - work, and the iteration step is "check which styles the first result used instead, and extend - the list if needed." -- **Verified 2026-08-08** against both guides, fetched as raw markdown (hashes as in I24); the - Opus 5.5 guide verified 2026-09-23 (hash as in I8-c). + defaults" (the generic-look negative, named patterns, and the iteration step). +- **As of 2026-08-08** (our probe of both guides as raw markdown, hashes as in I24); the Opus 5.5 + guide as of 2026-09-23 (hash as in I8-c). **Recheck trigger:** either guide's design section dropping the fixed-palette claim or the propose-options recommendation. @@ -1835,8 +1943,8 @@ gate is unmet). adjudicates that the line actually premises brevity on effort rather than merely co-locating the two. - **Remediate:** replace the effort clause with an explicit length or style instruction, the - documented control for response length ("To control response length, prompt for it explicitly"), - keeping any effort change only where its stated ground is thinking volume, cost, or latency. + documented control for response length, keeping any effort change only where its stated ground + is thinking volume, cost, or latency. - **Must NOT flag: effort lowered on thinking-volume, cost, or latency grounds**, since "reduce effort to cut thinking cost on mechanical work" states the property the docs confirm; this row fires only on the length premise. @@ -1847,26 +1955,35 @@ gate is unmet). - **Must NOT flag: `effortLevel` settings keys and `effort:` frontmatter as such**, on the same discriminator as I21 and I17: a config value implements a choice without stating the premise; this row audits instruction text, including instruction text that lives in a config file. -- **Source:** Opus 5 prompting guide: "The effort parameter controls how much the model thinks - rather than how much it says: lowering effort can reduce thinking volume without reliably - shortening the visible response. To control response length, prompt for it explicitly." - Corroborated by Effort, whose Opus 5 section states it unhedged: "Effort controls thinking - volume, not visible response length: on Claude Opus 5, changing effort does not reliably shorten - responses, so prompt for length instead." The guide's "can" hedge is quoted as written. The - detection needs only the negative half (not reliably shortening), which both pages state. -- **Verified 2026-08-08** against the live guide raw-`.md` (11,225 bytes, MD5 +- **Source:** Opus 5 prompting guide, "Response length and verbosity" (effort governs thinking + volume, not response length, with a hedge). Corroborated by Effort, "Recommended effort levels + for Claude + Opus 5", which states it unhedged. The detection needs only the negative half (effort does not + reliably shorten the response), which both pages state. +- **As of 2026-08-08** (our probe: the live guide raw-`.md`, 11,225 bytes, MD5 `8579d63fc9f793784b8c56320fd74e71`, byte-identical to the 2026-07-25 corpus capture) and the effort page's Opus 5 section, both fetched that day. **Recheck trigger:** either page restating the property model-agnostically or a second model guide stating it (gate met → unscope), or either statement disappearing from its page. +- **Re-justified 2026-10-01 against the current models:** `opus-5` stays (see "Tokens of models + that are no longer current"). We do not widen it. A model joins this row's scope only when its + guide states that effort does not reliably shorten the response. On 2026-10-01 neither the Opus + 5.5 nor the Sonnet 5.5 effort section stated that. We read both as tying higher effort to longer + output, which runs against this row's premise rather than with it, so widening to either model + would flag a line its guide supports. Pointer: [Opus + 5.5 guide, calibrate + effort](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#calibrate-effort) + and [Sonnet 5.5 guide, calibrate + effort](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#calibrate-effort). + As of: 2026-10-01. Recheck trigger: either section starts stating that effort does not reliably + shorten the response, or stops tying higher effort to longer output. ### I28: Over-aggressive trigger emphasis and blanket tool defaults Tier `mechanical` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surfaces: all. Unscoped. The claim sits on the model-agnostic best-practices page, its Migration considerations restate it -generation-wide ("Claude 4.6 models are more proactive and may overtrigger on instructions that -were needed for previous models"), and no later model guide reverses it; the Sonnet 5 and Opus 4.8 -literalism sections ("interprets prompts literally and explicitly") corroborate the mechanism. +generation-wide (overtriggering on instructions written for earlier models), and no later model +guide reverses it; the Sonnet 5 and Opus 4.8 literalism sections corroborate the mechanism. - **Detect:** two arms of one defect, prompting written against undertriggering that no longer exists: @@ -1883,19 +2000,30 @@ literalism sections ("interprets prompts literally and explicitly") corroborate fact, however emphatically set. **Must NOT flag: a document *about* the pattern**, on the same audience test I8-b applies. This row is the canonical instance. - **Remediate:** for arm 1, normal conditional phrasing: "Use this tool when …". For arm 2, replace - the blanket default with the condition it was standing in for: "Use [tool] when it would enhance - your understanding of the problem." Verify per Deletion tiers (a consequential removal needs a + the blanket default with the condition it was standing in for: "Use [tool] when ." Verify per Deletion tiers (a consequential removal needs a closed watch, an editorial one does not); watch for overtriggering receding, not just continued triggering. -- **Source:** prompting best practices, "Tool usage": prompts "designed to reduce undertriggering - on tools or skills … may now overtrigger. The fix is to dial back any aggressive language. Where - you might have said 'CRITICAL: You MUST use this tool when…', you can use more normal prompting - like 'Use this tool when…'"; "Overthinking and excessive thoroughness": "Replace blanket - defaults with more targeted instructions … Instructions like 'If in doubt, use [tool]' will - cause overtriggering"; "Migration considerations": "Tune anti-laziness prompting". -- **Verified 2026-08-08** against that page, fetched as raw markdown. **Recheck trigger:** those +- **Source:** prompting best practices, "Tool usage" (dialing back aggressive trigger language, with + a worked example), "Overthinking and excessive thoroughness" (targeted instructions over blanket + defaults), and "Migration considerations" (anti-laziness prompting). +- **As of 2026-08-08** (that page read as raw markdown). **Recheck trigger:** those three sections changing, or any model guide stating that a current model undertriggers and needs emphasis restored, which would re-open the scoping question. +- **That trigger fired on 2026-10-01, and the arms stand.** Two current guide sections on tool + and search triggering (pointers below) were re-read that day; neither asks for emphasis or a + blanket default back. So the row stays unscoped and keeps both arms, with one fence added: + **Must NOT flag a search instruction scoped to a named class of facts**, even when it overrides + the model's own judgment that no search is needed. That is a targeted condition, the arm-2 + remediation's own shape, not a blanket default; a line discouraging tool use is I36's. + Pointer: [Sonnet 5.5 guide, tool use in + chat and knowledge + work](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#tool-use-in-chat-and-knowledge-work) + and [Fable 5.1 guide, search triggering at low + effort](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#search-triggering-at-low-effort). + As of: 2026-10-01. **Recheck trigger, restated:** a model guide asking for forced-compliance + emphasis or a blanket tool default to be restored on a current model, which re-opens the scoping + question; a guide describing under-triggering and prescribing a targeted condition does not. - **Routes to the findings relay.** I28 and I29 (scanner-fed) and I30 to I33 (lane-fed, admitted through `--from-lane`) are the checks in this catalog whose findings reach `review:fanout`'s apply relay, behind `--persist-findings`. I28's two arms carry one @@ -1907,8 +2035,8 @@ literalism sections ("interprets prompts literally and explicitly") corroborate including the body-scope fence, are [context/persist-findings.md](../context/persist-findings.md). - **The remediation is a downgrade, never a deletion.** The directive survives verbatim and only its volume changes. A proposal that removes the instruction rather than its shouting has misread - the check. The Source's own worked example replaces `"CRITICAL: You MUST use this tool when…"` - with `"Use this tool when…"`, keeping the instruction and dropping the shout. + the check. The Source's own worked example drops a leading `CRITICAL: You MUST` and keeps the + instruction, dropping only the shout. **One byte may legitimately differ: sentence-initial capitalization.** Where the emphasis is a *leading* wrapper, dropping it promotes the next word to sentence-initial position, so `…MUST resolve the item id` becomes `Resolve the item id`. That is forced by the edit, not a @@ -1975,17 +2103,18 @@ I19 from benchmark figures to every dated claim about a harness, a tool, an upst measurement. - **Detect:** a claim carrying an as-of date ("verified 2026-07-13", "as of 2.1.240") but no - observable event that would re-open it. The four-part record is claim, basis, as-of date, and - recheck trigger; a stamp with three of the four is the finding. + observable event that would re-open it. The record is the decision, a pointer to the exact + source section, an as-of date, and a recheck trigger beside the decision; a dated stamp that + lacks the trigger is the finding, whether or not its pointer is present. - **Must NOT flag:** a dated stamp whose trigger lives in a named owner record the site points at ("recheck per `reference/parent-contract.md`"); a CHANGELOG entry or ADR, which are history by design; a date that is data (a release date in a table) rather than a verification stamp. - **Must NOT flag: a moving ref or an undated version literal** (`main`, a branch name, "requires 2.1.200"). With no as-of date there is no stamp for this row to judge. When the evidence in hand shows the literal is stale, report it under [Out-of-catalog defects](#out-of-catalog-defects). -- **Remediate:** add the trigger as an observable event (a release note naming the flag, a fetch - no longer carrying the quoted span, a version floor moving), or point the site at the dated - owner record. +- **Remediate:** add the trigger as an observable event (a release note naming the flag, the + pointed-at section changing the value the decision rests on, a version floor moving), or point + the site at the dated owner record. --- @@ -2003,7 +2132,7 @@ phrasing costs the same wherever it sits. never saw. - **Must NOT flag:** a structural contrast between two current alternatives ("separate body Bash calls rather than pre-compute lines" names two present mechanisms); a CHANGELOG or ADR; a dated - four-part record whose trigger legitimately names the prior state. + record (pointer, as-of date, recheck trigger) whose trigger legitimately names the prior state. - **Remediate:** state the current rule and its reason in the present tense; move the history to the CHANGELOG or an ADR. @@ -2033,11 +2162,12 @@ surface (CLAUDE.md, a natively read AGENTS.md, rules, a user or project skill or `~/.claude/skills/synced/`), or gated by settings, environment, plan, or host, since none of those rosters can be enumerated from files; a capability named by class ("a visualization capability") that deliberately avoids a binding; a reference inside a fenced example. -- **Stamp:** the reserved `anthropic-skills` namespace and the `~/.claude/skills/synced/` download - directory are from ("Claude Code reserves the name - `anthropic-skills` ... for skills synced from claude.ai"; "Claude Code downloads your account's - skills into `~/.claude/skills/synced/`"). **Verified 2026-09-27.** **Recheck trigger:** either - span leaving that page, or a release note moving synced skills to another namespace or directory. +- **Stamp:** we treat `anthropic-skills` as the namespace reserved for skills synced from + claude.ai, and `~/.claude/skills/synced/` as their download directory. Pointer: + and + . **As of 2026-09-27.** + **Recheck trigger:** that page stops reserving the namespace or naming the directory, or a + release note moves synced skills to another namespace or directory. - **Remediate:** name the skill that exists, or describe the capability by class per the seam-phrasing convention; never leave a route to nowhere. @@ -2095,12 +2225,67 @@ Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surface - **Must NOT flag:** the instruction on a long-chat or project surface for short back-and-forth follow-ups, which is the shape the guide recommends; a document *about* the pattern, on the audience test I8-b applies. -- **Source:** Opus 5.5 guide, "Thinking instructions in chat system prompts": "Leave it out where - you want the model to keep re-examining its earlier work, for example in long analyses, or in - agentic tasks where a later step can reveal a mistake in an earlier one", and the instruction "may - also make the model less likely to point out a mistake in an earlier answer on its own". - **Verified 2026-09-23** against the guide's raw `.md` (hash as in I8-c). **Recheck trigger:** that - section dropping the carve-out, or a second model guide stating it. +- **Source:** Opus 5.5 guide, "Thinking instructions in chat system prompts" (where to leave the + instruction out, and its cost to self-correction). + **As of 2026-09-23** (our probe: the guide's raw `.md`, hash as in I8-c). **Recheck trigger:** + that section dropping the carve-out, or a second model guide stating it. + +### I36: Tool-discouraging language + +Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surfaces: all · Model scope: +`sonnet-5-5` (a single guide states the claim; promotion gate unmet). + +- **Detect:** a line that restricts tool or search use as a standing policy, naming no particular + tool and giving no reason tied to one, or that ranks the model's recall above checking a source. + The line must sit on a component that has, or hands work to an agent that has, a search or + retrieval tool. +- **Remediate:** where the component's work depends on facts that change, propose replacing the + line with the guide's form at the pointer; otherwise propose deleting it. That form is targeted, + not a blanket default, so the replacement does not reach I28. +- **Must NOT flag:** a limit on one named tool that carries its own reason (side effects, cost, a + rate limit, a slow or destructive tool), which is the surface's to set. I11's steering from an MCP + tool to an equivalent CLI, which redirects tool use rather than discouraging it. A surface with + no tool that could check a fact. A document *about* the pattern, on the audience test I8-b + applies. +- **Adjacent rows:** I28 flags the opposite calibration, emphasis written to force a trigger; one + line is never both. I17-d covers reduced tool reach with thinking off, on `sonnet-5`. +- **Source:** Sonnet 5.5 guide, "Tool use in chat and knowledge work". Pointer: [Sonnet 5.5 guide, + tool use in chat and knowledge + work](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#tool-use-in-chat-and-knowledge-work). + As of: 2026-10-01. Recheck trigger: a second model guide stating the claim, which meets the + promotion gate and unscopes the row, or the section dropping it. + +### I37: Harness text after every tool result + +Tier `behavioral` · Authority `ANTHROPIC-DOCS` · Severity `warning` · Surfaces: all, chiefly hook +instruction text and text that configures a harness · Model scope: `sonnet-5-5` (a single guide +states the claim; promotion gate unmet). + +- **Detect:** in an interactive session (one where the user can type while a turn runs), any + configuration or instruction that adds model-visible text of any kind after every tool result + without a condition. In Claude Code the + mechanical form is a `PostToolUse` hook whose matcher covers every tool and which returns + `additionalContext` on every call; elsewhere it is an instruction to a harness to append text on + every step. For why the frequency matters, follow the pointer. +- **Remediate:** gate the text on the condition that needs the model's attention (a narrower + matcher, a failure, a threshold) and delete any figure nothing acts on. For a harness built on + the API, place the text as the pointer's section directs rather than as this row restates it. +- **Must NOT flag:** the API's own task-budget feature, which this row does not cover. A hook that + fires on a narrow condition or only occasionally. A run where nobody can type mid-turn, such as a + non-interactive `-p` run or an unattended lane. Hook output that reaches only the user, such as a + status line or a transcript notice. A document *about* the pattern, on the audience test I8-b + applies. +- **Adjacent rows:** I23 (scoped to Fable) covers a model deciding to stop or hand off on a budget + it was shown; this row covers where and how often harness text lands. Their scopes do not + overlap, so one countdown never draws both on one target. +- **Source:** Sonnet 5.5 guide, "Mid-turn user messages". Which hook output reaches the model is + read from hooks, "PostToolUse decision control". Pointer: [Sonnet 5.5 guide, mid-turn user + messages](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#mid-turn-user-messages-and-task-budgets) + and [hooks: PostToolUse decision + control](https://code.claude.com/docs/en/hooks#posttooluse-decision-control). + As of: 2026-10-01. Recheck trigger: a second model guide covering the same topic, which meets + the promotion gate and unscopes the row, or that section changing which kinds of per-step text + it covers. --- @@ -2121,8 +2306,8 @@ more aggressive. is not the posture in the places where being wrong is expensive. - **Report every withholding** in the run's own section, naming the check it moderated and the ground. A silently suppressed finding reads as coverage. -- **Source:** none. The "except in highly important areas" carve-out appears on no official page, - and it is the calibration knob the de-prescription guidance (I8's Fable 5 source) leaves unset. +- **Source:** none. A high-consequence carve-out appears on no official page, and it is the + calibration knob the de-prescription guidance (I8's Fable 5 source) leaves unset. --- diff --git a/plugins/harness-config/skills/audit-instructions/reference/finding-identity.md b/plugins/harness-config/skills/audit-instructions/reference/finding-identity.md index c2c8ec68bd..7ed6f7f465 100644 --- a/plugins/harness-config/skills/audit-instructions/reference/finding-identity.md +++ b/plugins/harness-config/skills/audit-instructions/reference/finding-identity.md @@ -103,6 +103,8 @@ prints `#REFUSED`, the row, and the reason, and is never given an id. | I33 | `I33.spoke-self-description` | 1 | | I34 | `I34.maintainer-rationale-in-yaml` | 1 | | I35 | `I35.settled-answers-instruction` | 1 | +| I36 | `I36.tool-discouraging-language` | 1 | +| I37 | `I37.harness-text-after-every-tool-result` | 1 | A new catalog check lands here and in `scripts/finding-ids.sh` in the same change; `scripts/finding-ids.test.sh` fails when the two tables or the catalog's check headings disagree. diff --git a/plugins/harness-config/skills/audit-instructions/scripts/finding-ids.sh b/plugins/harness-config/skills/audit-instructions/scripts/finding-ids.sh index 96e033a155..8586f6c0f3 100755 --- a/plugins/harness-config/skills/audit-instructions/scripts/finding-ids.sh +++ b/plugins/harness-config/skills/audit-instructions/scripts/finding-ids.sh @@ -119,6 +119,8 @@ claim_template() { I33) echo "I33.spoke-self-description" ;; I34) echo "I34.maintainer-rationale-in-yaml" ;; I35) echo "I35.settled-answers-instruction" ;; + I36) echo "I36.tool-discouraging-language" ;; + I37) echo "I37.harness-text-after-every-tool-result" ;; *) return 1 ;; esac } @@ -183,16 +185,14 @@ fi # Home-directory scope. The user surface is $HOME-wide, not # ${CLAUDE_CONFIG_DIR:-~/.claude} alone, because Claude Code reads instruction # files that sit outside the config directory and under $HOME. -# Claim: "Claude Code loads `CLAUDE.md` and `CLAUDE.local.md` from your -# current working directory and every directory above it", and reads -# "every `AGENTS.md` and `.claude/AGENTS.md` in your working directory -# and the directories above it" where no CLAUDE.md counts; imports -# accept "Both relative and absolute paths", and a user-scope file's -# imports load without the approval dialog. -# Basis: https://code.claude.com/docs/en/memory ("How CLAUDE.md files load", -# "When Claude Code reads AGENTS.md", "Import additional files"). +# We treat CLAUDE.md, CLAUDE.local.md and the AGENTS.md files as loading +# from the working directory and each ancestor, and an import as able to +# name any absolute path, so the surface set reaches past the config dir. +# Pointer: https://code.claude.com/docs/en/memory ("How CLAUDE.md files +# load", "When Claude Code reads AGENTS.md", "Import additional +# files"). # As of: 2026-09-29. -# Recheck trigger: the ancestor-loading or import-path sentences change, or a +# Recheck trigger: the ancestor-loading or import-path sections change, or a # release note adds an instruction file type outside *.md, a skill's # own files and the settings and hooks JSON above. instruction_shape() { diff --git a/plugins/harness-config/skills/audit-permission-state/SKILL.md b/plugins/harness-config/skills/audit-permission-state/SKILL.md index 330d2f0aac..4cab1cfcb8 100644 --- a/plugins/harness-config/skills/audit-permission-state/SKILL.md +++ b/plugins/harness-config/skills/audit-permission-state/SKILL.md @@ -19,11 +19,13 @@ from one it could not read, there is no `claude permissions` subcommand or machi the merged allow/ask/deny set, and none of it exists outside a live session. This skill computes that locally, off a live session, as line records a script can read. -> **Verification.** Claim: no CLI surface exports the merged allow/ask/deny set. Basis: -> [CLI reference](https://code.claude.com/docs/en/cli-reference) lists no `permissions` subcommand and -> no standalone `config` subcommand; the two JSON surfaces it documents, `claude auto-mode defaults` -> and `claude auto-mode config`, print classifier rules, not permission rules. As of 2026-09-12. -> Recheck when a release note mentions a permissions export or a `/permissions` export action. +We found no CLI surface that exports the merged allow/ask/deny set, so this skill computes it. + +- **Pointer**: for the CLI's commands, see + [CLI commands](https://code.claude.com/docs/en/cli-reference#cli-commands). +- **As of**: 2026-09-12 +- **Recheck trigger**: a release note mentions a permissions export or a `/permissions` export + action. It answers a question the siblings do not. `audit-permission-grants` asks whether the grants you **wrote** are durable and portable; `audit` asks whether your config files are **correct**. This @@ -46,12 +48,10 @@ could actually open, and what each one holds. Two native surfaces act on the same permission rules this skill reports, so "fix my permissions" can mean any of the three. -- **`fewer-permission-prompts` (bundled skill)**: scans transcripts for common read-only Bash and - MCP calls and adds a prioritized allowlist to project `.claude/settings.json`. The model and the - person can both invoke it. -- **`/permissions` (built-in command, alias `/allowed-tools`).** An interactive dialog to view, - add, and remove rules by scope, review recent auto mode denials, and edit classifier rules on its - Auto mode tab. It is reserved for the person to run; the model does not invoke it. +- **`fewer-permission-prompts` (bundled skill)**: writes an allowlist into project settings to cut + prompts. Either the model or the person may run it. +- **`/permissions` (built-in command).** The person's interactive editor for rules by scope and + for auto mode's classifier rules. This skill never runs it. - **This skill (marketplace plugin).** Computes the effective merged set across every scope, off a live session, with source, precedence, auto mode drops, and dead config. It writes nothing. @@ -75,9 +75,9 @@ when they resolve, never that they are present. The four-part records live in holds including under `--oracle`. It is not the same as writing nothing at all: `--oracle` spawns a real `claude -p` session, and a session rewrites `~/.claude.json` and adds project, session-env, security, subagent and backup state under your config directory. The flag prints that before it -spawns anything. Every other action writes nothing anywhere. Managed policy is read-only by -construction: those are admin-write OS locations or a claude.ai Owner role, so a plugin could not -author them even if it wanted to. +spawns anything. Every other action writes nothing anywhere. This skill treats managed policy as +read-only by construction: its sources are admin-write OS locations or a claude.ai Owner role, so a +plugin could not author them even if it wanted to. ## Arguments @@ -144,32 +144,31 @@ inert scopes= outranked_by= one per beaten en Two mechanics decide those records, and conflating them produces confident wrong answers: -- **Rules merge across scopes rather than override**, so the same rule in the same list at two scopes - has no winner. Both are live, and `scopes=` names every contributor. Never report one of them as - having overridden the other. -- **Kind is decided by evaluation order, deny then ask then allow, from any scope, in both +- **The merge treats rules as merging across scopes, never overriding**, so the same rule in the + same list at two scopes has no winner. Both are live, and `scopes=` names every contributor. + Never report one of them as having overridden the other. +- **The merge decides kind by evaluation order, deny then ask then allow, from any scope, in both directions.** A user-level deny blocks a project-level allow just as a project-level deny blocks a user-level allow. Scope rank does not enter into it. This is what answers "why is my allow rule ignored": the `inert` record names the rule that beat it. -- **A rule that is a bare tool name reaches every call of that tool.** A whole-tool deny removes the - tool from context entirely, so every other rule naming it is inert, other denies included; - `EndConversation` is the documented exception. A whole-tool ask prompts for every call, so no scoped - allow for that tool applies. Both print a `NOTE:` naming the tool. +- **The merge treats a bare tool name as reaching every call of that tool.** A whole-tool deny + makes every other rule naming that tool inert, other denies included, with the one exception + `reference/criteria.md` names. A whole-tool ask leaves no scoped allow for that tool in effect. + Both print a `NOTE:` naming the tool. -`reference/criteria.md` maps every `precedence_basis` token to the sentence it follows from, and -states the two standing bounds the run prints. +`reference/criteria.md` maps every `precedence_basis` token to the docs section it follows from, +with that record's as-of date and recheck trigger, and states the two standing bounds the run +prints. ## Phase 3: What entering auto mode drops -On entering auto mode, broad allow rules that grant arbitrary code execution are **silently dropped**, -and restored when the session leaves auto mode again. This stage says which of yours survive. +This stage predicts which of your broad allow rules auto mode sets aside on entry, with no notice, +until the session leaves auto mode again, and which carry over. **It describes a transition most run shapes never make, so state the precondition when you report -it.** Auto mode is the built-in starting mode in one of the seven documented run shapes: a Pro, Max, -or Team plan in a terminal or the VS Code extension. Every other shape, `claude -p` and the Agent SDK -among them, starts in Manual and never makes this transition. The run prints the full list as a -`DIFF-NOTE`; carry it rather than presenting the diff as unconditional. -`reference/criteria.md` §"The auto-mode entry diff" holds the dated record. +it.** The run prints, as a `DIFF-NOTE`, the run shapes we treat as starting in auto mode; carry it +rather than presenting the diff as unconditional. `reference/criteria.md` §"The auto-mode entry +diff" and §"The precondition: which sessions enter auto mode at all" hold the pointers. Stage: `automode-entry-diff.sh`, fed the merge. @@ -181,20 +180,20 @@ entry-diff kept scopes= carries over entry-diff summary allow_before= dropped= suspended= kept= ``` -- **Only allow rules change on entry.** Deny and ask are evaluated before the classifier in every - mode, so they are not part of this diff. Do not report them as "surviving". +- **The diff covers allow rules only.** We treat deny and ask as evaluated before the classifier + in every mode, so they are not part of this diff. Do not report them as "surviving". - **Neither label is permanent.** `dropped` and `suspended` both describe what is in force while auto mode is active; the rules are restored when the session leaves it, and nothing edits a settings file. The two labels are kept apart because the remedies differ: a `dropped` rule is fixable by narrowing that rule, while `suspended` is a global switch no rule edit reaches. -- **`class` names the documented reason**: `blanket`, `interpreter-wildcard`, `package-manager-run`, +- **`class` names the reason**: `blanket`, `interpreter-wildcard`, `package-manager-run`, `agent`, or `monitor`. The three shell shapes come from `lib/permission-patterns.sh`, the vocabulary `audit-permission-grants` check P1 also scans with; `agent` and `monitor` are - whole-tool classes this script tests on the tool token. `Monitor` allow rules joined the dropped - set upstream in v2.1.236, because Claude Code runs Monitor commands through the shell. -- **`autoMode.classifyAllShell` inverts the answer wholesale.** When true it suspends *every* Bash and - PowerShell allow rule, so narrow rules do **not** carry over. It is resolved only from the scopes - the classifier reads, so a project- or local-scope copy is reported inert rather than obeyed. + whole-tool classes this script tests on the tool token, `monitor` from v2.1.236. +- **`autoMode.classifyAllShell` inverts the answer wholesale.** The stage treats it, when true, as + suspending *every* Bash and PowerShell allow rule, so narrow rules do **not** carry over. It + resolves the key only from the scopes it treats the classifier as reading, so a project- or + local-scope copy is reported inert rather than obeyed. `reference/criteria.md` holds the pointers. - **`--oracle` is opt-in and priced.** It spawns a real `claude -p` session to corroborate the prediction. Measured cost: your settings files are untouched, but `~/.claude.json` is rewritten and project, session-env, security and subagent state appear under your config directory. A capture @@ -213,12 +212,12 @@ lint summary findings= checks_run= status= ``` Eleven checks: three `C2-*` dead-config gates, `C5-disableType`, and seven `C6-*` rules-that-cannot-match, including a malformed Tool(content) rule and an uncompilable Read/Edit path. -`reference/criteria.md` maps each to the sentence it follows from and lists the legitimate rule shapes -the checks are written NOT to flag. +`reference/criteria.md` maps each to the docs section it follows from and lists the legitimate rule +shapes the checks are written NOT to flag. -- **`C5-disableType` is the one to read first.** `disableAutoMode` must be the **string** `"disable"`; - a boolean is valid JSON, is accepted, and does nothing, so the operator believes auto mode is - locked out when it is not. +- **`C5-disableType` is the one to read first.** We treat only the **string** `"disable"` as + locking auto mode out; a boolean is valid JSON and the check flags it, because the operator + believes auto mode is locked out when it is not. - **The three `C2` gates stay separate findings.** Different scope sets, different version histories: an operator who fixed one and saw the count drop would reasonably believe they had fixed all three. - **`findings=0` is a clean bill only under `status=read`.** Under `status=incomplete` a scope could @@ -235,8 +234,9 @@ not permission rules the harness matches. Independent of the pipeline, it reads bash "${CLAUDE_PLUGIN_ROOT}/skills/audit-permission-state/scripts/automode-block-lint.sh" [--critique] ``` -- **`C4-defaults`**: a customized section that omits `"$defaults"`. Customizing **replaces** the - built-in list rather than adding to it, so the finding names how many built-in entries are gone. +- **`C4-defaults`**: a customized section that omits `"$defaults"`. We treat such a section as + **replacing** the built-in list rather than adding to it, so the finding names how many built-in + entries are gone. `reference/criteria.md` §"The `autoMode` block lane" holds the pointer. - **`C2b-contradiction`**: the same subject in `allow` and in a deny section. - **`C3-shadowed`**: an entry an earlier `hard_deny` already forecloses, so it can never fire. - **`--critique` surfaces `claude auto-mode critique`, wrapped and never replaced.** It owns the @@ -260,37 +260,35 @@ is clean" and "the block was never read" is the whole point. An administrator deploys managed policy believing it is policy. Some of it is; some is not, and nothing surfaces which. Stage: `managed-conformance.sh`, fed the inventory. -- **`managed enforced deny `**: the strongest thing an administrator can write. No level, - command line included, can override a managed permission rule, and a tool denied at any level - cannot be allowed at another. -- **`managed loosenable rule …`**: the interaction that surprises people. "Managed is highest" and - "deny before ask before allow, **from any scope**" are both true: a lower-scope deny beats a managed - allow without ever overriding it. -- **`managed loosenable autoMode`**: a managed `autoMode` section is **additive, not a policy - boundary**. A developer cannot remove entries it provides, but a developer-added `allow` can - override an organization `soft_deny`. Permissions, hooks, MCP, sandbox-filesystem and - sandbox-network each got an exclusivity lock; auto mode did not. +- **`managed enforced deny `**: the strongest thing an administrator can write. We treat a + managed permission rule as outranked by no level, command line included, and a tool denied at + any level as allowed at none. +- **`managed loosenable rule …`**: the interaction that surprises people. Managed settings rank + highest, and evaluation still runs deny, then ask, then allow **from any scope**, so a + lower-scope deny beats a managed allow without ever overriding it. +- **`managed loosenable autoMode`**: we report a managed `autoMode` section as **additive, not a + policy boundary**, since a developer's own entries can loosen it. - **`managed loosenable lockout`**: `disableAutoMode` set to anything but the string `"disable"`. +`reference/criteria.md` §"Managed policy, and what it does not buy" holds the pointers for all +four. + **This report never prescribes.** It says what the consumer's own policy does and does not achieve; every rule string it prints came from a file it read. It ships no security floor of its own. -**Completeness is bounded on every run.** Server-managed settings are fetched at sign-in and cached -at `~/.claude/remote-settings.json`. The cache is user-writable and can be stale, so "managed" means -the local admin surfaces only; the cache is not folded in and is not the live policy. The live -delivery has no local path. A surface that could not be read gets its own note saying so, because an -administrator reading silence as "no policy deployed" is the failure this report exists to prevent. -The note routes that diagnosis to `/status` (Setting sources, and the Organization policy line for a -policy that did not load, a policy-helper failure, or a credential that is signed in but not the one -in use) and to `claude doctor`, which shows the same Organization policy line. - -**Record.** Claim: `/status` and `claude doctor` carry an Organization policy line that says why -the organization's policy could not be loaded, and `/status` marks the credential that is not in -use. Basis: and the -`/status` row of , plus the managed-settings page's -statement that `claude doctor`'s Organization policy line says where the policy loaded from or why -it did not (Claude Code v2.1.261 or later). As of: 2026-09-28. Recheck: those pages drop the -Organization policy line or stop naming `/status` as the place a managed source is shown. +**Completeness is bounded on every run.** "Managed" means the local admin surfaces only. We do not +fold the server-managed cache at `~/.claude/remote-settings.json` into the effective set: it is +user-writable and can be stale, so it is not the live policy, and the live delivery has no local +path. A surface that could not be read gets its own note saying so, because an administrator +reading silence as "no policy deployed" is the failure this report exists to prevent. The note +routes that diagnosis to the Organization policy line in `/status` and `claude doctor`. + +- **Pointer**: for where a managed source and a policy that failed to load are shown, see + [Read the source in /status](https://code.claude.com/docs/en/managed-settings#read-the-source-in-/status) + and the `/status` row of [All commands](https://code.claude.com/docs/en/commands#all-commands). +- **As of**: 2026-09-28 +- **Recheck trigger**: those pages drop the Organization policy line or stop naming `/status` as the + place a managed source is shown. ## Reading the output honestly @@ -306,41 +304,47 @@ collapse it in the report: - **Every scope and every managed surface emits a record on every OS**, including the ones that do not apply here (`not-applicable`). A surface missing from the output is a defect in this reader, not evidence about the machine. -- **`managed` means the LOCAL managed surfaces.** Server-managed settings are cached at - `~/.claude/remote-settings.json`. The cache is not the live policy, and the script does not fold - it into the effective set. The failure read is the Organization policy line in `/status`. The - script says so on every run; carry it into the report rather than implying completeness. -- **An `ask` finding names where the contract lives.** The quote is on the auto mode config page: - content-scoped ask rules always force a prompt, even in auto mode, and the classifier cannot - auto-approve a match. v2.1.257 fixed the compound-command and subshell miss only. #42797 is - closed. #83766 is still open. Say so when reporting an `ask` result, and point at - `permissions.deny` where the outcome must hold regardless. See `reference/criteria.md`. +- **`managed` means the LOCAL managed surfaces.** The script does not fold the server-managed cache + at `~/.claude/remote-settings.json` into the effective set, and treats the Organization policy + line in `/status` as the failure read. The script says so on every run; carry it into the report + rather than implying completeness. +- **An `ask` finding names where the contract lives**: the auto mode config page, and the open + upstream issue #83766 against it. Carry the caveat `reference/criteria.md` §"Ask rules under auto + mode" words, with its pointer and dated record, when reporting an `ask` result, and point at + `permissions.deny` where the outcome must hold regardless. - **`invalid-json` is not `absent`.** A malformed settings file contributes no rules to the - inventory. A managed settings file, drop-in, MDM plist, or HKLM value that cannot be parsed - refuses startup (exit 1) and names the source, from v2.1.259. A user, project, or local file - shows a Settings Error; after continue, `/status` names the file, and an unparsable user - `settings.json` pauses the retention sweep and warns in `/status`. Report the parse failure, not - an empty scope, and do not describe a managed parse failure as silent non-enforcement. + inventory. Report the parse failure, not an empty scope, and do not describe a managed parse + failure as silent non-enforcement: we treat an unparsable managed source as stopping Claude Code + at startup (from v2.1.259) and an unparsable user, project, or local file as a reported settings + error. + - **Pointer**: for managed sources that fail to parse, see + [Find entries Claude Code dropped](https://code.claude.com/docs/en/managed-settings#find-entries-claude-code-dropped); + for other settings files, see + [Fix a broken settings file](https://code.claude.com/docs/en/settings#fix-a-broken-settings-file). + - **As of**: 2026-10-01 + - **Recheck trigger**: either section changes what Claude Code does when a settings source fails + to parse. ## Scopes -Five, and the two easy to get wrong: `local` resolves **through worktrees to the main checkout**, so -a reader anchored on the worktree root looks where the file is not; `startdir-local` is a -pre-v2.1.211 copy that is **not** a fallback, since permission rules from both files stay in effect. -`managed` is four admin surfaces per OS, not one file, plus a `remote-cache` record for -`~/.claude/remote-settings.json` that is not folded into the effective set. `reference/criteria.md` §Scopes has the full table -and the dated record for the `pre-v2.1.211` boundary. - -**Four documented conditions keep the local file beside `.claude/settings.json` instead**, and the -reader resolves all four: outside a git repository, repository root is the home directory, on Windows, -and repository root or its `.git`/`.claude` not owned by the current user. The basis line names which -applied. A fifth case is stated rather than detected, because it is a helper's behavior and not a -session property: the Agent SDK's `resolveSettings()` always reads from the starting directory. - -**A cloud session reads a different scope set, and the run says so.** The operator's own user and -local settings are not read there, and only server-managed settings arrive, so a user-scope record in -a cloud session describes the container. `CLAUDE_CODE_REMOTE` is the documented detection and the only -entrypoint variable this reader branches on. `reference/criteria.md` §Scopes holds both dated records. +Five, and the two easy to get wrong: the reader resolves `local` **through worktrees to the main +checkout**, so a reader anchored on the worktree root looks where the file is not; it treats +`startdir-local` as a pre-v2.1.211 copy that is **not** a fallback, keeping permission rules from +both files in effect. `managed` is four admin surfaces per OS, not one file, plus a `remote-cache` +record for `~/.claude/remote-settings.json` that is not folded into the effective set. +`reference/criteria.md` §Scopes has the full table and the dated record for the `pre-v2.1.211` +boundary. + +**The reader resolves every condition we treat as keeping the local file beside +`.claude/settings.json` instead**, and the basis line names which applied. One further case, the +Agent SDK's `resolveSettings()` helper, is stated rather than detected, because it is a helper's +behavior and not a session property. `reference/criteria.md` §"The four start-directory +conditions, and the one that is not detectable" holds the list and its pointer. + +**A cloud session reads a different scope set, and the run says so.** We treat a user-scope record +in a cloud session as describing the container, not the operator. The reader detects a cloud +session by `CLAUDE_CODE_REMOTE`, the only entrypoint variable it branches on. +`reference/criteria.md` §Scopes holds both dated records. ## Prerequisites diff --git a/plugins/harness-config/skills/audit-permission-state/reference/criteria.md b/plugins/harness-config/skills/audit-permission-state/reference/criteria.md index 5570d8581d..adcf7a8753 100644 --- a/plugins/harness-config/skills/audit-permission-state/reference/criteria.md +++ b/plugins/harness-config/skills/audit-permission-state/reference/criteria.md @@ -17,13 +17,18 @@ Version: 1.1.0 Last updated: 2026-08-28 -This file defines what `permission-merge.sh` may claim and on which documented mechanic each claim -rests. It exists because an *effective* permission set is a precedence claim, and a precedence claim -with no cited mechanic is folklore. The reader's own record contract lives in `SKILL.md`; the -per-check grant vocabulary lives in the sibling `audit-permission-grants`. Neither is restated here. - -Sources, both fetched 2026-08-11: §How scopes interact and - §Manage permissions and §Settings precedence. +This file defines what `permission-merge.sh` may claim and which documented mechanic each claim +points at. It exists because an *effective* permission set is a precedence claim, and a precedence +claim with no cited mechanic is folklore. Each rule below is our decision in our words with a +pointer to the section that documents the mechanic; read the wording there. The reader's own record +contract lives in `SKILL.md`; the per-check grant vocabulary lives in the sibling +`audit-permission-grants`. Neither is restated here. + +Pointers, both read 2026-08-11: +[settings precedence](https://code.claude.com/docs/en/settings#settings-precedence), and +[Manage permissions](https://code.claude.com/docs/en/permissions#manage-permissions) and +[Settings precedence](https://code.claude.com/docs/en/permissions#settings-precedence) on the +permissions page. --- @@ -34,83 +39,94 @@ Sources, both fetched 2026-08-11: §H | `managed` | Highest precedence. Four surfaces per OS, not one file. See below | | `user` | `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/settings.json`. Where Claude Code's own "Always allow" path writes, so it accumulates the most rules | | `project` | `.claude/settings.json` at the repository root | -| `local` | `.claude/settings.local.json`, resolved **through worktrees to the main checkout**, so anchoring on the worktree root looks where the file is not. Four documented conditions keep it beside `settings.json` instead, and the reader resolves all four: outside a git repository, when the repository root is the home directory, on Windows, and when the repository root or its `.git` or `.claude` entry is not owned by the current user | +| `local` | `.claude/settings.local.json`, resolved **through worktrees to the main checkout**, so anchoring on the worktree root looks where the file is not. Four documented conditions keep it beside `settings.json` instead, and the reader resolves all four: no enclosing git repository, a repository root equal to `$HOME`, a Windows host, and a repository root, `.git` or `.claude` entry owned by another user | | `startdir-local` | A pre-v2.1.211 copy left in the session's start directory. Not a fallback: when both exist the repository root wins on a shared key, **but permission rules from both stay in effect**, so both are live | This table is the dated owner record for the `pre-v2.1.211` boundary. Every other site in this -plugin that names the boundary points here rather than restating it. The settings page states it -directly: "Before v2.1.211, Claude Code kept the file in the starting directory. It still reads a -file an earlier version left there alongside the root file; where both set the same key, the root's -value applies, and permission rules from both files apply. The Agent SDK's `resolveSettings()` -helper always reads the file from the starting directory." Basis: -. Verified 2026-09-06 against Claude Code 2.1.263 and -that page as fetched that day. Recheck when the settings page names a different version, drops the -sentence, or a release note names where `settings.local.json` is read from. +plugin that names the boundary points here rather than restating it. The reader treats a +`settings.local.json` that a version before v2.1.211 left in the start directory as live beside the +root file: the root file wins any key both set, and the reader counts both files' permission rules. +It treats the Agent SDK's `resolveSettings()` helper as reading the start-directory file. + +- **Pointer**: for where `settings.local.json` is read from, see + [Where Claude Code keeps the local file in a git repository](https://code.claude.com/docs/en/settings#where-claude-code-keeps-the-local-file-in-a-git-repository). +- **As of**: 2026-09-06, Claude Code 2.1.263 +- **Recheck trigger**: that section names a different version boundary or drops the start-directory + case, or a release note names where `settings.local.json` is read from. ### The four start-directory conditions, and the one that is not detectable -The settings page states them in one sentence: the file stays with `.claude/settings.json` "outside a -git repository, when the repository root is your home directory, on Windows, or when the repository -root or its `.git` or `.claude` entry isn't owned by your user". All four are deterministic and the -reader resolves all four, naming which applied in the local-scope basis line. +The four conditions in the Scopes table's `local` row are the ones the same section documents. All +four are deterministic and the reader resolves all four, naming which applied in the local-scope +basis line. + +The Agent SDK helper above is **not** a fifth condition of the same kind. It is the +`resolveSettings()` helper, a standalone inspection function, not a class of running session. +Nothing observable inside a session distinguishes one that used the helper, and no documented +environment variable identifies an Agent SDK or headless run, so the reader states the limit rather +than guessing at it. -The Agent SDK sentence quoted above is **not** a fifth condition of the same kind. It describes the -`resolveSettings()` helper, a standalone inspection function, not a class of running session. Nothing -observable inside a session distinguishes one that used the helper, and no documented environment -variable identifies an Agent SDK or headless run, so the reader states the limit rather than guessing -at it. Basis: and -. Verified 2026-09-12. Recheck when either page documents an -entrypoint variable or changes the condition list. +- **Pointer**: [Where Claude Code keeps the local file in a git repository](https://code.claude.com/docs/en/settings#where-claude-code-keeps-the-local-file-in-a-git-repository) + and [Environment variables](https://code.claude.com/docs/en/env-vars). +- **As of**: 2026-09-12 +- **Recheck trigger**: either page documents an entrypoint variable or changes the condition list. ### Cloud sessions read a different scope set -"User and project local settings (`~/.claude/settings.json` and `.claude/settings.local.json`): not -read. Both stay on your machine, and the local file isn't in the clone." Only server-managed settings -reach a cloud session; a `managed-settings.json` file or MDM profile on the operator's device does -not. So a user-scope record in a cloud session describes the container's file, never the operator's, -and reporting it without that framing invites the wrong conclusion. +The reader treats user settings and project local settings (`~/.claude/settings.json` and +`.claude/settings.local.json`) as not read in a cloud session. Of the managed sources, the reader +counts only server-managed settings there; device files and MDM profiles stay on the operator's +machine. So a user-scope record in a cloud session describes the container's file, never the +operator's, and reporting it without that framing invites the wrong conclusion. -`CLAUDE_CODE_REMOTE` is the documented detection and the only entrypoint variable this reader branches -on: "Set automatically to `true` when Claude Code is running as a cloud session. Read this from a hook -or setup script to detect whether you are in a cloud session." `CLAUDE_CODE_ENTRYPOINT` is not -documented and is never read. Basis: ("Settings in cloud -sessions") and . Verified 2026-09-12. Recheck when the -cloud-session scope set changes or the variable is documented differently. +The reader detects a cloud session only by `CLAUDE_CODE_REMOTE` set to `true`, the only entrypoint +variable it branches on. `CLAUDE_CODE_ENTRYPOINT` is not documented and is never read. + +- **Pointer**: [Settings in cloud sessions](https://code.claude.com/docs/en/settings#settings-in-cloud-sessions) + and the `CLAUDE_CODE_REMOTE` row of [Environment variables](https://code.claude.com/docs/en/env-vars). +- **As of**: 2026-09-12 +- **Recheck trigger**: the cloud-session scope set changes or the variable is documented + differently. The managed scope is four surfaces. Two are the **portable core**, read on every OS: the per-OS `managed-settings.json` and its `managed-settings.d/` drop-in directory. Their merge order is documented rather than guessed, so the reader implements it instead of reporting an inventory: -Drop-in files merge as their own rule, not as a blanket "arrays concatenated, objects deep-merged". -`managed-settings.json` is the base; `*.json` files in the drop-in directory follow in alphabetical -order. A later single value replaces an earlier one, lists combine with duplicates removed, nested -blocks merge key by key, and `fallbackModel`, `modelPicker`, and a same-named -`extraKnownMarketplaces` or `managedMcpServers` entry are replaced whole. Hidden files are ignored. -Basis: [managed settings](https://code.claude.com/docs/en/managed-settings), "Split a file-based -policy across teams". Verified 2026-09-28. Recheck when that section changes how two drop-in files -combine a key. +The reader merges drop-in files by their own rule, not as a blanket "arrays concatenated, objects +deep-merged". `managed-settings.json` is the base; `*.json` files in the drop-in directory follow in +alphabetical order. The reader lets a later scalar overwrite an earlier one, unions arrays and +drops repeats, merges objects per key, and swaps in `fallbackModel`, `modelPicker`, and a +same-named `extraKnownMarketplaces` or `managedMcpServers` entry whole. It skips dotfiles. + +- **Pointer**: [Split a file-based policy across teams](https://code.claude.com/docs/en/managed-settings#split-a-file-based-policy-across-teams). +- **As of**: 2026-09-28 +- **Recheck trigger**: that section changes how two drop-in files combine a key. -Cross-source combination under `managedSourcesBehavior: "merge"` is by key kind. Lists union. Locks +The reader combines sources under `managedSourcesBehavior: "merge"` by key kind. Lists union. Locks take the strictest value. Restriction allowlists and values taken whole come from the highest source that sets them. `sandbox.credentials.awsPairs` and `sandbox.ripgrep` are values taken whole since v2.1.257. Provided MCP server names union, and the higher source's entry wins on a name clash. A -named set of keys is read from the highest source only. `env` merges per variable. Basis: -[settings reference](https://code.claude.com/docs/en/settings-reference) `managedSourcesBehavior`, -and [managed settings](https://code.claude.com/docs/en/managed-settings) "Compose every managed -source". Verified 2026-09-28. Recheck when that table gains or drops a key kind. +named set of keys is read from the highest source only. `env` merges per variable. + +- **Pointer**: [`managedSourcesBehavior`](https://code.claude.com/docs/en/settings-reference#managedsourcesbehavior) + and [Compose every managed source](https://code.claude.com/docs/en/managed-settings#compose-every-managed-source). +- **As of**: 2026-09-28 +- **Recheck trigger**: that table gains or drops a key kind. Two are **declared optional platform integrations**: the Windows policy registry keys and the macOS managed-preferences domain. Each is read where it is native and readable; where its tool is missing the surface reports `skipped` with a notice and every other result is unaffected. -`HKCU` is not a peer of `HKLM`. It is documented as lowest policy priority, used only when no -admin-level source exists, so the first key that **exists** ends the search and the rest are not -consulted. An existing key that yields nothing readable is reported unread, never as permission to -fall through. +`HKCU` is not a peer of `HKLM`. The reader treats it as lowest policy priority, used only when no +admin-level source exists (pointer: [Compose every managed source](https://code.claude.com/docs/en/managed-settings#compose-every-managed-source)), +so the first key that **exists** ends the search and the rest are not consulted. An existing key +that yields nothing readable is reported unread, never as permission to fall through. ## The one thing that is not a contest -> "Permission rules behave differently because they merge across scopes rather than override." +The reader treats permission rules as merging across scopes, never overriding (pointer: +[Settings precedence](https://code.claude.com/docs/en/permissions#settings-precedence) and +[Lists merge instead of overriding](https://code.claude.com/docs/en/settings#lists-merge-instead-of-overriding)). Every scope's rules are in effect at once. A rule text present in the **same list** at several scopes therefore has no winner and no loser. The entries are all live and identical in outcome. Naming one @@ -122,17 +138,15 @@ never a ranking. ## The one thing that is -> "Rules are evaluated in order: deny, then ask, then allow. The first match in that order determines -> the outcome, and rule specificity doesn't change the order." - -> "If a tool is denied at any level, no other level can allow it… The same holds across settings -> scopes: if user settings allow a permission and project settings deny it, the deny rule blocks it. -> The reverse is also true: a user-level deny blocks a project-level allow, because deny rules from -> any scope are evaluated before allow rules." +The reader evaluates deny, then ask, then allow, and takes the first match, whatever a rule's +specificity. A deny at any scope beats an allow at any other scope, in both directions, because deny +rules from every scope are evaluated before allow rules (pointer: +[Manage permissions](https://code.claude.com/docs/en/permissions#manage-permissions) and +[Settings precedence](https://code.claude.com/docs/en/permissions#settings-precedence)). The winner is decided by **kind**, and the mechanic is scope-independent in both directions. An -implementation that ranked scopes here would get the second sentence exactly backwards: `user` is the -lowest scope and its deny still wins. +implementation that ranked scopes here would get the cross-scope case exactly backwards: `user` is +the lowest scope and its deny still wins. ## `precedence_basis` vocabulary @@ -156,9 +170,9 @@ basis on it would read as a claim about what is in force, and instead names what ## Whole-tool rules -> "A bare tool name like `Bash` removes the tool from Claude's context entirely, so Claude never sees -> it… A scoped rule like `Bash(rm *)` leaves the tool available and blocks matching calls when Claude -> attempts them." +The reader treats a deny rule that is a bare tool name as removing the tool from the model's +context, and a scoped deny rule as leaving the tool available and blocking matching calls (pointer: +[Manage permissions](https://code.claude.com/docs/en/permissions#manage-permissions)). The tool token is the text before the first `(`; a rule that **is** its own token names the whole tool. That test needs no pattern matcher, so it is computed rather than caveated. @@ -166,8 +180,9 @@ tool. That test needs no pattern matcher, so it is computed rather than caveated - **A whole-tool deny makes every other rule for that tool inert**, whatever its kind. An inert deny is moot, not weakened. The tool is gone, so a second deny has nothing left to block. Reporting a scoped allow as effective underneath one would claim access to a tool that is not in context. -- **`EndConversation` is the documented exception**: "a deny rule can't remove it while any other tool - remains, and an ask rule never prompts for it." It is exempt from removal here. +- **`EndConversation` is the documented exception**: the reader never treats a deny rule as removing + it while any other tool remains, and never expects an ask rule to prompt for it. It is exempt from + removal here. - **A whole-tool ask outranks every scoped allow for that tool**, because it matches every call and ask is evaluated before allow. @@ -181,18 +196,18 @@ Neither is a limitation to apologize for; both change what a finding means. - **The command-line scope has no file.** Without `allowManagedPermissionRulesOnly`, `--settings`, `--allowedTools`, and `--disallowedTools` rank above local, project, and user settings, and no file reader can see them. The merge is the effective set the settings **files** define. With the lock - set in managed settings, `--allowedTools` is ignored, and `allow`, `ask`, and `deny` rules in user, - project, local, and `--settings` files are ignored. `--disallowedTools` and the session's deny and - ask rules still apply, including after a settings reload (v2.1.257+; before that version those - command-line and session rules were dropped at the first reload). The merge applies the lock only - when a managed `conf` record carries `allowManagedPermissionRulesOnly true`. Basis: - [settings reference](https://code.claude.com/docs/en/settings-reference#allowmanagedpermissionrulesonly). - Verified 2026-09-28. Recheck when that entry changes what the lock ignores. -- **Rules are compared by exact text, and the error direction is known.** "A broad deny rule like - `Bash(aws *)` blocks every matching call, including calls that also match a narrower allow rule like - `Bash(aws s3 ls)`." This merge does not evaluate pattern subsumption, so a narrow allow that a - broader deny blocks is still reported effective. It over-reports allow; it never over-reports - blocking. + set in managed settings, the merge drops `--allowedTools` and every rule the user, project, local + and `--settings` files carry, of all three kinds. It keeps `--disallowedTools` and deny or ask + rules added during the session, also across a settings reload (v2.1.257+; earlier versions lost + those command-line and session rules on the first reload). The merge applies the lock only + when a managed `conf` record carries `allowManagedPermissionRulesOnly true`. Pointer: + [`allowManagedPermissionRulesOnly`](https://code.claude.com/docs/en/settings-reference#allowmanagedpermissionrulesonly). + As of: 2026-09-28. Recheck trigger: that entry changes what the lock ignores. +- **Rules are compared by exact text, and the error direction is known.** A deny like + `Bash(aws *)` wins over each call it covers, a call that a narrower allow like `Bash(aws s3 ls)` + also covers included (pointer: [Manage permissions](https://code.claude.com/docs/en/permissions#manage-permissions)). + This merge does not evaluate pattern subsumption, so a narrow allow that a broader deny blocks is + still reported effective. It over-reports allow; it never over-reports blocking. - **A rule containing a literal newline or carriage return is reported, never split or stripped.** The records are line-oriented, so such a rule cannot be represented in one. Read line by line, a newline would yield two records, and the CRLF line-ending strip would silently delete a carriage @@ -207,39 +222,30 @@ Neither is a limitation to apologize for; both change what a finding means. ## The auto-mode entry diff -> "On entering auto mode, broad allow rules that grant arbitrary code execution are dropped: Blanket -> `Bash(*)` or `PowerShell(*)`; Wildcarded interpreters like `Bash(python*)`; Package-manager run -> commands; `Agent` allow rules; `Monitor` allow rules, because Claude Code runs Monitor commands -> through the shell. Narrow rules like `Bash(npm test)` carry over. Dropped rules are restored when -> you leave auto mode." -> -> Source: [permission-modes](https://code.claude.com/docs/en/permission-modes#eliminate-prompts-with-auto-mode), -> "How the classifier evaluates actions", re-fetched 2026-08-26. The `Monitor` category was added -> upstream in v2.1.236; before that version Monitor allow rules stayed in effect in auto mode. +The diff reports these allow rules as dropped on entering auto mode and restored on leaving it: +blanket `Bash(*)` or `PowerShell(*)`, wildcarded interpreters such as `Bash(python*)`, rules +letting a package manager run project scripts, any `Agent` allow rule, and any `Monitor` allow +rule (from v2.1.236; before that version a `Monitor` allow rule stayed in effect in auto mode). A +narrow rule such as `Bash(npm test)` carries over. + +- **Pointer**: [How auto mode evaluates actions](https://code.claude.com/docs/en/permission-modes#how-auto-mode-evaluates-actions). +- **As of**: 2026-08-26 +- **Recheck trigger**: a release note or that section adds, removes or narrows a dropped class. ### The precondition: which sessions enter auto mode at all -The starting-mode table lists seven run shapes and `auto` is the outcome in exactly one of them, so -the diff describes a transition the other six never make. Reporting it unconditionally hands a -headless run a verdict for a mode it never enters. +The diff applies only to a session whose starting mode resolves to `auto`. That turns on the CLI +version, how Claude Code is run, `disableAutoMode`, feature-flag fetching and the first session +after an upgrade, read at the pointer when the report is written; this file keeps no copy of that +table. Reporting the diff unconditionally hands a run that never enters auto mode a verdict for a +mode it never enters. -| How Claude Code runs | Starting mode | -| --- | --- | -| A Pro, Max, or Team plan, in a terminal or the VS Code extension | **`auto`** | -| `claude -p` or the Agent SDK | `default` | -| An Enterprise plan or a Claude Console API key | `default` | -| Bedrock, Google Cloud's Agent Platform, Microsoft Foundry, Claude Platform on AWS, apps gateway | `default` | -| Any settings file sets `disableAutoMode` to `"disable"` | `default` | -| Feature-flag fetching is off | `default` | -| First session after an install or upgrade | `default` | - -Basis: the starting-mode table in -. Verified 2026-09-12. Recheck when a release note -changes the starting mode, adds a run shape, or the table's plan conditions change. - -The vendor announcement that made auto mode the starting mode on those plans is dated 2026-08-14 -(); the reference documentation states the -standing condition rather than the date, and the condition is what the report needs. +- **Pointer**: for which mode a session starts in, see + [Which mode a session starts in](https://code.claude.com/docs/en/permission-modes#which-mode-a-session-starts-in) + (correlate with [the announcement](https://claude.com/blog/auto-mode-default-in-claude-code), 2026-08-14). +- **As of**: 2026-10-01 +- **Recheck trigger**: that section changes a starting mode, adds or removes a run shape, or moves + a version floor. Five documented classes, and every dropped rule is reported as exactly one of them: `blanket`, `interpreter-wildcard`, `package-manager-run`, `agent`, `monitor`. The shell-shape patterns are not @@ -259,21 +265,23 @@ to share, so the diff driver tests them on the tool token instead. version before acting on a `monitor` verdict. - **Only allow rules are in scope.** Deny and ask are evaluated before the classifier in every mode. -- **`autoMode.classifyAllShell` (v2.1.193+) inverts the carry-over answer.** When true it "suspend[s] - every Bash and PowerShell allow rule while auto mode is active", so a narrow `Bash(npm test)` does - **not** carry over. A diff that cannot see this key can be exactly wrong, which is why the reader - inventories it as a `conf` record. +- **`autoMode.classifyAllShell` (v2.1.193+) inverts the carry-over answer.** When true, the diff + treats every Bash and PowerShell allow rule as suspended while auto mode is active, so a narrow + `Bash(npm test)` does **not** carry over. A diff that cannot see this key can be exactly wrong, + which is why the reader inventories it as a `conf` record. Pointer: + [Route all shell commands through the classifier](https://code.claude.com/docs/en/auto-mode-config#route-all-shell-commands-through-the-classifier). - **The key is resolved only from scopes the classifier reads**: user settings, managed settings, and - inline `--settings`/SDK JSON. "The classifier doesn't read `autoMode` from project settings in - `.claude/settings.json` or `.claude/settings.local.json`." A project- or local-scope occurrence is - reported as having no effect, never obeyed. + inline `--settings`/SDK JSON, never `.claude/settings.json` or `.claude/settings.local.json`. A + project- or local-scope occurrence is reported as having no effect, never obeyed. Pointer: + [Where the classifier reads configuration](https://code.claude.com/docs/en/auto-mode-config#where-the-classifier-reads-configuration). - **A bare tool name is the broadest shell grant, not a surviving one.** `Bash` with no parentheses is strictly broader than `Bash(*)`, so it drops as `blanket`, the same treatment `Agent`'s bare form already had. Reporting it as kept would tell an operator their widest grant survives. - **Three scopes set `autoMode`; this reader can open two.** Inline `--settings` and Agent SDK JSON have no file, and `classifyAllShell` set there **inverts every shell verdict**. Every run states that bound, because the merge's command-line caveat covers rules and not this key. -- **`classifyAllShell: false` is the documented default**, not a type error. Only a value that is +- **The reader treats `classifyAllShell: false` as the default**, not a type error. Only a value that + is neither boolean is reported as malformed. - **The oracle is corroboration, not the read path.** `--oracle` spawns a real session to capture the harness's own `Ignoring dangerous permission … (bypasses classifier)` narration. Those are @@ -314,21 +322,21 @@ once, before a session, and naming the file the dead entry is in. different version histories, so a merged count would let an operator fix one and believe they had fixed all three. -| Check | Mechanic it follows from | +| Check | Firing rule, in our words, and the pointer to its mechanic | | --- | --- | -| `C2-autoMode` | "The classifier doesn't read `autoMode` from project settings in `.claude/settings.json` or `.claude/settings.local.json`." Before v2.1.207 it also read local settings, so a local-scope finding says so rather than implying it never worked | -| `C2-defaultMode` | **Claim:** `.claude/settings.json` and `.claude/settings.local.json` ignore `permissions.defaultMode` `auto` and `bypassPermissions`; user and managed settings read both, and `acceptEdits`, `plan`, `dontAsk`, `default`, and `manual` apply from any file in a terminal session. An ignored value still hides a user-scope one unless a higher-ranked settings file or `--permission-mode` sets a mode: `auto` falls to the built-in default, `bypassPermissions` to Manual, so the finding says to remove it from the named file. **Basis**, fetched 2026-09-29 as whole raw pages: [settings-reference#permissions-defaultmode](https://code.claude.com/docs/en/settings-reference#permissions-defaultmode), "`auto` and `bypassPermissions` don't take effect from project or local settings, so set them in `~/.claude/settings.json` instead. Before v2.1.257, `bypassPermissions` took effect from any file."; [permission-modes#which-mode-a-session-starts-in](https://code.claude.com/docs/en/permission-modes#which-mode-a-session-starts-in), "Claude Code then uses the built-in default rather than a `defaultMode` from `~/.claude/settings.json`" (for `auto`) and "If you set `"bypassPermissions"` in those two files, it doesn't take effect either, and the session starts in Manual mode. The other values apply from any settings file."; [permission-modes#start-in-a-different-mode](https://code.claude.com/docs/en/permission-modes#start-in-a-different-mode), "Sessions you start in a terminal honor every value except `auto` and `bypassPermissions`; sessions the VS Code extension starts don't read project settings for the starting permission mode" and "When more than one settings file sets `permissions.defaultMode`, settings precedence decides"; the [changelog](https://code.claude.com/docs/en/changelog) 2.1.257 entry, "Changed `defaultMode: "bypassPermissions"` in `.claude/settings.json` or `.claude/settings.local.json` to be ignored, like `"auto"`"; and, for the scopes that are read, [permission-modes#switch-permission-modes](https://code.claude.com/docs/en/permission-modes#switch-permission-modes), "`permissions.defaultMode: "bypassPermissions"` in user, `--settings`, or managed settings". **Version for `auto`: unverified, so the lint states none.** No page states a version for `auto` (checked: permission-modes, settings, settings-reference, auto-mode-config, managed-settings, changelog). **As of** 2026-09-29. **Recheck when** a fetch of settings-reference or permission-modes no longer carries those sentences, a page gives a version for `auto`, or the 2.1.257 changelog entry changes | -| `C2-planMode` | `useAutoModeDuringPlan` is "**Not read from shared project settings**". That names `.claude/settings.json` specifically, so a local-settings occurrence is **not** claimed dead, since doing so would assert a restriction no page states | -| `C5-disableType` | "set `permissions.disableBypassPermissionsMode` or `permissions.disableAutoMode` to `\"disable\"` in any settings file", the **string**. Checked at both documented key paths, in every scope; it is not managed-only | -| `C6-winPath` | "On Windows, paths are normalized to POSIX form before matching. `C:\Users\alice` becomes `/c/Users/alice`". Tested on the **shape**, a drive-letter or UNC prefix, never on the backslash character. Tested on the shape because a character test fails in both directions: the doubled JSON-source spelling is decoded away by `jq -r`, so a test on it is dead in the real pipeline, and a bare backslash is ordinary in shell rules (a regex, an escape, `\n`), so a test on it turns every rule into a severity-`error` finding and drowns the single true one. A UNC path gets its own message: the drive-letter remedy is wrong advice for it | -| `C6-contentField` | "You can't match a tool's primary content field this way: `command` for Bash and PowerShell, `file_path` for Read, Edit, and Write, `path` for Grep and Glob, `notebook_path` for NotebookEdit, and `url` for WebFetch… Claude Code ignores it and emits a startup warning" | -| `C6-allowParam` | "**Deny and ask rules** can match a top-level input parameter on any tool with `Tool(param:value)`… An allow rule for one parameter value wouldn't establish that the call is safe overall, so allow rules continue to use each tool's own specifier syntax." An operator writing one believes they narrowed a grant and has not. Fires only on parameters the page names for tools whose own syntax is a path or a command. `WebFetch(domain:host)` is the documented WebFetch form and `Bash(npm:*)` is a command prefix, so neither is distinguishable from a parameter by shape and neither fires | -| `C6-uncoveredPath` | "Claude Code checks file permissions against `Edit(path)` and `Read(path)` rules only. If you write a path rule for `Write`, `NotebookEdit`, `Glob`, or the legacy `MultiEdit` tool instead, Claude Code accepts the rule but never consults it, and warns at startup" (v2.1.210+; a `Glob` rule passed in `--allowedTools` is the stated exception) | -| `C6-colonStar` | "The `:*` form is only recognized at the end of a pattern. In a pattern like `Bash(git:* push)`, the colon is treated as a literal character". The mechanic is about **command-prefix** patterns, so a documented parameter form is exempt: "WebFetch rules use a `domain:` prefix… supports `*` wildcards", and firing on `WebFetch(domain:*.example.com)` called a documented, working rule broken. **Known gap:** in a deny or ask rule a mid-pattern `:*` with NO space after it, as in `Bash(git:*push)`, is not reported. It is structurally identical to the parameter form `Agent(model:*-haiku)`, so once the space is gone nothing in the rule text distinguishes them; the space was the only signal. The documented example is the space form, and the pages show no no-space mid-pattern rule anywhere. The exemption is by **grammar**, where in a deny or ask rule an `identifier:value` body is the parameter form, not by a list of parameter names: the page says parameter matching works "on any tool" for "any scalar parameter", so an allowlist could only ever chase it and would flag documented forms such as `Agent(model:*-haiku)` | -| `C6-malformed` | A rule is `Tool` or `Tool(specifier)`. Parentheses inside the specifier are literal, so `Edit(./Finance (2024)/**)` is one rule. Text after the closing parenthesis, as in `Bash(ls) x`, is a malformed Tool(content) rule: Claude Code reports it as invalid settings instead of matching it. An unclosed specifier is the same check. The lint used to keep only the token and treat `Bash(ls) x` as `Bash(ls)` | -| `C6-literalPath` | A Read or Edit path that is not a usable gitignore pattern, such as an unclosed `[`. A deny or ask rule guards that exact path. An allow rule approves nothing. One such deny used to fail every file edit; that is not the current behavior | -| `C6-malformed` | A rule is `Tool` or `Tool(specifier)`. Parentheses inside the specifier are literal, so `Edit(./Finance (2024)/**)` is one rule. Text after the closing parenthesis, as in `Bash(ls) x`, is a malformed Tool(content) rule: Claude Code reports it as invalid settings instead of matching it. An unclosed specifier is the same check. The lint used to keep only the token and treat `Bash(ls) x` as `Bash(ls)` | -| `C6-literalPath` | A Read or Edit path that is not a usable gitignore pattern, such as an unclosed `[`. A deny or ask rule guards that exact path. An allow rule approves nothing. One such deny used to fail every file edit; that is not the current behavior | +| `C2-autoMode` | Fires on an `autoMode` key in `.claude/settings.json` or `.claude/settings.local.json`, which the classifier does not read. Before v2.1.207 it also read local settings, so a local-scope finding says so rather than implying it never worked. Pointer: [Where the classifier reads configuration](https://code.claude.com/docs/en/auto-mode-config#where-the-classifier-reads-configuration) | +| `C2-defaultMode` | Fires on `permissions.defaultMode` `auto` or `bypassPermissions` in `.claude/settings.json` or `.claude/settings.local.json`, which we treat as ignored there (`bypassPermissions` from v2.1.257); user, `--settings` and managed settings read both, and `acceptEdits`, `plan`, `dontAsk`, `default`, and `manual` apply from any file in a terminal session. An ignored value still hides a user-scope one unless a higher-ranked settings file or `--permission-mode` sets a mode: `auto` falls to the built-in default, `bypassPermissions` to Manual, so the finding says to remove it from the named file. **Version for `auto`: unverified, so the lint states none.** No page stated a version for `auto` (checked: permission-modes, settings, settings-reference, auto-mode-config, managed-settings, changelog). **Pointer**: [`permissions.defaultMode`](https://code.claude.com/docs/en/settings-reference#permissions-defaultmode), [Common setups](https://code.claude.com/docs/en/permission-modes#common-setups), [Switch permission modes](https://code.claude.com/docs/en/permission-modes#switch-permission-modes), and the [2.1.257 changelog entry](https://code.claude.com/docs/en/changelog#2-1-257). **As of** 2026-09-29. **Recheck when** settings-reference or permission-modes changes which files honor `auto` or `bypassPermissions`, a page gives a version for `auto`, or the 2.1.257 changelog entry changes | +| `C2-planMode` | Fires on `useAutoModeDuringPlan` in `.claude/settings.json`, which we treat as not read from shared project settings. That covers `.claude/settings.json` only, so a local-settings occurrence is **not** claimed dead, since doing so would assert a restriction no page states. Pointer: [`useAutoModeDuringPlan`](https://code.claude.com/docs/en/settings-reference#useautomodeduringplan) | +| `C5-disableType` | Fires on `permissions.disableBypassPermissionsMode` or `permissions.disableAutoMode` set to anything other than the **string** `"disable"`, which is the only value we treat as a lock. Checked at both documented key paths, in every scope; it is not managed-only. Pointer: [`permissions.disableBypassPermissionsMode`](https://code.claude.com/docs/en/settings-reference#permissions-disablebypasspermissionsmode) and [`disableAutoMode`](https://code.claude.com/docs/en/settings-reference#disableautomode) | +| `C6-winPath` | Fires on a path rule whose path has a Windows drive-letter or UNC shape, since we treat Windows paths as normalized to POSIX form (`/c/...`) before matching, so the Windows spelling never matches. Pointer: [Read and Edit](https://code.claude.com/docs/en/permissions#read-and-edit). Tested on the **shape**, a drive-letter or UNC prefix, never on the backslash character. Tested on the shape because a character test fails in both directions: the doubled JSON-source spelling is decoded away by `jq -r`, so a test on it is dead in the real pipeline, and a bare backslash is ordinary in shell rules (a regex, an escape, `\n`), so a test on it turns every rule into a severity-`error` finding and drowns the single true one. A UNC path gets its own message: the drive-letter remedy is wrong advice for it | +| `C6-contentField` | Fires on a `Tool(param:value)` rule naming the tool's primary content field. The fields the lint checks, by tool: Bash and PowerShell `command`; Read, Edit, Write `file_path`; Grep, Glob `path`; NotebookEdit `notebook_path`; WebFetch `url`. We treat such a rule as ignored with a startup warning. Pointer: [Match by input parameter](https://code.claude.com/docs/en/permissions#match-by-input-parameter) | +| `C6-allowParam` | Fires on an **allow** rule in `Tool(param:value)` form, since we treat parameter matching as a deny-and-ask feature only; allow rules keep each tool's own specifier syntax. Pointer: [Match by input parameter](https://code.claude.com/docs/en/permissions#match-by-input-parameter). An operator writing one believes they narrowed a grant and has not. Fires only on parameters the page names for tools whose own syntax is a path or a command. `WebFetch(domain:host)` is the documented WebFetch form and `Bash(npm:*)` is a command prefix, so neither is distinguishable from a parameter by shape and neither fires | +| `C6-uncoveredPath` | Fires on a path rule whose tool is `Write`, `NotebookEdit`, `Glob`, or the retired `MultiEdit`, since we treat file access as decided by `Read(path)` and `Edit(path)` rules alone; such a rule is accepted, never consulted, and warned about at startup (v2.1.210+; a `Glob` rule passed in `--allowedTools` is the stated exception). Pointer: [Read and Edit](https://code.claude.com/docs/en/permissions#read-and-edit) | +| `C6-colonStar` | Fires on a mid-pattern `:*`, as in `Bash(git:* push)`, since we treat `:*` as recognized only at the end of a pattern and a mid-pattern colon as a literal character. Pointer: [Wildcard patterns](https://code.claude.com/docs/en/permissions#wildcard-patterns). The mechanic is about **command-prefix** patterns, so a documented parameter form is exempt: WebFetch's `domain:` prefix takes `*` wildcards ([WebFetch](https://code.claude.com/docs/en/permissions#webfetch)), and firing on `WebFetch(domain:*.example.com)` called a documented, working rule broken. **Known gap:** in a deny or ask rule a mid-pattern `:*` with NO space after it, as in `Bash(git:*push)`, is not reported. It is structurally identical to the parameter form `Agent(model:*-haiku)`, so once the space is gone nothing in the rule text distinguishes them; the space was the only signal. The documented example is the space form, and the pages show no no-space mid-pattern rule anywhere. The exemption is by **grammar**, where in a deny or ask rule an `identifier:value` body is the parameter form, not by a list of parameter names: we treat parameter matching as working for any scalar parameter on any tool ([Match by input parameter](https://code.claude.com/docs/en/permissions#match-by-input-parameter)), so an allowlist could only ever chase it and would flag documented forms such as `Agent(model:*-haiku)` | +| `C6-malformed` | We accept two rule shapes, a bare tool name or a tool name with one parenthesized specifier; a parenthesis within the specifier counts as an ordinary character, which makes `Edit(./Finance (2024)/**)` one rule. Text after the closing parenthesis, as in `Bash(ls) x`, is a malformed Tool(content) rule: Claude Code reports it as invalid settings instead of matching it. An unclosed specifier is the same check. The lint used to keep only the token and treat `Bash(ls) x` as `Bash(ls)` | +| `C6-literalPath` | A Read or Edit path that is not a usable gitignore pattern, such as an unclosed `[`. As a deny or ask rule it protects only that literal path; as an allow rule it grants nothing. One such deny used to fail every file edit; that is not the current behavior | +| `C6-malformed` | We accept two rule shapes, a bare tool name or a tool name with one parenthesized specifier; a parenthesis within the specifier counts as an ordinary character, which makes `Edit(./Finance (2024)/**)` one rule. Text after the closing parenthesis, as in `Bash(ls) x`, is a malformed Tool(content) rule: Claude Code reports it as invalid settings instead of matching it. An unclosed specifier is the same check. The lint used to keep only the token and treat `Bash(ls) x` as `Bash(ls)` | +| `C6-literalPath` | A Read or Edit path that is not a usable gitignore pattern, such as an unclosed `[`. As a deny or ask rule it protects only that literal path; as an allow rule it grants nothing. One such deny used to fail every file edit; that is not the current behavior | **`C5-disableType` is the highest-consequence check here.** A boolean is valid JSON, is accepted, and does nothing, so the operator believes auto mode is locked out and it is not. @@ -394,35 +402,34 @@ one feature. ## Ask rules under auto mode: carry this caveat on any `ask` finding -Any finding that rests on an `ask` rule prompting under auto mode carries this, named: - -> Content-scoped ask rules are evaluated before the classifier and always force a permission prompt, -> even in auto mode, because an explicit ask rule is your stated intent to be prompted for that -> action. The classifier cannot auto-approve a matching action. +Any finding that rests on an `ask` rule prompting under auto mode carries this caveat, named, in our +words: we treat a content-scoped ask rule as checked ahead of the classifier, so an action it +matches always reaches a human prompt in auto mode and never gets a classifier approval. -The sentence is on the [auto mode config page](https://code.claude.com/docs/en/auto-mode-config), not -the permissions page. The permissions page states the same mechanic for compound commands and -subshells: an ask rule such as `Bash(git clean *)` still prompts for `cd /tmp && git clean -f` or -`echo "$(git clean -f)"`, even in auto mode. That compound and subshell path was fixed in v2.1.257. +That mechanic is documented on the auto mode config page, not the permissions page. The permissions +page covers the related case of an ask rule matching one subcommand of a compound command or a +subshell, which still prompts in auto mode; that compound and subshell path was fixed in v2.1.257. It is not a claim that every ask-rule miss is fixed. -Upstream issue **#42797** ("Auto-mode ignores permissions.ask") is closed. **#83766** remains open and -still reports `permissions.ask` patterns auto-approved under `defaultMode: "auto"`. This plugin -follows the auto-mode config page, the source with a stated contract. A reader acting on an `ask` -finding should know #83766 is still open. +Upstream issue **#42797**, about auto mode ignoring `permissions.ask`, is closed. **#83766** remains +open and still reports `permissions.ask` patterns auto-approved under `defaultMode: "auto"`. This +plugin follows the auto-mode config page, the source with a stated contract. A reader acting on an +`ask` finding should know #83766 is still open. **What this changes in practice:** an `ask` rule is reported here as outranking an `allow`, and as surviving auto mode. Treat `ask` as a prompt you *expect*, not a guarantee you *rely on*, and use `permissions.deny` where the outcome must hold. This is not a defect in the reader. -**Basis:** [auto mode config](https://code.claude.com/docs/en/auto-mode-config) and -[permissions](https://code.claude.com/docs/en/permissions) (compound commands and subshells). Verified -2026-09-28. **Recheck when** #83766 closes or the auto-mode config page changes the "always force a -permission prompt" sentence. +- **Pointer**: [Add a human checkpoint](https://code.claude.com/docs/en/auto-mode-config#add-a-human-checkpoint) + and [Compound commands](https://code.claude.com/docs/en/permissions#compound-commands). +- **As of**: 2026-09-28 +- **Recheck trigger**: #83766 closes, or the auto-mode config page changes whether a content-scoped + ask rule always prompts. ## Managed policy, and what it does not buy -> "no other level, including command line arguments, can override a managed permission rule." +The reader treats a managed permission rule as one no lower scope, command line included, can +override (pointer: [Settings precedence](https://code.claude.com/docs/en/permissions#settings-precedence)). A managed rule cannot be removed by a lower scope. It does **not** follow that managed rules win every contest: a deny at any scope still beats an allow at managed, because deny is evaluated first @@ -436,15 +443,15 @@ any scope**. So a lower-scope deny changes the outcome of a managed allow withou | Verdict | What it rests on | | --- | --- | -| `enforced deny` | "If a tool is denied at any level, no other level can allow it." The strongest thing an administrator can write | +| `enforced deny` | a managed deny, which no other level can allow. The strongest thing an administrator can write | | `enforced allow` / `enforced ask` | the managed rule is highest and nothing beneath it outranks its kind | | `loosenable rule` | a lower scope carries an earlier-evaluated kind for the same rule text | -| `loosenable autoMode` | "A developer can extend `environment`, `allow`, `soft_deny`, and `hard_deny` with personal entries but can't remove entries that managed settings provide… a developer-added `allow` entry can override an organization `soft_deny` entry: the combination is additive, not a hard policy boundary." Permissions, hooks, MCP, sandbox-filesystem and sandbox-network each have an exclusivity lock; auto mode has none | +| `loosenable autoMode` | a managed `autoMode` block, which we treat as additive, not a policy boundary: developers may add their own entries to any of the four sections without being able to delete managed ones, and a developer's `allow` entry beats a managed `soft_deny` entry (pointer: [Where the classifier reads configuration](https://code.claude.com/docs/en/auto-mode-config#where-the-classifier-reads-configuration)). Permissions, hooks, MCP, sandbox-filesystem and sandbox-network each have an exclusivity lock; auto mode has none | | `enforced` / `loosenable lockout` | `disableAutoMode` is a real lock only when it carries the documented string `"disable"` | -The remedy the `autoMode` finding names is the page's own: "For actions that must never run regardless -of user intent or classifier configuration, use `permissions.deny` in managed settings, which… can't -be overridden." +The remedy the `autoMode` finding names is `permissions.deny` in managed settings, for an action that +must never run whatever the user or the classifier configuration says, since no lower scope can +override it (pointer: [Where the classifier reads configuration](https://code.claude.com/docs/en/auto-mode-config#where-the-classifier-reads-configuration)). **The report prescribes nothing.** It says what the consumer's policy does and does not achieve, and every rule string it prints came from a file it read, a property the suite asserts positively rather @@ -456,6 +463,6 @@ which keeps it neutral by construction rather than by restraint. report means the local admin surfaces; the cache is not folded in. The failure read is the Organization policy line in `/status`. A `skipped` or `unreadable` surface gets its own note stating that it is not evidence no policy is deployed there. An administrator reading silence as "no policy" -is the failure this report exists to prevent. Basis: -[server-managed settings](https://code.claude.com/docs/en/server-managed-settings). Verified -2026-09-28. Recheck when that page moves the cache path or the `/status` failure line. +is the failure this report exists to prevent. Pointer: +[Fetch and caching behavior](https://code.claude.com/docs/en/server-managed-settings#fetch-and-caching-behavior). +As of: 2026-09-28. Recheck trigger: that page moves the cache path or the `/status` failure line. diff --git a/plugins/harness-config/skills/audit-prompting-postures/evals/evals.json b/plugins/harness-config/skills/audit-prompting-postures/evals/evals.json index 0fa1c3ece6..faaf477b8a 100644 --- a/plugins/harness-config/skills/audit-prompting-postures/evals/evals.json +++ b/plugins/harness-config/skills/audit-prompting-postures/evals/evals.json @@ -98,6 +98,42 @@ "Reports a reversal as a row with its corrected verdict instead of discarding it", "Ends with an attestation line that names surface batches and says the verdicts check ran" ] + }, + { + "id": 9, + "name": "p12-binds-only-a-component-pinned-at-xhigh-or-max", + "prompt": "/harness-config:audit-prompting-postures agents. I have two coding agents. One sets effort: max in its frontmatter and says nothing about when to stop once its checks pass. The other pins no effort level at all. Which of them is missing P12?", + "expected_output": "P12 is MISSING on the max-effort agent: it is code-changing, it pins max, and it does not state where its run ends or rule out steps of its own after that point. The proposed addition carries the Sonnet 5.5 model condition, because the behavior P12 guards against is documented for that model only, and wording comes from a live fetch of the Sonnet 5.5 guide. P12 is NOT-APPLICABLE on the agent that pins no effort level, with the failed predicate named. A review the repository requires, such as a mandated fresh-context verifier, is not something the proposal tells the agent to drop.", + "files": [], + "expectations": [ + "Returns MISSING for P12 on the agent pinned at max, NOT-APPLICABLE on the unpinned one with its failed predicate", + "The proposed addition states the Sonnet 5.5 condition rather than applying unconditionally", + "Does not propose removing a review the repository requires" + ] + }, + { + "id": 10, + "name": "p13-done-claim-needs-a-runnable-check", + "prompt": "/harness-config:audit-prompting-postures skills. My fix-bug skill edits code and ends by telling the user the fix is done. It never says which test or build to run, or what to report if none can run. Is it missing anything?", + "expected_output": "P13 is MISSING: the skill is code-changing, it names no check that exercises the change for its done claim to depend on, and its done report does not name a check and its result. The proposed addition asks for both, carries no model condition because the catalog treats P13 as model-neutral, and takes its wording from a live fetch of the Sonnet 5.5 guide's verification section rather than from the catalog. P3 is judged separately and only if the flow involves making tests pass.", + "files": [], + "expectations": [ + "Returns MISSING for P13 and names both missing parts: the runnable check behind the done claim, and the check and its result in the done report", + "The proposal carries no model condition", + "Does not fold the finding into P3 or P5" + ] + }, + { + "id": 11, + "name": "p14-ideas-first-present-via-a-human-gate", + "prompt": "/harness-config:audit-prompting-postures skills. My brainstorm skill produces a list of options for the user and then asks which one to pursue before doing anything else. Does it need P14?", + "expected_output": "The skill is ideating, so P14 applies, and it is PRESENT: the skill delivers the options and stops at a human gate before building anything, which the catalog counts as satisfying P14. The verdict cites where the skill stops. No addition is proposed.", + "files": [], + "expectations": [ + "Classifies the skill as ideating and judges P14", + "Returns PRESENT, citing the human gate after the options step", + "Proposes no addition" + ] } ] } diff --git a/plugins/harness-config/skills/audit-prompting-postures/reference/postures.md b/plugins/harness-config/skills/audit-prompting-postures/reference/postures.md index 414d3164e4..12c8eea3a6 100644 --- a/plugins/harness-config/skills/audit-prompting-postures/reference/postures.md +++ b/plugins/harness-config/skills/audit-prompting-postures/reference/postures.md @@ -1,6 +1,6 @@ # Posture catalog -Eleven postures. Each row: the applicability predicate (which component purposes it binds), what +Fourteen postures. Each row: the applicability predicate (which component purposes it binds), what counts as present, and the guide pointer that owns the recommended wording. Pointers only. Wording is fetched live per SKILL.md Phase A; the recheck trigger for every row is a change to its cited section. @@ -8,35 +8,66 @@ cited section. Guide root: (sections cited by heading). Model subpages cited by page + heading where a row needs one. -Two model subpages exist, and a row that names the "Fable 5 subpage" means -. -Its sibling - -covers Claude Fable 5.1 and Claude Mythos 5.1 and carries its own headings, among them "Consider -all effort levels", "Finish the whole task", "Keep changes and tests to what the task asks for", -and "Let the lead agent keep working while subagents run". When the audited component targets -Fable 5.1, read the 5.1 sibling as well as the heading a row names, and cite whichever page -carries the wording the proposal uses. Every heading the rows below name is present on the Fable 5 -page. Verified 2026-09-06 against Claude Code 2.1.263 and both subpages as fetched that day. -Recheck when a row's cited heading disappears from the Fable 5 page, when a newer model subpage -appears beside these two, or when the best-practices page's model-guidance table gains a row. - -A row that names the "Opus 5.5 subpage" means -, -and the "Opus 5.5 usage guide" means the vendor blog - (published 2026-09-22), which states -the same run-shaping advice for CLAUDE.md and Claude Code sessions. Both are fetched lazily, like -the other subpages. The behaviors they describe (named stops, a finish line, a task file, a report -that leads with what the human owes) are model-neutral, so proposals citing them carry no model -condition. Verified 2026-09-23 against the subpage's raw `.md` (28,311 bytes); recheck when a -cited heading disappears from either page. +Rows cite the subpages of the current models first. A subpage name in a row means: + +- "Fable 5.1 subpage": + +- "Opus 5.5 subpage": + +- "Sonnet 5.5 subpage": + +- "Fable 5 subpage": + + (P5 only; see that row). +- "Opus 5 subpage" and "Opus 4.8 subpage": + + and + . + We keep them in P1 while Claude Code can still put a session on either model. Pointer: for the + models a session can fall back to, see + [model configuration: automatic model fallback](https://code.claude.com/docs/en/model-config#automatic-model-fallback). + As of: 2026-10-01. Recheck trigger: either model leaves Claude Code's model page. + +- **Pointer**: for which model each subpage covers, see the + [model-specific guidance table](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices#model-specific-guidance). +- **As of**: 2026-10-01 (each subpage read as raw markdown that day). +- **Recheck trigger**: a row's cited heading disappears from its subpage, or that table gains or + drops a row. + +A row's "correlate:" note names a heading in the Opus 5.5 usage guide +(correlate with , published 2026-09-22), +a vendor blog that is never the pointer; the Opus 5.5 subpage sections linked below are. Both are +fetched lazily, like the other subpages. We treat the behaviors the Opus 5.5 rows check (named +stops, a finish line, a task file, a report that leads with what the human owes) as model-neutral, +so proposals citing them carry no model condition. + +- **Pointer**: the Opus 5.5 subpage's + [capabilities relevant to prompting](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#capability-improvements) + (the rendered page gives this heading the id `capability-improvements`) and + [unattended agentic runs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#unattended-agentic-runs). +- **As of**: 2026-09-23 (our probe: the subpage's raw `.md`, 28,311 bytes; no artifact stored). + Both anchors re-matched against the rendered page on 2026-10-01. +- **Recheck trigger**: a cited heading disappears from either page. + +The Sonnet 5.5 subpage backs P2, P6, P8, P12, P13 and P14. We treat P13 and P14 as model-neutral, +so their proposals carry no model condition. P12 is the one model-conditional row: its proposal +carries the Sonnet 5.5 condition. + +- **Pointer**: the Sonnet 5.5 subpage's + [steer initiative and scope](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#steer-initiative-and-scope) + (P2, P6, P12, P14), + [mid-turn user messages](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#mid-turn-user-messages-and-task-budgets) + (P8) and + [verification on coding tasks](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#verification-on-coding-tasks) + (P13). +- **As of**: 2026-10-01 +- **Recheck trigger**: a cited heading disappears from that subpage, or another model's subpage + covers self-started review rounds, which re-opens P12's model condition. A row that names the "Sonnet 5.5 subpage" means . -It is fetched lazily, like the other subpages. The behaviors its cited headings address (a stop -rule that names when to ask, scope held to the request, a real check before "done") are -model-neutral, so proposals citing it carry no model condition. Verified 2026-10-01 against the -subpage's raw `.md` (27,412 bytes); recheck when a cited heading disappears. +It is fetched lazily, like the other subpages. Which of its rows carry a model condition is stated +above: only P12. ## Purpose classification vocabulary @@ -55,6 +86,8 @@ Classify each component by what its body has the model DO (multiple or none): - **multi-window**: spans sessions/windows via saved state, handoffs, or resumability - **parallelism-steering**: instructs when/how to parallelize tool calls - **user-gated**: interactive flow with genuine decision gates only the user can answer +- **ideating**: the body's deliverable is a set of proposals for the user to choose from, not the + built thing ## Postures @@ -67,17 +100,21 @@ Classify each component by what its body has the model DO (multiple or none): before accepting it and consolidate the results into one table. - **Pointer:** main page, "Subagent orchestration"; Opus 5 subpage, "Controlling subagent spawning"; Opus 4.8 subpage, "Controlling subagent spawning"; Opus 5.5 subpage, "Capabilities - relevant to prompting" (audits and migrations run with parallel subagents); Opus 5.5 usage - guide, "Ask it to split big work across subagents". + relevant to prompting" (audits and migrations run with parallel subagents); correlate: Opus 5.5 + usage guide, "Ask it to split big work across subagents". ### P2: Minimal-scope guardrail - **Predicate:** code-changing. -- **Present when:** the component bounds scope to what was asked (no unrequested features, - abstractions, defensive code, or cleanup beyond the task). -- **Pointer:** main page, "Overeagerness"; Fable 5 subpage, "Consider all effort levels" - (anti-overengineering block); Sonnet 5.5 subpage, "Steer initiative and scope" (unrequested - additions). +- **Present when:** the component limits its changes to what the request or approved plan covers, + and sends anything else it notices to its report rather than into the change. +- **Pointer:** main page, + [Overeagerness](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices#overeagerness); + Fable 5.1 subpage, + [Keep changes and tests to what the task asks for](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#keep-changes-and-tests-to-what-the-task-asks-for); + Sonnet 5.5 subpage, + [Steer initiative and scope](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#steer-initiative-and-scope) + (unrequested additions). ### P3: Anti-test-gaming guardrail @@ -98,8 +135,12 @@ Classify each component by what its body has the model DO (multiple or none): - **Predicate:** long-running. - **Present when:** the component ties progress/status claims to tool-result evidence and requires naming unverified work as unverified. -- **Pointer:** Fable 5 subpage, "Ground progress claims during long runs"; Sonnet 5.5 subpage, - "Verification on coding tasks" (a change reported done with no check that exercised it). +- **Pointer:** Fable 5 subpage, + [Ground progress claims during long runs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5#ground-progress-claims-during-long-runs). + On 2026-10-01 that was the only docs section covering this check: we read the Fable 5.1, Opus 5.5 + and Sonnet 5.5 subpages that day (raw markdown, no artifact stored) and none of them covers tying + progress claims to tool-result evidence. As of: 2026-10-01. Recheck trigger: a current model's + subpage starts covering the check (repoint there), or that Fable 5 section disappears. ### P6: Autonomy or checkpoint posture @@ -108,16 +149,20 @@ Classify each component by what its body has the model DO (multiple or none): CLAUDE.md or natively read AGENTS.md is long-running when the repo carries long-running components or its own text invites long runs. - **Present when:** an autonomous component states its finish line (what "done" observably is, or - that the dispatching brief must state it) and names both kinds of stop: keep going when a step - needs no input, with status notes in the same message as the next action rather than a summary - that names the next step, an offer to continue, or a list of non-blocking choices; stop and ask - only when nothing can move without the human, or before a destructive, hard-to-undo, or outward - action. The keep-going half never licenses turning permission prompts or P7's gates off. An - interactive one names the gates worth stopping at. -- **Pointer:** Fable 5 subpage, "Rare cases of early stopping" (autonomous) and "Strong - instruction following" (checkpoint block); Opus 5.5 subpage, "Unattended agentic runs"; Opus 5.5 - usage guide, "Say what 'done' looks like, then let it run" and "Tell it which stops you want"; - Sonnet 5.5 subpage, "Steer initiative and scope" (carrying work through at lower effort). + that the dispatching brief must state it) and names both kinds of stop: when it carries on + without asking, reporting progress alongside continued work rather than as a turn-ending + summary, and the cases where it must stop for the human (no way forward without them, or an + action that is destructive, hard to undo, or reaches outside the workspace). The keep-going half + never licenses turning permission prompts or P7's gates off. An interactive one names the gates + worth stopping at. +- **Pointer:** Fable 5.1 subpage, + [Finish the whole task](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#finish-the-whole-task) + (autonomous); Sonnet 5.5 subpage, + [Steer initiative and scope](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#steer-initiative-and-scope) + (carrying work through); Opus 5.5 subpage, + [Unattended agentic runs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#unattended-agentic-runs); + correlate: Opus 5.5 usage guide, "Say what 'done' looks like, then let it run" and "Tell it which + stops you want". ### P7: Destructive-action confirmation @@ -135,18 +180,20 @@ Classify each component by what its body has the model DO (multiple or none): ### P8: Context-budget reassurance - **Predicate:** context-surfacing. -- **Model condition:** the guide section this row points at scopes the underlying capability by - model: "Claude Sonnet 5, Claude Sonnet 4.6, Claude Sonnet 4.5, and Claude Haiku 4.5 feature - context awareness", main page, "Context awareness and multiwindow workflows" (fetched - 2026-08-12). Components here run on any consumer model, so per SKILL.md Gotchas +- **Model condition:** we treat the underlying capability as model-scoped. Components here run on + any consumer model, so per SKILL.md Gotchas ("Model-conditional postures stay conditional") the proposal must be model-neutral or carry that - same condition. Re-read the section's own model list on the run's live fetch rather than trusting - this one. The list is the guide's to change, and the recheck trigger for this row is a change to - it. + model condition. Read the model list on the run's live fetch; this file keeps no copy. + Pointer: for which models the capability covers, see + [context windows: context awareness](https://platform.claude.com/docs/en/build-with-claude/context-windows#context-awareness). + As of: 2026-10-01. Recheck trigger: that section's model list changes. - **Present when:** the surfaced figure is accompanied by do-not-wrap-up-early framing (or the component deliberately avoids surfacing raw countdowns at all, the stronger form). -- **Pointer:** main page, "Context awareness and multiwindow workflows"; Fable 5 subpage, "Rare - cases of context-budget concern". +- **Pointer:** main page, + [Context awareness and multiwindow workflows](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices#context-awareness-and-multiwindow-workflows); + Sonnet 5.5 subpage, + [Mid-turn user messages](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#mid-turn-user-messages-and-task-budgets) + (countdowns after tool results). ### P9: Multi-window state guidance @@ -157,8 +204,8 @@ Classify each component by what its body has the model DO (multiple or none): task list in a file, ticked as items finish and extended with new ones found, read instead of the scrollback. An existing ledger or state file that does this satisfies it. - **Pointer:** main page, "Workflows across multiple context windows" and "State management best - practices"; Opus 5.5 subpage, "Unattended agentic runs"; Opus 5.5 usage guide, "Keep the task - list in a file". + practices"; Opus 5.5 subpage, "Unattended agentic runs"; correlate: Opus 5.5 usage guide, "Keep + the task list in a file". ### P10: Parallel-tool-call steering @@ -174,5 +221,47 @@ Classify each component by what its body has the model DO (multiple or none): decisions, changes to approve), then what changed and what was found, for example under the headings "Blocked on me", "Changed", "Found". An existing report shape that puts the human's items first satisfies it; adapt that shape rather than adding a second one. -- **Pointer:** Opus 5.5 subpage, "Capabilities relevant to prompting" (communication); Opus 5.5 - usage guide, "Read what it needs from you first". +- **Pointer:** Opus 5.5 subpage, "Capabilities relevant to prompting" (communication); correlate: + Opus 5.5 usage guide, "Read what it needs from you first". + +### P12: No self-started review rounds at xhigh or max effort + +- **Predicate:** code-changing or orchestrating, AND the component pins `xhigh` or `max` effort (its + `effort:` frontmatter) or its text sets one of those levels for its own work. A component that + pins nothing is NOT-APPLICABLE, whatever level a session might run it at. +- **Model condition:** the only page section behind this row is in the Sonnet 5.5 subpage, so per + SKILL.md Gotchas the proposal carries that model's condition, for example "when running on + Sonnet 5.5 at `xhigh` or `max`". Never propose it unconditionally. +- **Present when:** the component, for runs at `xhigh` or `max`, states where its run ends and adds + no step of its own after that point. For the steer this checks for, see the pointer. A check the + repository requires (a mandated fresh-context verifier, a merge gate) is part of the run, not a + step after its end. +- **Pointer:** Sonnet 5.5 subpage, + [Steer initiative and scope](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#steer-initiative-and-scope) + (its paragraph on the top two effort levels). + +### P13: Runnable check behind a done claim + +- **Predicate:** code-changing. +- **Present when:** the component names a check that executes the change (a repository command, a + verification skill, or a gate), makes a done claim depend on it, and has its report name the + check and its result. For the steer this checks for, see the pointer. Handing the check to a + named skill or gate satisfies the row. We treat this as model-neutral, so the proposal carries + no model condition. +- **Why a new row and not P3 or P5:** P3 binds only a flow that makes tests pass and asks what the + fix targets. P5 binds long-running components and asks about progress claims in general. This + row binds every code-changing component and asks one thing: that "done" rests on a check that + ran, named in the report. +- **Pointer:** Sonnet 5.5 subpage, + [Verification on coding tasks](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#verification-on-coding-tasks). + +### P14: Ideas-first on open-ended requests + +- **Predicate:** ideating. +- **Present when:** in the component's text, the step that delivers its proposals is followed by a + wait for the user's choice, and every step that writes or edits files comes after that wait. An + explicit end of turn, a human gate, or a choose-then-proceed instruction each counts as the wait. + We treat this as model-neutral, so the proposal carries no model condition. +- **Pointer:** Sonnet 5.5 subpage, + [Steer initiative and scope](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#steer-initiative-and-scope) + (open-ended requests). diff --git a/plugins/harness-config/skills/audit/reference/audit-checklist.md b/plugins/harness-config/skills/audit/reference/audit-checklist.md index 5b504c4afa..b918137b2b 100644 --- a/plugins/harness-config/skills/audit/reference/audit-checklist.md +++ b/plugins/harness-config/skills/audit/reference/audit-checklist.md @@ -92,11 +92,11 @@ taken, the finding is stated conditionally, not asserted. | --- | --- | --- | | All hook scripts exist on disk | error | Resolve `$CLAUDE_PROJECT_DIR` to the project root, check file exists | | Hook scripts are readable | error | `[[ -r ]]` | -| `timeout` is a seconds value, not milliseconds | warning | The [hooks reference](https://code.claude.com/docs/en/hooks) states for `timeout`: "Seconds before canceling. Defaults: 600 for `command`, `http`, and `mcp_tool`; 30 for `prompt`; 60 for `agent`." Flag a **recognizably millisecond-scale** value: a round thousands multiple such as `30000` or `120000`, which read as seconds are 8 h and 33 h. Do NOT flag merely-large values: the page documents defaults, not a maximum, so a deliberately long-running hook may legitimately exceed 600. When the value is large but not millisecond-shaped, corroborate against the hook's expected runtime before reporting anything | +| `timeout` is a seconds value, not milliseconds | warning | We read `timeout` as seconds. Flag a **recognizably millisecond-scale** value: a round thousands multiple such as `30000` or `120000`, which read as seconds are 8 h and 33 h. Do NOT flag merely-large values: the reference sets per-type defaults, not a maximum, so a deliberately long-running hook may legitimately run longer than its default. When the value is large but not millisecond-shaped, corroborate against the hook's expected runtime before reporting anything. Pointer: for the unit and the per-type defaults, see [hooks: common fields](https://code.claude.com/docs/en/hooks#common-fields). As of: 2026-10-01. Recheck trigger: that field's unit changes, or the reference adds a maximum | | Timeouts are reasonable (5-15s for simple formatters, 30s for slow-startup tools such as pwsh) | warning | Compare against known good values. These figures are this skill's judgment, not a documented limit | | Matcher takes its intended evaluation path | warning | Classify the matcher by its characters, confirm exact-match vs regex matches intent, anchor regex-path matchers with `^…$` | -| Shell-form path placeholders are quoted | warning | Same page: "Prefer exec form for any hook that references a path placeholder. In shell form, wrap each placeholder in double quotes." Flag a shell-form hook whose project/plugin placeholder is unquoted, in **either** spelling, braced (`${CLAUDE_PROJECT_DIR}`, `${CLAUDE_PLUGIN_ROOT}`, `${CLAUDE_PLUGIN_DATA}`) or bare-dollar (`$CLAUDE_PROJECT_DIR`), since both reach the shell and an unquoted path breaks on a space either way. Report exec form as the page's preference, but do **not** flag shell form itself: the page endorses omitting `args` when the hook needs pipes, `&&`, redirects, or a `.cmd`/`.bat` shim, and quoted shell form is a documented, correct spelling | -| Exec-form `command` resolves on every platform the repo targets | error | Same page: "On Windows, exec form requires `command` to resolve to a real executable such as a `.exe`." Windows-only constraint: `bash` and `sh` are real executables on macOS/Linux, so flag only for a repo that runs on Windows. There `bash` resolves to the WSL relay `System32\bash.exe` and the launch fails; a failed launch is a non-blocking error, so a gate hook silently enforces nothing. A `.sh` path or a `.cmd`/`.bat` shim as `command` is the same defect: exec form has no shell to honor a shebang. Do not accept bare `bash` with the script in `args` as the fix. Fixes per the page: `"command": "node"` with the script path in `args`, or shell form with `"shell": "bash"`, which Claude Code routes through Git Bash instead of a PATH lookup. **Claim:** the quoted Windows sentence, and that exec form spawns with no shell. **Basis:** [hooks](https://code.claude.com/docs/en/hooks) "Exec form and shell form", the quoted span, fetched as raw `hooks.md` on 2026-09-28 (330,813 bytes, SHA-256 `57e3b47d55acfbae3dcdc112866c8c0f75528d8b5c4fca9bfcdaa904d4728218`). **As of:** 2026-09-28. **Recheck:** that Windows sentence changes, or [anthropics/claude-code#90495](https://github.com/anthropics/claude-code/issues/90495) closes (args dropped, the hook still routed through `bash.exe`). This marketplace's `scripts/check-exec-form-windows-probe.sh` checks its own rows; a non-Windows skip does not clear #90495. | +| Shell-form path placeholders are quoted | warning | Flag a shell-form hook whose project/plugin placeholder is unquoted, in **either** spelling, braced (`${CLAUDE_PROJECT_DIR}`, `${CLAUDE_PLUGIN_ROOT}`, `${CLAUDE_PLUGIN_DATA}`) or bare-dollar (`$CLAUDE_PROJECT_DIR`), since both reach the shell and an unquoted path breaks on a space either way. Do **not** flag shell form itself, and do not flag quoted shell form. For a hook that names a path placeholder, the report may suggest exec form as an option; which form fits a given hook is the pointer's to say, not this row's. Pointer: for the quoting rule, see [hooks: reference scripts by path](https://code.claude.com/docs/en/hooks#reference-scripts-by-path); for when each form fits, see [hooks: exec form and shell form](https://code.claude.com/docs/en/hooks#exec-form-and-shell-form). As of: 2026-10-01. Recheck trigger: either section changes its quoting rule or its form preference | +| Exec-form `command` resolves on every platform the repo targets | error | Flag only for a repo that runs on Windows: an exec-form hook whose `command` is not a real executable that runs the hook there. `bash`, `sh`, a `.sh` path, and a `.cmd`/`.bat` shim are examples, not the whole set. On macOS and Linux the same hook is not a finding. We rate it `error` because a gate hook that does not launch enforces nothing. Do not accept bare `bash` with the script in `args` as the fix. The fixes we accept: `"command": "node"` with the script path in `args`, or shell form with `"shell": "bash"`. For why each of these fails or works on Windows, follow the pointer. Pointer: for the Windows exec-form rule, see [hooks: exec form and shell form](https://code.claude.com/docs/en/hooks#exec-form-and-shell-form). As of: 2026-09-28 (our probe: raw `hooks.md`, 330,813 bytes, SHA-256 `57e3b47d55acfbae3dcdc112866c8c0f75528d8b5c4fca9bfcdaa904d4728218`). Recheck trigger: that section's Windows rule changes, or [anthropics/claude-code#90495](https://github.com/anthropics/claude-code/issues/90495) closes (args dropped, the hook still routed through `bash.exe`). This marketplace's `scripts/check-exec-form-windows-probe.sh` checks its own rows; a non-Windows skip does not clear #90495 | | No duplicate hooks (same script registered twice for same event) | info | Compare commands within each event | | Hook events are valid per official docs | error | The engine reads the Event table of the [hooks reference](https://code.claude.com/docs/en/hooks) each run (`hook-event` rows). Read the page yourself only for a `not-inspectable` row, live rather than from a recalled event list | | Hook-suppression levers are read and reported | info | Read `disableAllHooks` from the settings-declared layer and `allowManagedHooksOnly` / `strictPluginOnlyCustomization` from the managed layer. Report each as set or unset with the hooks it switches off. `info` because the reading is state, not a defect: a repo may set any of them deliberately. It is not optional, though: **B.1–B.3's third narrowing may not downgrade a missing deny rule on the strength of a hook one of these has already disabled**, so an unread lever means the narrowing is unavailable rather than assumed clear | @@ -108,14 +108,14 @@ taken, the finding is stated conditionally, not asserted. | --- | --- | --- | | Every plugin's marketplace exists in `extraKnownMarketplaces` | error | Split `@marketplace` suffix, verify marketplace key exists | | No enabled plugin depends on a disabled plugin | warning | The engine merges user, project and local `enabledPlugins` and reads each enabled plugin's `.claude-plugin/plugin.json` `dependencies` (direct only). A merged `false` that an enabled plugin declares is `dependency-disabled`; every other `false` is an inventory row. See [validation-categories.md](../context/validation-categories.md) "E.1" | -| Every enabled plugin has a component definition, and none has two | error | Read the [Strict mode section](https://code.claude.com/docs/en/plugins/marketplace-reference#strict-mode) and the entry rules above it on marketplace-reference; do not restate them from memory. **Claim:** a plugin with no `plugin.json` is defined by its marketplace entry whatever `strict` says, so a bare root `marketplace.json` is not by itself a finding; an entry that declares component fields beside a `plugin.json` under `strict: false` fails to load with `Plugin has conflicting manifests`. Error conditions: a plugin with no `plugin.json` and no entry declaring its components (nothing defines what loads), and that conflict. **Basis:** the section's "The entry is the manifest" row and conflict message, pinned in `doc-citations.tsv`, fetched as raw markdown 2026-09-29. **As of:** 2026-09-29. **Recheck:** `check-doc-citations.sh` fails on either span, or the section's result table changes | +| Every enabled plugin has a component definition, and none has two | error | Read the [Strict mode section](https://code.claude.com/docs/en/plugins/marketplace-reference#strict-mode) and the entry rules above it on marketplace-reference; do not restate them from memory. Do not flag a plugin with no `plugin.json` whose marketplace entry declares its components, whatever `strict` says; a bare root `marketplace.json` is not by itself a finding. Error conditions: a plugin with no `plugin.json` and no entry declaring its components (nothing defines what loads), and an entry that declares component fields beside a `plugin.json` under `strict: false`. For what each case does at load, follow the pointer. Pointer: that section's result table and its conflict error, both spans pinned in `doc-citations.tsv`. As of: 2026-09-29. Recheck trigger: `check-doc-citations.sh` fails on either span, or the section's result table changes | ## F. Environment Variables | Check | Severity | How to verify | | --- | --- | --- | -| Env vars in settings.json are documented CC vars | warning | Read `code.claude.com/docs/en/env-vars` and search it for each env var name. WebSearch alone is insufficient. **Read it verbatim, not through a summarizer.** The page is long and a summarizing fetch truncates it, then reports the rows past the cutoff as absent: `curl https://code.claude.com/docs/en/env-vars.md` to a file and grep the file, per the [fetch route](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/upstream-drift/README.md#reading-the-basis-the-fetch-route). A truncated read supports NO finding. Say so and move on. **Absence from this page is not "unrecognized" either:** it is not the whole env-var surface, and `CLAUDE_CODE_ENTRYPOINT`, `CLAUDE_CODE_ENHANCED_TELEMETRY_BETA`, and `CLAUDE_CODE_EXPERIMENTAL_OBSERVER_AGENTS` are all real and all absent from it (verified 2026-08-10 on a full verbatim read; recheck when `env-vars` gains a row for any of them). A name missing here is at most "not documented on `env-vars`". Check `monitoring-usage`, `mcp`, and `settings` before writing anything stronger | -| No secrets in settings.json (tokens, keys, passwords) | error | Scan for patterns: `ghp_`, `eyJ`, `sk-`, `AKIA`, common token prefixes. The engine's `SECRET_RE` is a literal list of vendor token shapes, not derived from the docs: neither `settings-reference` nor `env-vars` states any credential format (a full verbatim read of both found no `ghp_`, `github_pat_`, `eyJ`, `sk-`, `AKIA`, or `xox` shape), and the token shapes belong to GitHub, AWS, Slack and others. Verified 2026-09-29 against the Claude Code 2.1.284 docs; recheck when either page documents a credential format, then derive the pattern from it and keep the literal as the fallback. Never narrow the literal to match a doc | +| Env vars in settings.json are documented CC vars | warning | Read `code.claude.com/docs/en/env-vars` and search it for each env var name. WebSearch alone is insufficient. **Read it verbatim, not through a summarizer.** The page is long and a summarizing fetch truncates it, then reports the rows past the cutoff as absent: `curl https://code.claude.com/docs/en/env-vars.md` to a file and grep the file, per the [fetch route](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/conventions/upstream-drift/README.md#reading-the-basis-the-fetch-route). A truncated read supports NO finding. Say so and move on. **Absence from this page is not "unrecognized" either:** it is not the whole env-var surface, and `CLAUDE_CODE_ENTRYPOINT`, `CLAUDE_CODE_ENHANCED_TELEMETRY_BETA`, and `CLAUDE_CODE_EXPERIMENTAL_OBSERVER_AGENTS` are all real and all absent from it (our probe, a full verbatim read; no artifact stored. Pointer: [env-vars: variables](https://code.claude.com/docs/en/env-vars#variables). As of: 2026-08-10. Recheck trigger: `env-vars` gains a row for any of them). A name missing here is at most "not documented on `env-vars`". Check `monitoring-usage`, `mcp`, and `settings` before writing anything stronger | +| No secrets in settings.json (tokens, keys, passwords) | error | Scan for patterns: `ghp_`, `eyJ`, `sk-`, `AKIA`, common token prefixes. The engine's `SECRET_RE` is a literal list of vendor token shapes, not derived from the docs: neither `settings-reference` nor `env-vars` states any credential format (our probe: a full verbatim read of both found no `ghp_`, `github_pat_`, `eyJ`, `sk-`, `AKIA`, or `xox` shape), and the token shapes belong to GitHub, AWS, Slack and others. Never narrow the literal to match a doc. Pointer: [settings-reference](https://code.claude.com/docs/en/settings-reference), the whole page, since the probe is a negative over every section, and [env-vars: variables](https://code.claude.com/docs/en/env-vars#variables); no artifact stored. As of: 2026-09-29, Claude Code 2.1.284. Recheck trigger: either page documents a credential format; then derive the pattern from it and keep the literal as the fallback | | Secrets are in settings.local.json only | error | settings.local.json is gitignored | ## G. Skill-listing budget @@ -130,10 +130,10 @@ measurement at all*, and the rest only apply once one exists. | No measurement is reported as "not measured", never as "no overflow" | error | An unmeasured category that reports clean is the failure mode this row exists to block, the same defect as a passing check that never ran | | Any overflow finding names the budget constant it was measured against | warning | `skillListingBudgetFraction` (**default `0.01`**) × context window × ~4 chars/token, or `SLASH_COMMAND_TOOL_CHAR_BUDGET` when set (**documented fallback 8,000 chars**). Confirm both in Phase 3.1 against [settings-reference](https://code.claude.com/docs/en/settings-reference) and [env-vars](https://code.claude.com/docs/en/env-vars), since they are upstream-owned. Without the constant the report cannot say *how far* over | | Roster composition counted before any lever is recommended | warning | Count listing entries by origin: plugin skills, project skills (`.claude/skills/`), user skills (`${CLAUDE_CONFIG_DIR:-~/.claude}/skills/`). This decides which levers exist | -| Every recommended lever is reachable for the origin it targets | error | `skillOverrides` reaches project and user skills only: *"Overrides don't apply to plugin skills, which you manage through `/plugin`."* ([settings-reference](https://code.claude.com/docs/en/settings-reference), verified 2026-09-27; recheck when `doc-citations.tsv` fails on that span). On a plugin-heavy roster the levers are `/plugin` and upstream description trimming. Recommending `skillOverrides` for a plugin skill is a lever the operator cannot pull | -| No `skillOverrides` entry targets a plugin skill (engine: `G/skill-override-plugin`) | warning | For each `skillOverrides` key holding `:` in the user, project and local settings files, the engine takes the text before the first `:` and looks it up among the plugin names in the installed registry and every `enabledPlugins` key. A match is inert: *"Plugin skills are not affected by `skillOverrides`"* ([skills](https://code.claude.com/docs/en/skills)). The reachable levers are `enabledPlugins` or `/plugin` for the whole plugin, or `disable-model-invocation` in the skill's frontmatter, which is the plugin author's to set. A colon key whose prefix names no plugin is a `skip` row, never clean and never a finding: it may be a nested directory-qualified skill (`apps/web:deploy`) or a claude.ai-synced skill (`anthropic-skills:`), and the docs do not settle whether an override reaches it. A key naming no skill anywhere is not detected: bundled and synced skill names are not enumerable from files | -| `skillOverrides` in the user dir's `settings.local.json` is flagged (engine: `G/skill-override-home-local`) | info | The engine reads `/settings.local.json` whatever the project root. That file is the project-local settings file for sessions started in the home directory only ([settings](https://code.claude.com/docs/en/settings) scope table: Project local is `.claude/settings.local.json`, "You, in this one project only"; User is `~/.claude/settings.json`), so its entries reach no other project. A user-wide override belongs in `settings.json`. The `/skills` menu saves to this file from a home-rooted session, so the entry may be intended. Present but unreadable or invalid is `not-inspectable` | -| Per-entry text within the per-skill cap | warning | Combined `description` + `when_to_use` ≤ `skillListingMaxDescChars` (**default `1536`**, verified 2026-08-31; confirm in Phase 3.1 against the [frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference), which is upstream-owned, and a moved default re-derives this row). This is a per-entry cap, independent of the shared budget above | +| Every recommended lever is reachable for the origin it targets | error | We treat `skillOverrides` as reaching project and user skills only, never plugin skills. On a plugin-heavy roster the levers are `/plugin` and upstream description trimming. Recommending `skillOverrides` for a plugin skill is a lever the operator cannot pull. Pointer: [settings-reference: `skillOverrides`](https://code.claude.com/docs/en/settings-reference#skilloverrides), the span pinned in `doc-citations.tsv`. As of: 2026-09-27. Recheck trigger: `check-doc-citations.sh` fails on that span | +| No `skillOverrides` entry targets a plugin skill (engine: `G/skill-override-plugin`) | warning | For each `skillOverrides` key holding `:` in the user, project and local settings files, the engine takes the text before the first `:` and looks it up among the plugin names in the installed registry and every `enabledPlugins` key. A match is inert, since the override does not reach plugin skills (pointer: [skills: override skill visibility from settings](https://code.claude.com/docs/en/skills#override-skill-visibility-from-settings); as of: 2026-09-27; recheck trigger: that section changes whether overrides reach plugin skills). The reachable levers are `enabledPlugins` or `/plugin` for the whole plugin, or `disable-model-invocation` in the skill's frontmatter, which is the plugin author's to set. A colon key whose prefix names no plugin is a `skip` row, never clean and never a finding: it may be a nested directory-qualified skill (`apps/web:deploy`) or a claude.ai-synced skill (`anthropic-skills:`), and the docs do not settle whether an override reaches it. A key naming no skill anywhere is not detected: bundled and synced skill names are not enumerable from files | +| `skillOverrides` in the user dir's `settings.local.json` is flagged (engine: `G/skill-override-home-local`) | info | The engine reads `/settings.local.json` whatever the project root. We treat that file as the project-local settings file for sessions started in the home directory only, so its entries reach no other project (pointer: the scope table in [settings: settings files and who they affect](https://code.claude.com/docs/en/settings#settings-files-and-who-they-affect), its project-local and user rows; as of: 2026-10-01; recheck trigger: that table changes who a project-local or user settings file reaches). A user-wide override belongs in `settings.json`. The entry may be intended: we treat the `/skills` menu as writing to this file from a home-rooted session (pointer: [skills: override skill visibility from settings](https://code.claude.com/docs/en/skills#override-skill-visibility-from-settings); as of: 2026-10-01; recheck trigger: that section changes which file the menu writes). Present but unreadable or invalid is `not-inspectable` | +| Per-entry text within the per-skill cap | warning | Combined `description` + `when_to_use` ≤ `skillListingMaxDescChars`, which this row keeps at **`1536`** as its own setting. Confirm it in Phase 3.1. This is a per-entry cap, independent of the shared budget above. Pointer: for the default, see the [frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference). As of: 2026-08-31. Recheck trigger: the default moves, which re-derives this row | ### Measuring it in a repository (in-repo proxy, not the real population) @@ -158,19 +158,23 @@ accepts into the file and then does not apply the way its author expects. How lo differs per row, so each row says so rather than the section claiming a blanket silence. Two of these rows also have an authoring-time path: the declared schema section A checks for -constrains `effortLevel` by `enum` and `fallbackModel` by `maxItems`, so an editor validating -against it flags them before the file is ever loaded. The rows stay for two reasons: the schema is advisory, +constrains the values of `effortLevel` and `fallbackModel`, so an editor validating against it +flags them before the file is ever loaded. The rows stay for two reasons: the schema is advisory, so the harness still reads a file that violates it, and section A checks that `$schema` is present, not what the values are. Where the schema and the harness disagree, the row says which is -which. **Claim:** the SchemaStore document types `effortLevel` as an `enum` of `low`, `medium`, -`high`, `xhigh` and `fallbackModel` as an array with `maxItems` 3. **Basis:** -`https://json.schemastore.org/claude-code-settings.json`, fetched 2026-09-29. **As of:** -2026-09-29. **Recheck:** either differs on a fresh fetch. +which. We rely on the SchemaStore document constraining both keys; for the constraints, read the +document. + +- **Pointer**: `https://json.schemastore.org/claude-code-settings.json`, its + `properties.effortLevel` and `properties.fallbackModel` entries (a JSON document has no section + anchors, so the entries are named by path). +- **As of**: 2026-09-29 +- **Recheck trigger**: either type differs on a fresh fetch. The engine reads no schema: it never fetches the SchemaStore document, so a limit that only the schema -states (the raw `maxItems` of `fallbackModel`) is a `skip` row, not a finding, and the engine applies no -number that a page it read does not state. Fetching the schema would add a second network dependency -whose version can drift from the installed CLI, and the schema's `enum` and `maxItems` are the +states (a raw array-length limit on `fallbackModel`) is a `skip` row, not a finding, and the engine +applies no number that a page it read does not state. Fetching the schema would add a second network +dependency whose version can drift from the installed CLI, and the schema's constraints are the authoring-time path the paragraph above already assigns to an editor. Apply `jq` recipes to `settings.json` and `~/.claude/settings.json`. For `settings.local.json`, @@ -185,19 +189,20 @@ section F resolves environment variables against their own page. | Check | Severity | How to verify | | --- | --- | --- | | `effortLevel` is a value its **Type** bullet documents | warning | `jq '.effortLevel'`. The accepted values are the **Type** bullet of the [`effortLevel` section](https://code.claude.com/docs/en/settings-reference#effortlevel) on settings-reference; the engine reads that bullet each run and reports any other value. Report it as a value the page does not document; the page does not state what runs instead, so do not assert a level | -| `fallbackModel` keeps no more distinct models than the page says the chain keeps | warning | The cap is the number the [`fallbackModel` section](https://code.claude.com/docs/en/settings-reference#fallbackmodel) states ("at most three distinct allowed models"); the engine reads it each run and applies no other. A section that states none gives one `skip` row. Two tests can disagree: the declared schema caps RAW array length, while the page caps the chain after duplicate removal, so with a cap of 3 `["sonnet","haiku","sonnet","opus"]` has 4 raw entries and exactly 3 distinct. Only the distinct count is a finding; a raw count above the cap is a `skip` row because the schema is not read. Dedupe in place with `jq '.fallbackModel \| reduce .[] as $m ([]; if index($m) then . else . + [$m] end)'`, since `unique` would sort away the order the chain is tried in, and name any entry past the cap as at risk of being ignored. Not "dead": allowlist-excluded entries are also dropped when the chain is read, and the page does not state whether that dropping happens before or after the cap | +| `fallbackModel` keeps no more distinct models than the page says the chain keeps | warning | The cap is the number the [`fallbackModel` section](https://code.claude.com/docs/en/settings-reference#fallbackmodel) states for distinct allowed models; the engine reads it each run and applies no other. A section that states none gives one `skip` row. Two tests can disagree: the declared schema caps RAW array length, while the page caps the chain after duplicate removal, so with a cap of 3 `["sonnet","haiku","sonnet","opus"]` has 4 raw entries and exactly 3 distinct. Only the distinct count is a finding; a raw count above the cap is a `skip` row because the schema is not read. Dedupe in place with `jq '.fallbackModel \| reduce .[] as $m ([]; if index($m) then . else . + [$m] end)'`, since `unique` would sort away the order the chain is tried in, and name any entry past the cap as at risk of being ignored. Not "dead": allowlist-excluded entries are also dropped when the chain is read, and the page does not state whether that dropping happens before or after the cap. Pointer: for the cap and the allowlist drop, see the [`fallbackModel` section](https://code.claude.com/docs/en/settings-reference#fallbackmodel) and [model configuration: fallback model chains](https://code.claude.com/docs/en/model-config#fallback-model-chains); for the raw-length limit, the SchemaStore document's `properties.fallbackModel`. As of: 2026-10-01 (both pages state a cap and neither orders the allowlist drop against it). Recheck trigger: either page drops or changes its cap, or states whether the allowlist drop runs before the cap | | Any other string-valued key whose **Type** bullet lists its values holds one of them | warning | `jq 'to_entries[] \| select(.value \| type == "string")'` over `settings.json` and `~/.claude/settings.json`; `settings.local.json` is not read by value. The engine takes the values from the `string, one of:` bullets on settings-reference, treats an entry such as `custom:` as matching any text in its placeholder, and gives a key with no such list no row. A list holding a bullet that is not a literal value, such as the strftime pattern in the `timeFormat` list, is open: its literals are not the whole set, so the key has no row. `effortLevel` and `disableDeepLinkRegistration` keep their own rows. A key whose Type bullet states its values inline in a sentence is not parsed, so it has no row | -| `availableModels` does not mix a family wildcard with a specific entry of that family | warning | The engine fixes the families as `opus`, `sonnet`, `haiku` and `fable`; the record below the table says why. An entry naming a specific model "disables that family's wildcard entry": `["sonnet", "claude-sonnet-4-5"]` permits only Sonnet 4.5, not every Sonnet. Reached the same way by a Mantle ID and by an `ANTHROPIC_CUSTOM_MODEL_OPTION` value embedding a family name. This one is not silent: an alias narrowed to an older permitted version shows "a notice naming both the requested and substituted models". So the finding is that the allowlist is narrower than its author meant, not that nothing surfaces. Report the models they most likely still expect to be selectable | -| `enforceAvailableModels: true` is paired with a non-empty `availableModels` | error | `jq 'select(.enforceAvailableModels == true) \| .availableModels'`. The finding requires the flag to be `true` AND the list unset or empty. An explicit `false` is someone turning enforcement off on purpose and is never a finding, so gate on the value rather than the key's presence. When it does fire: the key "has no effect when `availableModels` is unset or empty", so an administrator who set it believes the Default option is constrained when it is not. That is an enforcement bypass, which this skill's severity guide rates `error`. Both keys belong in the highest-precedence managed source, and managed sources do not merge; that placement is not decidable from the files this skill reads, so report the pairing, not the placement | - -The engine keeps the family list fixed because no page it reads gives a list to parse. **Claim:** the -Model config aliases table lists `opus`, `sonnet`, `haiku` and `fable` in one column beside `default`, -`best`, `opusplan` and the `[1m]` variants, with no column or note marking which entries name a -family, and the page names the family aliases only inside sentences, such as "a model family alias, -`opus`, `sonnet`, `haiku`, or `fable`" in Restrict model selection. **Basis:** -, fetched 2026-09-30. **As of:** -2026-09-30, Claude Code 2.1.285. **Recheck:** that table gains a marker for family aliases, the page -states the family aliases as a list, or the engine starts reading the page. +| `availableModels` does not mix a family wildcard with a specific entry of that family | warning | The engine fixes the families as `opus`, `sonnet`, `haiku` and `fable`; the record below the table says why. Flag a list that holds a family's wildcard and also an entry naming a specific model of that family; we read the pair as narrower than its author meant. A Mantle ID, or an `ANTHROPIC_CUSTOM_MODEL_OPTION` value embedding a family name, counts as such an entry. This row is not silent at run time (the pointer says how Claude Code tells the user), so the finding is the narrowing itself. For what the list then permits, follow the pointer. Report the models they most likely still expect to be selectable. Pointer: [model configuration: merge behavior](https://code.claude.com/docs/en/model-config#merge-behavior). As of: 2026-10-01. Recheck trigger: that section changes how a specific entry interacts with its family's wildcard | +| `enforceAvailableModels: true` is paired with a non-empty `availableModels` | error | `jq 'select(.enforceAvailableModels == true) \| .availableModels'`. The finding requires the flag to be `true` AND the list unset or empty. An explicit `false` is someone turning enforcement off on purpose and is never a finding, so gate on the value rather than the key's presence. When it does fire: we treat the key as having no effect without a non-empty `availableModels`, so the constraint its author set is not in force. That is an enforcement bypass, which this skill's severity guide rates `error`. Where both keys must live is a managed-settings placement question this skill cannot decide from the files it reads, so report the pairing, not the placement. Pointer: [settings-reference: `enforceAvailableModels`](https://code.claude.com/docs/en/settings-reference#enforceavailablemodels). As of: 2026-10-01. Recheck trigger: that section changes what the key does with an empty or unset list | + +The engine keeps the family list fixed because no page it reads gives a list to parse: our probe of +the model-config page found no table or list that marks which aliases name a family (no artifact +stored). + +- **Pointer**: for the aliases, see + [model configuration: model aliases](https://code.claude.com/docs/en/model-config#model-aliases). +- **As of**: 2026-09-30, Claude Code 2.1.285 +- **Recheck trigger**: that table gains a marker for family aliases, the page states the family + aliases as a list, or the engine starts reading the page. Model IDs in `modelOverrides` are not validated here: unknown keys are ignored rather than rejected, and deciding whether a key is a real Anthropic model ID means resolving it against @@ -205,18 +210,20 @@ rejected, and deciding whether a key is a real Anthropic model ID means resolvin ### `bashOutputMaxChars` -**Claim:** `bashOutputMaxChars` raises how many characters of a successful Bash or PowerShell -command Claude receives inline, clamped from 4000 to 128000. Unset, Claude receives up to 30000 -characters. The key requires Claude Code v2.1.261 or later. Do not raise it by default: inline -output is context cost. `taskOutputMaxChars` was removed in v2.1.277 with the `TaskOutput` tool -and has no effect on current versions. **Basis:** -[settings-reference](https://code.claude.com/docs/en/settings-reference#bashoutputmaxchars) and -the `taskOutputMaxChars` warning on the same page. **As of:** 2026-09-28. **Recheck trigger:** -that page changes the clamp, the default, or restores `taskOutputMaxChars`. +We treat `bashOutputMaxChars` as a context-cost dial: a set value is the operator's choice, and we +never recommend raising it, because inline command output is context the session pays for. We +treat `taskOutputMaxChars` as dead on Claude Code v2.1.277 and later, our setting for this row. + +- **Pointer**: for the key's range, default and minimum version, see + [settings-reference: `bashOutputMaxChars`](https://code.claude.com/docs/en/settings-reference#bashoutputmaxchars), + and the `taskOutputMaxChars` note on the same page. +- **As of**: 2026-09-28 +- **Recheck trigger**: that page changes the range or the default, or restores + `taskOutputMaxChars`. | Check | Severity | How to verify | | --- | --- | --- | -| `bashOutputMaxChars` is unset, or an integer the page's clamp accepts | info | `jq '.bashOutputMaxChars'`. Report a set value as a context-cost choice, not as a defect. Do not recommend raising it. Report `taskOutputMaxChars` as dead on Claude Code v2.1.277 and later | +| `bashOutputMaxChars` is unset, or an integer in the page's range | info | `jq '.bashOutputMaxChars'`. Report a set value as a context-cost choice, not as a defect. Do not recommend raising it. Report `taskOutputMaxChars` as dead on Claude Code v2.1.277 and later | ### `effort:` and `model:` frontmatter on skills and agents @@ -231,51 +238,61 @@ pointing at the other is the shape a hand-off takes when nobody closes it. | Check | Severity | How to verify | | --- | --- | --- | -| A component's `effort:` pin names the model it was calibrated against, or an event that re-opens it | info | Read the frontmatter of every `skills/*/SKILL.md` and `agents/*.md` in scope. The effort scale is calibrated **per model**, so the same level name does not carry the same underlying value across models, and a level measured against one model and carried to the next is a pin nobody re-measured. The property is stated unqualified at , so it holds for every model rather than being a per-model quirk. **Do not flag a pin at the resolved model's own default level.** That pin encodes no measurement that could go stale. Resolve the default from the same page when the audit runs rather than assuming it: as of 2026-09-28 it is `high` everywhere effort is supported except Opus 5.5 and Sonnet 5.5, which default to `medium`, and Opus 4.7, which defaults to `xhigh`. **Recheck trigger:** `high` ceasing to be the general default, or the exception set changing. Report the missing re-derivation, never the level itself. Which level is right is the author's call and this check has no opinion on it | +| A component's `effort:` pin names the model it was calibrated against, or an event that re-opens it | info | Read the frontmatter of every `skills/*/SKILL.md` and `agents/*.md` in scope. We treat a pin as calibrated to the model it was set on, so a pin that names neither that model nor an event that re-opens it is flagged. **Do not flag a pin at the resolved model's own default level.** That pin encodes no measurement that could go stale. The default levels this row applies, our setting when no fetch is possible: `medium` on Opus 5.5 and Sonnet 5.5, `xhigh` on Opus 4.7, `high` on every other model with effort support. Confirm them against the pointer when the audit runs. Report the missing re-derivation, never the level itself; whether the level is high enough is the next row's question, not this one's. Pointer: for where each model's default is set, see [model configuration: adjust effort level](https://code.claude.com/docs/en/model-config#adjust-effort-level); for per-model calibration, see [choose an effort level](https://code.claude.com/docs/en/model-config#choose-an-effort-level). As of: 2026-10-01. Recheck trigger: a model's default on that page differs from this row's setting | +| A code-changing or verifying component's `effort:` is not pinned below `medium` | warning | Read the same frontmatter. Flag an agent or skill whose `effort:` is pinned below `medium` (today that is `low`) when its body has the model change code (edit, implement, refactor, fix) or verify a change (run the build, the tests, a review, or a check that a change works). The floor is the marketplace's [Effort floor](https://github.com/melodic-software/claude-code-plugins/blob/main/docs/plugin-philosophy.md#effort-floor): `medium` or above for that work on every model that supports effort. Classify by what the body has the model do, not by its name or tool grants. **Do not flag:** a component with no `effort:` pin, which inherits the session level and is not a pin; a read-only, chat-like, or mechanical component (lookups, formatting, listing) pinned `low`; or a model the pin runs on that has no effort support, where the key has no effect. Propose `medium` as the floor and leave any higher level to the author. Pointer: the floor's own record in that section, which points at the upstream effort pages. As of: 2026-10-01. Recheck trigger: that section changes the floor level or the kinds of work it covers | | A component's `effort:` and `model:` are consistent with each other | info | A definition setting `model:` without `effort:` inherits the session's level, and the two together are what a spawn actually runs on. Flag only the combination the author is unlikely to have intended: a cheap `model:` tier paired with a top effort level, or the reverse, with no stated reason. Report the mismatch, never a preferred pairing | -**Claim:** a subagent definition's own `effort` overrides the session level rather than yielding to -it, so the frontmatter value is what ships. **Basis:** the [subagents -reference](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields): "Effort level -when this subagent is active. Overrides the session effort level. Default: inherits from session." -The same row admits `max`, which the `effortLevel` settings key does not, so a level valid here is -not evidence it is valid in a settings file. **As of:** 2026-08-08, fetched as raw markdown. -**Recheck trigger:** that field's precedence or its accepted-level list changing on the page. +These rows read a set `effort:` as the level the component runs at and an absent one as the +session level. A level the frontmatter field accepts is not evidence that the `effortLevel` +settings key accepts it; check each against its own pointer. + +- **Pointer**: for the field's precedence and accepted levels, see the + [subagents frontmatter fields](https://code.claude.com/docs/en/sub-agents#supported-frontmatter-fields); + for the settings key, see + [settings-reference: `effortLevel`](https://code.claude.com/docs/en/settings-reference#effortlevel). +- **As of**: 2026-10-01 (both pages read as raw markdown) +- **Recheck trigger**: the field's precedence or either accepted-level list changing. ## I. Deep-link registration -Claude Code "registers the `claude-cli://` handler with your operating system on macOS, Linux, and -Windows when you send your first prompt of an interactive session", not at install, and -"starting `claude` and exiting without sending a prompt doesn't register the handler" -([deep links](https://code.claude.com/docs/en/deep-links)). Registration "writes to user-level -locations only" (`~/Applications/Claude Code URL Handler.app`, a -`claude-code-url-handler.desktop` under `$XDG_DATA_HOME/applications`, -`HKEY_CURRENT_USER\Software\Classes\claude-cli`). Whether the handler is in fact registered on a -given machine is workstation state, not configuration: these rows read the setting that governs -registration, never those locations. +These rows read the setting that governs whether Claude Code registers the `claude-cli://` +handler, never the places the handler is registered: whether it is in fact registered on a given +machine is workstation state, not configuration. For when and where Claude Code registers it, +follow the pointer. + +- **Pointer**: for when registration happens and where it writes, see + [deep links: registration and supported platforms](https://code.claude.com/docs/en/deep-links#registration-and-supported-platforms). +- **As of**: 2026-09-29 +- **Recheck trigger**: that section changes when registration happens or where it writes. Like two of section H's rows, the value check has an authoring-time path: the declared schema -types this key `"type": "string", "enum": ["disable"]`, so an editor validating -against it flags a boolean before the file is ever loaded. The row stays for the same reasons those -do: the schema is advisory, the harness still reads a file that violates it, and section A checks -that `$schema` is present, not what the values are. **Claim:** the SchemaStore document types the -key `"type": "string", "enum": ["disable"]` and words registration as "on startup". **Basis:** -`https://json.schemastore.org/claude-code-settings.json`, fetched 2026-09-29. **As of:** -2026-09-29. **Recheck:** either differs on a fresh fetch. +constrains this key's value, so an editor validating against it flags a boolean before the file is +ever loaded. The row stays for the same reasons those do: the schema is advisory, the harness still +reads a file that violates it, and section A checks that `$schema` is present, not what the values +are. The SchemaStore document and the deep-links page (linked above) disagree on when registration +happens. + +- **Pointer**: `https://json.schemastore.org/claude-code-settings.json`, its + `properties.disableDeepLinkRegistration` entry (a JSON document has no section anchors, so the + entry is named by path), and the deep-links section linked above. +- **As of**: 2026-09-29 +- **Recheck trigger**: the key's type or the registration wording differs on a fresh fetch. **Reach.** Read the key by value from `.claude/settings.json` and `~/.claude/settings.json`. `check-structure.sh` does not report it, so a `settings.local.json` or managed-settings occurrence is outside what this skill's safe-read rule surfaces. Record it as not inspectable rather than reporting the key as absent. The managed gap is not one a file read would close: server-managed delivery, an MDM plist, and Windows registry policy are all managed sources with no file in the path -this skill resolves, and the first non-empty source wins for a key like this one, which is on none of -the page's per-key exception lists. Nothing about the managed layer is decidable here, present or -absent. Both rows below are built only on what the readable scopes show. +this skill resolves, and which of them wins for this key is the managed-settings page's to say +(pointer: [managed settings: precedence within the managed +tier](https://code.claude.com/docs/en/managed-settings#precedence-within-the-managed-tier); as of: +2026-10-01; recheck trigger: that section moves or changes which managed source wins). +Nothing about the managed layer is decidable here, present or absent. Both rows below are built only on what the readable scopes show. | Check | Severity | How to verify | | --- | --- | --- | -| `disableDeepLinkRegistration`, **when present**, is a value its **Type** bullet documents | warning | `jq 'select(has("disableDeepLinkRegistration")) \| .disableDeepLinkRegistration'`. Gate on `has(…)`, never on the value being non-`null`: a bare `jq '.disableDeepLinkRegistration'` returns `null` for an absent key and for an explicit `null` alike, and an absent key is a consumer accepting the default registration on purpose, never a finding. The check fires only on a key that is present and not a documented value. The accepted value is the **Type** bullet of the [`disableDeepLinkRegistration` section](https://code.claude.com/docs/en/settings-reference#disabledeeplinkregistration) on settings-reference; the engine reads that bullet each run. Boolean `true` is the likely author error, the key reading as a flag; the declared schema's `enum` and `type` already flag it too, so a schema-aware editor catches it first. Report that the documented prevention is not invoked, so nothing exempts the machine from the default first-prompt registration above. Do **not** assert what the harness does with an unrecognized value, or that anything surfaces when it reads one, since neither page states either | -| Where enforcement is required, a `disableDeepLinkRegistration` already set to `"disable"` is not left sitting in a scope that cannot enforce it | warning | The deep-links page: "To prevent registration entirely, set `disableDeepLinkRegistration` to `"disable"` in `settings.json`. To enforce this across an organization so users cannot re-enable it, set it in [managed settings](https://code.claude.com/docs/en/server-managed-settings) instead." Only managed settings enforce, so no scope this audit reads by value can satisfy such a requirement. Both halves of the gate are therefore observable: the finding requires a declared enforcement requirement (the consuming repo's own rules, or the run's stated policy context) AND the key present with `"disable"` in `.claude/settings.json` or `~/.claude/settings.json`, a visible attempt lodged in a scope that cannot deliver it. User scope is the settings page's lowest layer, below project and local, so that entry is overridable as well as unenforcing. Report **the visible placement**, never the system: say that this entry does not enforce the requirement and that whether a managed source separately carries the key is outside this audit's reach, since server-managed delivery, MDM plist, and registry policy have no file on the path it resolves. Route the administrator to the one documented check: "Run `/status` to see which managed source is active" ([server-managed settings](https://code.claude.com/docs/en/server-managed-settings)). Deliberately `warning`, not the `error` its `enforceAvailableModels` sibling carries: a bypass is precisely what cannot be proven from here, and managed settings may already enforce this correctly. It becomes `error` only once an administrator confirms no managed source carries the key. Two cases that are **not** findings by design: the key absent from every readable scope (nothing visible to report on), and someone setting it in their own `~/.claude/settings.json` with no enforcement requirement in play, which is the documented single-machine usage. The two rows are sequential, not simultaneous: a key that is present but wrongly valued fails this row's `"disable"` clause and draws only the row above, and correcting the value in that same scope is what brings it into this gate. So say so when both conditions are in view, rather than reporting a placement finding the gate does not yet support | +| `disableDeepLinkRegistration`, **when present**, is a value its **Type** bullet documents | warning | `jq 'select(has("disableDeepLinkRegistration")) \| .disableDeepLinkRegistration'`. Gate on `has(…)`, never on the value being non-`null`: a bare `jq '.disableDeepLinkRegistration'` returns `null` for an absent key and for an explicit `null` alike, and an absent key is a consumer accepting the default registration on purpose, never a finding. The check fires only on a key that is present and not a documented value. The accepted value is the **Type** bullet of the [`disableDeepLinkRegistration` section](https://code.claude.com/docs/en/settings-reference#disabledeeplinkregistration) on settings-reference; the engine reads that bullet each run. Boolean `true` is the likely author error, the key reading as a flag; the declared schema (record above) flags it too, so a schema-aware editor catches it first. Report that the documented prevention is not invoked, so nothing exempts the machine from the default registration the section pointer above covers. Do **not** assert what the harness does with an unrecognized value, or that anything surfaces when it reads one, since neither page states either. Pointer: that settings-reference section and the deep-links section linked above, read whole for this negative. As of: 2026-10-01. Recheck trigger: either page states what Claude Code does with a value other than the documented one | +| Where enforcement is required, a `disableDeepLinkRegistration` already set to `"disable"` is not left sitting in a scope that cannot enforce it | warning | We treat `"disable"` in any `settings.json` as preventing registration on that machine, and only a managed-settings entry as enforcing it across an organization (pointer: [deep links: registration and supported platforms](https://code.claude.com/docs/en/deep-links#registration-and-supported-platforms); as of: 2026-09-29; recheck trigger: that section changes where the key enforces). Only managed settings enforce, so no scope this audit reads by value can satisfy such a requirement. Both halves of the gate are therefore observable: the finding requires a declared enforcement requirement (the consuming repo's own rules, or the run's stated policy context) AND the key present with `"disable"` in `.claude/settings.json` or `~/.claude/settings.json`, a visible attempt lodged in a scope that cannot deliver it. We treat a user-scope entry as overridable by project and local scope as well as unenforcing (pointer: [settings: settings precedence](https://code.claude.com/docs/en/settings#settings-precedence); as of: 2026-10-01; recheck trigger: that section moves user scope above project or local). Report **the visible placement**, never the system: say that this entry does not enforce the requirement and that whether a managed source separately carries the key is outside this audit's reach, since server-managed delivery, MDM plist, and registry policy have no file on the path it resolves. Route the administrator to `/status`, the one documented way to see which managed source is active (pointer: [server-managed settings: settings precedence](https://code.claude.com/docs/en/server-managed-settings#settings-precedence); as of: 2026-10-01; recheck trigger: that section names a different command for seeing the active managed source). Deliberately `warning`, not the `error` its `enforceAvailableModels` sibling carries: a bypass is precisely what cannot be proven from here, and managed settings may already enforce this correctly. It becomes `error` only once an administrator confirms no managed source carries the key. Two cases that are **not** findings by design: the key absent from every readable scope (nothing visible to report on), and someone setting it in their own `~/.claude/settings.json` with no enforcement requirement in play, which is the documented single-machine usage. The two rows are sequential, not simultaneous: a key that is present but wrongly valued fails this row's `"disable"` clause and draws only the row above, and correcting the value in that same scope is what brings it into this gate. So say so when both conditions are in view, rather than reporting a placement finding the gate does not yet support | ## J. Known-issues fix versions diff --git a/plugins/harness-memory/.claude-plugin/plugin.json b/plugins/harness-memory/.claude-plugin/plugin.json index 8900599fc5..34e1751d5c 100644 --- a/plugins/harness-memory/.claude-plugin/plugin.json +++ b/plugins/harness-memory/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "harness-memory", - "version": "1.0.1", + "version": "1.0.2", "description": "Keeps a repo's Claude Code memory layer healthy and under your control, against criteria derived from official Claude Code documentation. The audit skill checks the instruction/memory layer (CLAUDE.md, a root AGENTS.md, CLAUDE.local.md, .claude/rules/, auto-memory) with a deterministic script-backed spine plus judgment-tier checks. The stateless skill inspects, disables, and (confirm-gated) purges Claude-written auto memory across all settings scopes. Boundary: harness-memory owns the health of CLAUDE.md, AGENTS.md, CLAUDE.local.md, .claude/rules/ and auto-memory (structure, size, placement, index integrity); harness-config audit-instructions judges whether instruction text across those files and skills, agents and hooks still fits the current model, and runs no memory-file hygiene checks.", "author": { "name": "Melodic Software", diff --git a/plugins/harness-memory/CHANGELOG.md b/plugins/harness-memory/CHANGELOG.md index 8f5db5b77e..edcb53fa04 100644 --- a/plugins/harness-memory/CHANGELOG.md +++ b/plugins/harness-memory/CHANGELOG.md @@ -3,6 +3,24 @@ All notable changes to the `harness-memory` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [1.0.2] - 2026-10-02 + +### Changed + +- **The `audit` and `stateless` reference files hold our decision plus a pointer per topic.** + `criteria.md` and `official-guidance.md` in both skills no longer restate documentation text. + The `audit` load-model table is labeled this audit's working model, with a pointer record to the + memory page sections it rests on, and the `update` action refreshes the decision and the pointer + together. +- The `${...}` placeholder rule behind the `stateless` purge and `audit` update spokes is stated as + our decision, with a pointer record to where each variable resolves, instead of the page's + wording. +- `stateless` no longer says `disable` takes effect only next session. It says what auto memory + already loaded stays in the current context, with pointers to the settings pages that say which + edits reach a running session. +- The audit's subagent-memory note and the stateless reference's directory layout point at the + docs sections instead of copying their scope table and file tree. + ## [1.0.1] - 2026-10-02 ### Fixed diff --git a/plugins/harness-memory/skills/audit/SKILL.md b/plugins/harness-memory/skills/audit/SKILL.md index 3c21269603..17b831e92c 100644 --- a/plugins/harness-memory/skills/audit/SKILL.md +++ b/plugins/harness-memory/skills/audit/SKILL.md @@ -27,18 +27,30 @@ the `audit` and `audit-automation-gaps` skills in the `harness-config` plugin). ## Scope -| Entity | Location | Loaded | Audited here | +| Entity | Location | Load model this audit uses | Audited here | |--------|----------|--------|-------------| | Project instructions | `CLAUDE.md` | Every session, full | Yes | -| Project instructions in `AGENTS.md` | `AGENTS.md` and `.claude/AGENTS.md` at the root | Every session, full, both files, when no `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` in the root or any directory above it displaces them (the user root's own `~/.claude/CLAUDE.md` does not count); under a `@AGENTS.md` shim it loads as that file's import instead | Yes, as the project instructions (the C-checks) | +| Project instructions in `AGENTS.md` | `AGENTS.md` and `.claude/AGENTS.md` at the root | Audited only where discovery reports Claude Code reads it; `scripts/lib/agents-md.sh` holds the condition and its record | Yes, as the project instructions (the C-checks) | | Local overrides | `CLAUDE.local.md` | Every session, full | Yes | | Rules | `.claude/rules/**/*.md` | Every session (unconditional) or on-demand (path-scoped) | Yes | | **User instructions** | `${CLAUDE_CONFIG_DIR:-~/.claude}/CLAUDE.md` | Every session, full, in **every** project | Yes | | **User rules** | `${CLAUDE_CONFIG_DIR:-~/.claude}/rules/**/*.md` | Same as project rules, in every project | Yes | -| Auto-memory | `~/.claude/projects//memory/` | First 200 lines / 25KB of MEMORY.md | Yes | -| Nested `AGENTS.md` | `**/AGENTS.md` below the root | On a Read in that directory, unless a `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` on its path is read instead; then only through one that imports or symlinks it | Reachability only (N1); content is not audited | +| Auto-memory | `~/.claude/projects//memory/` | The M1 budget: first 200 lines / 25KB of MEMORY.md | Yes | +| Nested `AGENTS.md` | `**/AGENTS.md` below the root | Reachable unless a `CLAUDE.md`-family file on its path displaces it without importing or symlinking it (N1) | Reachability only (N1); content is not audited | | Settings, hooks, MCP, agents, skills | Various | Various | No. Use `harness-config`'s `audit` / `audit-automation-gaps` | +The load-model column is this audit's working model, not a restatement of the docs; each check in +[reference/criteria.md](reference/criteria.md) carries the pointer it rests on. + +- **Pointer**: for how each file loads, see + [How CLAUDE.md files load](https://code.claude.com/docs/en/memory#how-claude-md-files-load), + [Organize rules with `.claude/rules/`](https://code.claude.com/docs/en/memory#organize-rules-with-claude/rules/), + [When Claude Code reads AGENTS.md](https://code.claude.com/docs/en/memory#when-claude-code-reads-agents-md) + and auto memory's [How it works](https://code.claude.com/docs/en/memory#how-it-works). +- **As of**: 2026-10-01 +- **Recheck trigger**: a Claude Code release note or a change to one of those sections alters which + files load at session start, on demand, or within the auto-memory limits. + Auto memory's effective enabled/disabled state must be resolved before auditing it, not assumed from a single scope: [`${CLAUDE_PLUGIN_ROOT}/skills/stateless/context/status.md`](../stateless/context/status.md), "Resolve the effective state". @@ -85,18 +97,23 @@ plus the provenance classification each finding carries) yields byte-identical f repo state; its **judgment tier** (C2-C9, R1-R4, M3-M4) applies fixed criteria with model reading, so findings vary in wording though not in criteria. Label those "judgment candidate" in the report. Criteria derive from official Claude -Code documentation (sourced quotes in [reference/official-guidance.md](reference/official-guidance.md)); -refresh both via the `update` action. +Code documentation: [reference/official-guidance.md](reference/official-guidance.md) holds, per +topic, the audit's decision and a pointer to the docs section behind it, never the docs' text. +Refresh both via the `update` action. ## Script paths The `context/` and `reference/` files write each bundled script as `/scripts/.sh`, where `` is this skill's directory: `${CLAUDE_SKILL_DIR}`. Put that path in place of the -placeholder before running a command. Those files arrive through the Read tool as plain bytes, so a -`${…}` token in them would reach the Bash tool unsubstituted, and the Bash tool's environment has no -`CLAUDE_PLUGIN_ROOT` to expand it from. Basis: the plugins reference, "Where each variable resolves", -and the skills page, "Available string substitutions", both verified 2026-09-27; recheck when either -table adds supporting files to where a `${…}` reference resolves. +placeholder before running a command. We never put a `${…}` token in those files: we do not rely on +one being substituted in a file read through the Read tool, or on the Bash tool's environment +carrying `CLAUDE_PLUGIN_ROOT`. + +- **Pointer**: for where each `${…}` reference resolves, see + [Where each variable resolves](https://code.claude.com/docs/en/plugins-reference#where-each-variable-resolves) + and [Available string substitutions](https://code.claude.com/docs/en/skills#available-string-substitutions). +- **As of**: 2026-09-27 +- **Recheck trigger**: either table adds supporting files to where a `${…}` reference resolves. ## Audit mode (default) @@ -129,12 +146,19 @@ It prints `/`, the scheme `harness-config and `audit-prompting-postures` uses. Run it and use its output as the key. Pass `--explain` when the report should say which rung produced its key. -**Why the key exists.** `${CLAUDE_PLUGIN_DATA}` resolves to `~/.claude/plugins/data/{id}/`, keyed to -the plugin identifier and nothing else. No project, checkout, worktree, or session segment -([plugins reference](https://code.claude.com/docs/en/plugins-reference), § Persistent data directory). -A fixed `audit/last-audit.md` is therefore **one file per machine**. Losing reports is the smaller -half; the larger half is the read. `report` mode would serve whatever that file currently holds and -`fix` mode would act on it, so on a machine with two repositories, project B can be shown project A's +**Why the key exists.** We treat `${CLAUDE_PLUGIN_DATA}` as one directory per plugin per machine, +with no project, checkout, worktree, or session segment. A fixed `audit/last-audit.md` is therefore +**one file per machine**. + +- **Pointer**: for the plugin data directory, see + [Environment variables](https://code.claude.com/docs/en/plugins-reference#environment-variables). +- **As of**: 2026-10-01 +- **Recheck trigger**: that section adds a project, worktree, or session segment to the data + directory's path. + +Losing reports is the smaller half; the larger half is the read. `report` mode would serve whatever +that file currently holds and `fix` mode would act on it, so on a machine with two repositories, +project B can be shown project A's findings and offered edits derived from another repository's memory layer. That is a wrong answer served, not merely an artifact lost, which is why an append-only history does not close it and the *path* has to carry project identity. diff --git a/plugins/harness-memory/skills/audit/context/update.md b/plugins/harness-memory/skills/audit/context/update.md index 7e32f3e6bd..bd95434657 100644 --- a/plugins/harness-memory/skills/audit/context/update.md +++ b/plugins/harness-memory/skills/audit/context/update.md @@ -30,9 +30,11 @@ Research must cover: ## Step 2: Diff against current guidance Read [../reference/official-guidance.md](../reference/official-guidance.md) and compare against -research findings: +research findings. Each section there holds the audit's decision and a pointer to a docs section, +never the docs' text, so compare each decision against the section its pointer names: -1. Identify changed guidance (quotes no longer matching) +1. Identify changed guidance (a decision the pointed-at section no longer supports, or an + anchor that moved) 2. Identify new guidance (topics not covered) 3. Identify removed/deprecated guidance @@ -47,7 +49,8 @@ diff as a contribution/issue against the plugin's repository so the shipped crit With that framing, the content updates are: -1. `reference/official-guidance.md`: new/changed quotes, dates, source URLs +1. `reference/official-guidance.md`: new or changed decisions in our words, pointers, as-of dates + and recheck triggers. Never copy the page's text into the file, quoted or paraphrased. 2. `reference/criteria.md`: check thresholds or severity levels needing adjustment, version number, "Last updated" date @@ -66,6 +69,6 @@ Present findings as actionable suggestions, not automatic changes. Output a summary of what changed: -- Guidance quotes: N updated, N added, N removed +- Guidance records: N updated, N added, N removed - Ecosystem suggestions: N items - Next action: suggest re-running the audit with updated criteria diff --git a/plugins/harness-memory/skills/audit/reference/criteria.md b/plugins/harness-memory/skills/audit/reference/criteria.md index 97093d2d02..2182f3779f 100644 --- a/plugins/harness-memory/skills/audit/reference/criteria.md +++ b/plugins/harness-memory/skills/audit/reference/criteria.md @@ -18,6 +18,11 @@ command. This file defines every check the audit runs. Each check has a severity, description, and instructions for evaluation. The audit applies checks per-entity-type (CLAUDE.md, rules, memory). +Every firing rule here is our decision, in our words; a doc-derived check names the docs section it +rests on as a **Pointer** and stores none of the page's text. Unless a check says otherwise, each +pointer's as-of date is 2026-09-20 (the "Last updated" date above) and its recheck trigger is a +Claude Code release note or docs change touching the section it points at. + To refresh this file against current official guidance, run the skill's `update` action. --- @@ -25,14 +30,18 @@ To refresh this file against current official guidance, run the skill's `update` ## Checks for CLAUDE.md and CLAUDE.local.md Every C-check here applies equally to each project root `AGENTS.md` that discovery emits as an -`agents-md` surface (`AGENTS.md` and `.claude/AGENTS.md`, both of which load at session start and -between which the doc states no precedence): that file IS the project instructions for the -session, so the same budget, content and currency criteria govern it. Discovery emits it only where Claude Code reads it, which +`agents-md` surface (`AGENTS.md` and `.claude/AGENTS.md`; we audit both and assume no precedence +between them): that file IS the project instructions for the session, so the same budget, content +and currency criteria govern it. Discovery emits it only where Claude Code reads it, which is why a repo under a one-line `@AGENTS.md` shim has no such row (the import already counts inside the CLAUDE.md's expanded figure) and a displaced `AGENTS.md` has none either. Cite the finding -against `AGENTS.md`, not against a CLAUDE.md that is not there. Basis: -code.claude.com/docs/en/memory, "When Claude Code reads AGENTS.md", fetched 2026-09-20; the -condition and its recheck trigger are recorded in `scripts/lib/agents-md.sh`. +against `AGENTS.md`, not against a CLAUDE.md that is not there. + +- **Pointer**: for when Claude Code reads `AGENTS.md`, see + [When Claude Code reads AGENTS.md](https://code.claude.com/docs/en/memory#when-claude-code-reads-agents-md); + the condition discovery applies is recorded in `scripts/lib/agents-md.sh`. +- **As of**: 2026-09-20 +- **Recheck trigger**: the recheck trigger recorded in `scripts/lib/agents-md.sh` fires. ### C1: Line Budget [FAIL] @@ -53,18 +62,22 @@ condition and its recheck trigger are recorded in `scripts/lib/agents-md.sh`. 5. Report the expanded figure and, when imports contributed, the per-file breakdown: a one-line `CLAUDE.md` importing a 300-line `AGENTS.md` is a 301-line file for this check -**Why**: Official docs: "Target under 200 lines per CLAUDE.md file. Longer files consume more context -and reduce adherence." Files over 200 lines cause Claude to ignore instructions. Imports count -because "imported files still load and enter the context window at launch" and "splitting into -`@path` imports helps organization but doesn't reduce context" (code.claude.com/docs/en/memory); -a raw line count of the root file alone passes a layer the loader treats as one file. +**Why**: We take the docs' per-file size target as this check's 200-line budget, and read a file +over it as an adherence risk. We count `@` imports because we treat imported files as loading at +launch, so a split saves nothing; a raw line count of the root file alone passes a layer the loader +treats as one file. + +- **Pointer**: for the size target, see + [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions); + for imports, see [Import additional files](https://code.claude.com/docs/en/memory#import-additional-files). + +**Diagnostic**: We read a rule Claude keeps ignoring as a sign the file is too long. When the audit +was prompted by a rule being ignored, add a C1 WARN naming that symptom even when steps 4-6 pass. +The branch is prompt-conditioned, so it belongs to the judgment tier. Label it "judgment +candidate" in the report; steps 1-6 remain the deterministic spine, unaffected. -**Diagnostic**: The symptom-first tell for this check: "If Claude keeps doing something you don't -want despite having a rule against it, the file is probably too long and the rule is getting lost" -(code.claude.com/docs/en/best-practices). When the audit was prompted by a rule being ignored, add a -C1 WARN citing this tell even when steps 4-6 pass. The branch is prompt-conditioned, so it belongs -to the judgment tier. Label it "judgment candidate" in the report; steps 1-6 remain the -deterministic spine, unaffected. +- **Pointer**: for the ignored-rule symptom, see + [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md). **Allowances**: Complex monorepos using `.claude/rules/` extensively may justify overages, and a repo may document a deliberate exemption in its own rules (see SKILL.md "Consumer-convention extension @@ -72,7 +85,7 @@ seam"). Report overage and justification together. ### C2: Deletion Test [WARN per line] -**What**: For each line, ask: "Would removing this cause Claude to make mistakes?" If not, cut it. +**What**: Flag each line whose removal would not lead Claude into a mistake. **How to check**: @@ -90,8 +103,10 @@ seam"). Report overage and justification together. skill"); group findings by H1/H2 section, and collapse a section whose every line flags into one section-level finding -**Why**: Official docs: "For each line, ask: 'Would removing this cause Claude to make mistakes?' If -not, cut it. Bloated CLAUDE.md files cause Claude to ignore your actual instructions!" +**Why**: This is the docs' per-line pruning test, applied line by line; we treat surplus lines as +diluting the ones that matter. + +- **Pointer**: [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md). ### C3: Content Placement [WARN] @@ -104,10 +119,10 @@ not, cut it. Bloated CLAUDE.md files cause Claude to ignore your actual instruct | Always-on project conventions | CLAUDE.md | | Machine-specific config/preferences | CLAUDE.local.md | | Language/framework-specific rules | `.claude/rules/` (path-scoped when that fits) | -| Subdirectory-specific conventions | Nested `CLAUDE.md` in that subdirectory, which loads on demand when Claude reads files there (ancestors of cwd load in full at launch); post-compaction re-injection priced below (code.claude.com/docs/en/memory) | -| One-off steering for the current conversation | A conversational `@`-mention of the file, which includes the file's full content in the conversation (code.claude.com/docs/en/common-workflows, "Reference files and directories"); distinct from `@path` imports *in* CLAUDE.md, which load at launch every session (priced in the imports row below). **Provenance**: the one-conversation scope and the cheaper-than-any-permanent-pointer framing are inferred, not doc-stated, and the placement posture is a repo extension (that doc is outside this file's `Source:` set), so the `update` action must not overwrite this row | +| Subdirectory-specific conventions | Nested `CLAUDE.md` in that subdirectory, which we treat as on-demand for that subdirectory; post-compaction re-injection priced below. Pointer: [Choose where to put CLAUDE.md files](https://code.claude.com/docs/en/memory#choose-where-to-put-claude-md-files) | +| One-off steering for the current conversation | A conversational `@`-mention of the file (pointer: [Reference files and directories](https://code.claude.com/docs/en/common-workflows#reference-files-and-directories)); distinct from `@path` imports *in* CLAUDE.md, which load at launch every session (priced in the imports row below). **Provenance**: the one-conversation scope and the cheaper-than-any-permanent-pointer framing are inferred, not doc-stated, and the placement posture is a repo extension (that doc is outside this file's `Source:` set), so the `update` action must not overwrite this row | | Reference material needed sometimes | Skills: the body loads on demand; a new skill's listing entry does not (priced below) | -| Learnings Claude discovered while working, not instructions you authored | Auto memory: Claude writes it; you do not hand-author entries, and asking Claude to remember something lands here rather than in CLAUDE.md. Available only while auto memory is enabled (gated below) | +| Learnings Claude discovered while working, not instructions you authored | Auto memory, which Claude writes rather than you; a request to remember something also belongs here rather than in CLAUDE.md. Available only while auto memory is enabled (gated below) | | Deterministic enforcement | Hooks (guaranteed execution) | | Compile-time/build-time rules | Analyzers, linters, architecture tests | | Information that changes frequently | Neither: keep it out | @@ -115,11 +130,13 @@ not, cut it. Bloated CLAUDE.md files cause Claude to ignore your actual instruct Flag content in the wrong layer. WARN severity because moving content is a judgment call. -**Auto memory is a destination only while it is enabled. Resolve that before routing to it.** It is -on by default, but `autoMemoryEnabled` and `CLAUDE_CODE_DISABLE_AUTO_MEMORY` can turn it off, and -Claude then neither writes nor loads auto-memory files -(). Recommending that accumulated learnings leave `CLAUDE.md` -for auto memory in that state deletes them from every future session instead of relocating them. +**Auto memory is a destination only while it is enabled. Resolve that before routing to it.** We +treat it as on by default and as switched off by `autoMemoryEnabled` or +`CLAUDE_CODE_DISABLE_AUTO_MEMORY`, with nothing written or loaded while off. Recommending that +accumulated learnings leave `CLAUDE.md` for auto memory in that state deletes them from every +future session instead of relocating them. + +- **Pointer**: [Enable or disable auto memory](https://code.claude.com/docs/en/memory#enable-or-disable-auto-memory). Rather than reading a single scope, resolve the **effective** state with the algorithm the sibling `stateless` skill already owns in @@ -139,37 +156,40 @@ nothing. Reproduced first-party on Claude Code 2.1.219 (2026-07-24); no official with doc-sourced text, and it needs re-verification on a current version rather than a doc re-fetch. **Price the move with the recommendation.** Moving content out of an always-loaded surface trades -per-session cost for post-compaction absence, and the trade differs by destination: path-scoped rules -and nested CLAUDE.md are re-injected only when a matching file is read again, while root CLAUDE.md, -unscoped rules, and auto memory are re-injected from disk. Read the destination's row in -[official-guidance.md](official-guidance.md), "Compaction by steering method", before recommending a -move, and state the cost alongside it. A rule that must persist across compaction stays unscoped or +per-session cost for post-compaction absence, and the trade differs by destination. Read the +destination's row in [official-guidance.md](official-guidance.md), "Compaction by steering method", +which holds the audit's per-destination model and its pointer, before recommending a move, and +state the cost alongside it. A rule that must persist across compaction stays unscoped or in the project-root CLAUDE.md. A recommendation that omits this proposes a silent behavior change in long sessions. -A **new** skill carries a second cost the compaction table does not show: the body defers, but the -listing entry it adds is always in context: `name` plus the combined `description` and -`when_to_use`, truncated at 1,536 characters. The saving is the body minus that entry rather than +A **new** skill carries a second cost the compaction table does not show: the body defers, but we +price the listing entry it adds as always in context: `name` plus the combined `description` and +`when_to_use`, at up to 1,536 characters. The saving is the body minus that entry rather than the whole body. Moving content into a skill that **already exists** adds no listing entry and does not carry -this cost. The only field that keeps a description out of context is `disable-model-invocation: true`, -which also makes the skill user-invocable only; `user-invocable: false` does not, and `skillOverrides` -does not reach plugin skills at all. State the entry as a cost of the recommended move. Whether the -target's listing budget is oversubscribed is a separate question this check does not answer. - -**Why**: Official docs: "For domain knowledge or workflows that are only relevant sometimes, use -skills instead. Claude loads them on demand without bloating every conversation." And: "Unlike -CLAUDE.md instructions which are advisory, hooks are deterministic." On imports: "splitting into -`@path` imports helps organization but doesn't reduce context, since imported files load at launch" -(code.claude.com/docs/en/memory). The per-destination compaction behavior is quoted with its sources -in [official-guidance.md](official-guidance.md) rather than restated here. On the listing entry: -"skill descriptions are loaded into context so Claude knows what's available, but full skill content -only loads when invoked", the combined `description` and `when_to_use` text "is truncated at 1,536 -characters in the skill listing to reduce context usage", and "Plugin skills are not affected by -`skillOverrides`" (quoted from - and that page's -invocation-control and visibility-override sections; verified 2026-08-31; recheck trigger: a -fetch of that page no longer carrying these quoted spans re-derives this paragraph and the -listing-entry cost above). +this cost. Price the entry as zero only for a skill with `disable-model-invocation: true` (which +also leaves it user-invocable only); never on the strength of `user-invocable: false`, or of a +`skillOverrides` entry against a plugin skill. State the entry as a cost of the recommended move. +Whether the target's listing budget is oversubscribed is a separate question this check does not +answer. + +**Why**: We place sometimes-needed domain knowledge in skills, deterministic enforcement in hooks, +and never count an `@path` split as a saving. The per-destination compaction model and its pointer +live in [official-guidance.md](official-guidance.md), not here. + +- **Pointer**: for skills and hooks as alternatives to CLAUDE.md, see + [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md) + and [Set up hooks](https://code.claude.com/docs/en/best-practices#set-up-hooks); for imports, see + [Import additional files](https://code.claude.com/docs/en/memory#import-additional-files). +- **Pointer** (the listing-entry cost and the zero-cost rule): for the listing cap, invocation + control and visibility overrides, see + [Frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference), + [Control who invokes a skill](https://code.claude.com/docs/en/skills#control-who-invokes-a-skill) + and [Override skill visibility from settings](https://code.claude.com/docs/en/skills#override-skill-visibility-from-settings). +- **As of** (the listing-entry pointer): 2026-08-31 +- **Recheck trigger** (the listing-entry pointer): a fetch of those sections no longer supports the + 1,536-character cap, the zero-cost rule, or the plugin-skill exclusion above; re-derive this + paragraph and the listing-entry cost from them. ### C4: Specificity [WARN] @@ -177,13 +197,16 @@ listing-entry cost above). **How to check**: -1. Scan for vague instructions: "format code properly", "keep things organized", "follow best practices", "write clean code" +1. Scan for vague instructions, such as "tidy up the formatting", "keep things organized", + "follow best practices", "write clean code" 2. Scan for instructions without actionable verbs or concrete outcomes 3. WARN for each vague instruction 4. Include a suggested rewrite -**Why**: Official docs examples: "Use 2-space indentation" instead of "Format code properly". "Run -`npm test` before committing" instead of "Test your changes." +**Why**: We want each instruction concrete enough to check: a named value or command, such as +"indent with 4 spaces" or "run `make check` before pushing", rather than a quality adjective. + +- **Pointer**: [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions). ### C5: Non-obvious Only [WARN] @@ -202,14 +225,15 @@ listing-entry cost above). 3. Flag framework documentation that should be linked, not copied 4. WARN per instance -**Provenance**: the KEEP branch is a **repo extension, not doc-derived**. The official -include/exclude table states no navigation posture (checked 2026-08-17 against -code.claude.com/docs/en/memory), so the `update` action must not overwrite it with doc-sourced -text. +**Provenance**: the KEEP branch is a **repo extension, not doc-derived**: we found no navigation +posture in the docs' include/exclude table (checked 2026-08-17), so the `update` action must not +overwrite it. -**Why**: Official include/exclude table: Exclude "Anything Claude can figure out by reading code", -"Standard language conventions Claude already knows", "Detailed API documentation (link to docs -instead)." +**Why**: Steps 1-3 apply the exclude side of the docs' include/exclude table, in our words. + +- **Pointer**: for the include/exclude table, see + [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md). +- **As of**: 2026-08-17 ### C6: Consistency [FAIL] @@ -225,7 +249,7 @@ when both sides are in the `discover-instruction-surfaces` population? 3. Check for redundancy (same instruction in multiple files) 4. Compare **user**-scope surfaces against project ones. Both load together, so a user↔project contradiction is a live conflict (see Step 3 in `context/audit.md`) -5. FAIL for contradictions (Claude picks one arbitrarily) +5. FAIL for contradictions (which side Claude follows is undefined) 6. WARN for redundancy (wastes context budget) **Boundary**: This check owns instruction-content conflicts whose **both** anchors are in the @@ -233,7 +257,10 @@ discover-instruction-surfaces population. Nested `CLAUDE.md` files, auto-memory, skills, agents, and output styles are outside that population, so those pairs belong to `harness-config:audit-instructions` I15 (and its precedence / co-residency adjudication), not here. -**Why**: Official docs: "If two rules contradict each other, Claude may pick one arbitrarily." +**Why**: A contradiction leaves which instruction Claude follows undefined, so we fail it rather +than warn. + +- **Pointer**: [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions). ### C7: Currency [FAIL] @@ -252,8 +279,8 @@ skills, agents, and output styles are outside that population, so those pairs be 6. WARN for stale counts **Navigation-section note**: stale pointers are the standing cost of the curated navigation -sections C5's KEEP branch permits, "a stale highway is worse than no highway": a pointer that -outlives its target misroutes every future session. This check's missing-file FAIL is what keeps +sections C5's KEEP branch permits, and a stale one is worse than none: a pointer that outlives its +target misroutes every future session. This check's missing-file FAIL is what keeps that posture honest, so give C5-kept navigation entries particular attention here. **Why**: Stale references cause Claude to hallucinate or waste time looking for nonexistent files. @@ -307,17 +334,19 @@ which are not repo-scoped. A wrong build command is not a C7 finding today, because a command is none of the three things C7 checks. Report a wrong command under C9 only, and do not double-report it. -**Why**: Official docs list "build and test commands" first among what project memory is for -(code.claude.com/docs/en/memory), and `/init` populates them by analyzing the codebase, so without -the statement, they are inferred every session rather than read. This check fires on a CLAUDE.md -that exists but omits them. Absent commands make every verification loop start by guessing how to -run the check. +**Why**: We treat the repo's build and test commands as core project-instruction content: without +a statement of them, every session infers them rather than reads them. This check fires on a +CLAUDE.md that exists but omits them. Absent commands make every verification loop start by +guessing how to run the check. + +- **Pointer**: [Set up a project CLAUDE.md](https://code.claude.com/docs/en/memory#set-up-a-project-claude-md). -**Counter-evidence, and why step 0 exists**: the same page's CLAUDE.md-vs-auto-memory table puts -"Build commands" in the *auto memory* column's "Use for" cell, against CLAUDE.md's "Coding -standards, workflows, project architecture". The page states both, so the honest reading is that -the commands must be *reachable*, not that they must sit in CLAUDE.md specifically. Step 0 is what -keeps this check from flagging a repo that followed the other half of the same page. +**Counter-evidence, and why step 0 exists**: the same page's CLAUDE.md-vs-auto-memory comparison +also bears on where build commands belong (pointer: +[CLAUDE.md vs auto memory](https://code.claude.com/docs/en/memory#claude-md-vs-auto-memory)). We +therefore require the commands to be *reachable* on some loaded surface, not to sit in CLAUDE.md +specifically. Step 0 is what keeps this check from flagging a repo that put them on another loaded +surface. --- @@ -407,20 +436,20 @@ instruction files are never reported as a missing Claude shim), and for each ask stays depth-1: this check is about the pointer, not the nested file's content. FAIL per displaced, unimported file; the fix is a one-line `@AGENTS.md` `CLAUDE.md` beside it. -**Why**: Official docs: Claude reads `AGENTS.md` "only when you have no `CLAUDE.md` in your working -directory or above it", counting "a `CLAUDE.md`, `.claude/CLAUDE.md`, or `CLAUDE.local.md` in your -working directory or any directory above it"; it attaches "a subdirectory's `AGENTS.md`, when Claude -opens a file there with the Read tool and that subdirectory has none of the three `CLAUDE.md` files -of its own"; and where reading `AGENTS.md` directly is unavailable, the page says to "import it from -a `CLAUDE.md`" (code.claude.com/docs/en/memory, "AGENTS.md", "When Claude Code reads AGENTS.md", -"When AGENTS.md support is unavailable"; verified 2026-09-19; recheck trigger: a fetch of that page -no longer stating which file names count for that check). - -The earlier basis for this check was the same page's sentence "Claude Code reads `CLAUDE.md`, not -`AGENTS.md`. If your repository already uses `AGENTS.md` for other coding agents, create a -`CLAUDE.md` that imports it", verified 2026-09-08. That recheck trigger fired: the sentence is gone -from the page as fetched 2026-09-19, and direct `AGENTS.md` reading shipped in v2.1.277. It is -quoted here unchanged as the superseded basis, never as a current claim. +**Why**: The script models a nested `AGENTS.md` as read directly only when no `CLAUDE.md`, +`.claude/CLAUDE.md` or `CLAUDE.local.md` on its path displaces it, and otherwise as reachable only +through an import or symlink from one of those files. An import from a `CLAUDE.md` is also the +fix where direct reading is unavailable, so the fix line names it in every case. + +- **Pointer**: for when an `AGENTS.md` is read and the fallback where support is unavailable, see + [When Claude Code reads AGENTS.md](https://code.claude.com/docs/en/memory#when-claude-code-reads-agents-md) + and [When AGENTS.md support is unavailable](https://code.claude.com/docs/en/memory#when-agents-md-support-is-unavailable). +- **As of**: 2026-09-19 +- **Recheck trigger**: a fetch of those sections changes which file names displace an `AGENTS.md`. + +The earlier basis for this check (verified 2026-09-08) was the page's former CLAUDE.md-only +reading model. Its trigger fired: the page had changed when fetched 2026-09-19, and direct +`AGENTS.md` reading shipped in v2.1.277. That basis is superseded and is no current claim. --- @@ -446,11 +475,14 @@ of a local edit; `local` keeps the ordinary fix line. RD1 does this itself; the **What**: Is MEMORY.md under 200 lines / 25KB? **How to check**: Count lines and file size on the content that loads. Strip YAML frontmatter and -block-level HTML comments first, since they are removed before the index is loaded and don't count -toward the limits. The SKILL.md pre-computed context already reports both post-strip figures +block-level HTML comments first: we count only loaded content toward the limits. The SKILL.md +pre-computed context already reports both post-strip figures (`memory-dir-stats.sh --memory-lines` / `--memory-bytes`); use them rather than re-measuring the raw -file. Only the first 200 loaded lines (or 25KB) load at session start, and anything beyond is -silently dropped. +file. This check sets the limit at the first 200 loaded lines or 25KB and treats anything beyond +as not loaded at session start. + +- **Pointer**: for the auto-memory load limits, see + [How it works](https://code.claude.com/docs/en/memory#how-it-works). Four readings the strip applies, so a hand count matches the reported figures: @@ -474,11 +506,10 @@ Four readings the strip applies, so a hand count matches the reported figures: 4. Byte counts measure LF-normalized content, so a CRLF index reports about one byte per line under its on-disk size, well under 1% of the 25KB cap. -**Provenance**: the strip rule itself is doc-derived (code.claude.com/docs/en/memory, "How it -works"). The four readings are not. The doc states the fenced-code carve-out for CLAUDE.md only and -is silent on it for MEMORY.md, and says nothing about unterminated blocks, unbounded blocks, -partial lines, or line endings. They are this plugin's reading, chosen so that no input silently -under-reports and leaves this `[FAIL]` gate unable to fire. Where a reading has to guess, it guesses +**Provenance**: the strip rule itself is doc-derived (pointer above). The four readings are not: +they cover cases we did not find the docs to settle for MEMORY.md (fenced code, unterminated or +unbounded blocks, partial lines, line endings). They are this plugin's reading, chosen so that no +input silently under-reports and leaves this `[FAIL]` gate unable to fire. Where a reading has to guess, it guesses toward counting: an over-count can only make the gate fire early on a file near its limit, while an under-count stops it firing at all. The `update` action must not overwrite them. diff --git a/plugins/harness-memory/skills/audit/reference/official-guidance.md b/plugins/harness-memory/skills/audit/reference/official-guidance.md index 8a73c7716d..4b8f85d7f6 100644 --- a/plugins/harness-memory/skills/audit/reference/official-guidance.md +++ b/plugins/harness-memory/skills/audit/reference/official-guidance.md @@ -24,9 +24,13 @@ - [Compaction by steering method (June 2026)](#compaction-by-steering-method-june-2026) - [No official scoring rubric](#no-official-scoring-rubric) -Last researched: 2026-06-20; code.claude.com/docs/en/memory re-verified 2026-08-10 (the other -sources below were not re-checked on that date) -Sources: [Steering Claude Code (June 18, 2026)](https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more), code.claude.com/docs/en/memory, code.claude.com/docs/en/hooks, code.claude.com/docs/en/best-practices, code.claude.com/docs/en/sub-agents, howborisusesclaudecode.com +Each section states what this audit does, in our words, and points at the section of the official +page that covers the topic. Read the page there for its wording; this file stores none of it. + +Last researched: 2026-06-20; the memory page pointers re-verified 2026-08-10 (the other sources +below were not re-checked on that date). Unless a section says otherwise, each pointer's as-of date +is 2026-06-20 and its recheck trigger is a Claude Code release note or docs change touching the +section it points at. Refresh this file from current official docs via the skill's `update` action. @@ -34,100 +38,74 @@ Refresh this file from current official docs via the skill's `update` action. ## Size and adherence -> "**Size**: target under 200 lines per CLAUDE.md file. Longer files consume more context and reduce adherence." -> -> code.claude.com/docs/en/memory - - - -> "Files over 200 lines consume more context and may reduce adherence." -> -> code.claude.com/docs/en/memory (troubleshooting section) - - - -> "Bloated CLAUDE.md files cause Claude to ignore your actual instructions!" -> -> code.claude.com/docs/en/best-practices - - - -> "If Claude keeps doing something you don't want despite having a rule against it, the file is probably too long and the rule is getting lost." -> -> code.claude.com/docs/en/best-practices - - +We flag a CLAUDE.md over 200 lines as an adherence risk, and treat a rule the model keeps ignoring +as a sign the file is too long. A third-party guide sets its own, looser line count; we use the +official target. -> "Less than 300 lines is best, and shorter is even better." -> -> humanlayer.dev/blog/writing-a-good-claude-md +- **Pointer**: [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions), + [My CLAUDE.md is too large](https://code.claude.com/docs/en/memory#my-claude-md-is-too-large), + [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md); + third-party: [humanlayer.dev, writing a good CLAUDE.md](https://humanlayer.dev/blog/writing-a-good-claude-md). ## Context injection clarification -> "CLAUDE.md content is delivered as a user message after the system prompt, not as part of the system prompt itself. Claude reads it and tries to follow it, but there's no guarantee of strict compliance, especially for vague or conflicting instructions." -> -> code.claude.com/docs/en/memory (troubleshoot section) +We treat CLAUDE.md content as context the model reads, not enforced configuration and not part of +the system prompt, so a finding never promises strict compliance. Where an instruction must sit at +the system-prompt level, we point to `--append-system-prompt`. - - -> "For instructions you want at the system prompt level, use `--append-system-prompt`." -> -> code.claude.com/docs/en/memory +- **Pointer**: [Troubleshoot memory issues](https://code.claude.com/docs/en/memory#troubleshoot-memory-issues), + the "Claude isn't following my CLAUDE.md" entry. ## The deletion test -> "Keep it concise. For each line, ask: 'Would removing this cause Claude to make mistakes?' If not, cut it." -> -> code.claude.com/docs/en/best-practices +For each line, the audit asks whether removing it would make Claude make mistakes; a line that +fails the test is a cut candidate. + +- **Pointer**: [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md). ## What to include vs exclude -Official include/exclude table (code.claude.com/docs/en/best-practices): +The audit keeps in CLAUDE.md what Claude cannot infer and that holds every session, and flags as +removal candidates what Claude can infer from the code or that changes often: -| Include | Exclude | +| Keep | Flag | |---------|---------| -| Bash commands Claude can't guess | Anything Claude can figure out by reading code | -| Code style rules that differ from defaults | Standard language conventions Claude already knows | -| Testing instructions and preferred test runners | Detailed API documentation (link to docs instead) | -| Repository etiquette (branch naming, PR conventions) | Information that changes frequently | -| Architectural decisions specific to your project | Long explanations or tutorials | -| Developer environment quirks (required env vars) | File-by-file descriptions of the codebase | -| Common gotchas or non-obvious behaviors | Self-evident practices like "write clean code" | - -## Build and test commands +| Shell commands for this repo that the model would not work out | Facts the model can read off the code | +| Style choices that depart from the language's defaults | A language's ordinary conventions | +| How this repo runs its tests, and with which runner | Full API reference (link to it) | +| This repo's branch, commit and pull request habits | Facts that go stale quickly | +| Design decisions this project made | Tutorials and long background | +| Setup quirks of the dev environment, such as env vars it needs | A file-by-file tour of the codebase | +| Traps and surprising behavior | Advice any engineer already follows | -> "Create this file and add instructions that apply to anyone working on the project: build and test commands, coding standards, architectural decisions, naming conventions, and common workflows." -> -> code.claude.com/docs/en/memory, "Set up a project CLAUDE.md" +- **Pointer**: the include and exclude table in + [Write an effective CLAUDE.md](https://code.claude.com/docs/en/best-practices#write-an-effective-claude-md). -Build and test commands lead the list of what project memory is for. The inference cost of omitting -them is stated on the same page, in what `/init` does instead: +## Build and test commands -> "Claude analyzes your codebase and creates a file with build commands, test instructions, and project conventions it discovers." -> -> code.claude.com/docs/en/memory +We expect a project CLAUDE.md to carry the build and test commands, because they lead the list of +what project memory is for and `/init` generates them. A project CLAUDE.md that omits them leaves +those commands to be discovered per session rather than read. Backs C9. -So a project CLAUDE.md that omits them leaves those commands to be discovered per session rather -than read. Backs C9. +- **Pointer**: [Set up a project CLAUDE.md](https://code.claude.com/docs/en/memory#set-up-a-project-claude-md). ## @import syntax -> "CLAUDE.md files can import additional files using `@path/to/import` syntax. Imported files are expanded and loaded into context at launch alongside the CLAUDE.md that references them." -> -> code.claude.com/docs/en/memory +The audit treats `@path/to/import` content as loaded at launch with the file that imports it, so an +import never saves context. Details the audit relies on: -Key details: +- Relative and absolute paths both work; a relative path resolves from the importing file. +- Imports recurse, up to a maximum of 4 hops. +- An external import asks for approval the first time; a declined import stays disabled. +- Typical uses: README, package.json, personal preferences + (`@~/.claude/my-project-instructions.md`). -- Both relative and absolute paths allowed. Relative paths resolve relative to the file containing the import -- Imported files can recursively import other files, max depth 4 hops -- First-time approval dialog for external imports; if declined, stays disabled -- Use for README, package.json, personal preferences (`@~/.claude/my-project-instructions.md`) +- **Pointer**: [Import additional files](https://code.claude.com/docs/en/memory#import-additional-files). ## claudeMdExcludes setting -> "In large monorepos, ancestor CLAUDE.md files may contain instructions that aren't relevant to your work. The `claudeMdExcludes` setting lets you skip specific files by path or glob pattern." -> -> code.claude.com/docs/en/memory +In a large monorepo, the audit recommends `claudeMdExcludes` to skip ancestor CLAUDE.md files that +do not apply: ```json { @@ -138,161 +116,170 @@ Key details: } ``` -- Patterns matched against absolute file paths using glob syntax -- Configurable at any settings layer (user, project, local, managed policy). Arrays merge across layers -- Managed policy CLAUDE.md cannot be excluded +- Patterns are globs matched against absolute file paths. +- Any settings layer may set it (user, project, local, managed policy); the audit combines every + layer's patterns into one list. +- A managed policy CLAUDE.md cannot be excluded. -## Skills vs CLAUDE.md +- **Pointer**: [Exclude specific CLAUDE.md files](https://code.claude.com/docs/en/memory#exclude-specific-claude-md-files). -> "CLAUDE.md is loaded every session, so only include things that apply broadly. For domain knowledge or workflows that are only relevant sometimes, use skills instead. Claude loads them on demand without bloating every conversation." -> -> code.claude.com/docs/en/best-practices +## Skills vs CLAUDE.md - +The audit recommends moving domain knowledge or workflows needed only sometimes out of CLAUDE.md +and unscoped rules into a skill, which loads on demand. -> "Rules load into context every session or when matching files are opened. For task-specific instructions that don't need to be in context all the time, use skills instead, which only load when you invoke them or when Claude determines they're relevant to your prompt." -> -> code.claude.com/docs/en/memory (rules section) +- **Pointer**: [Create skills](https://code.claude.com/docs/en/best-practices#create-skills), + [Set up rules](https://code.claude.com/docs/en/memory#set-up-rules). ## Hooks vs CLAUDE.md -> "Unlike CLAUDE.md instructions which are advisory, hooks are deterministic and guarantee the action happens." -> -> code.claude.com/docs/en/best-practices +The audit recommends a hook, not a CLAUDE.md line, for anything that must happen every time: a +CLAUDE.md line is advisory and a hook is deterministic. + +- **Pointer**: [Set up hooks](https://code.claude.com/docs/en/best-practices#set-up-hooks). ## InstructionsLoaded hook -> "Use the `InstructionsLoaded` hook to log exactly which instruction files are loaded, when they load, and why. This is useful for debugging path-specific rules or lazy-loaded files in subdirectories." -> -> code.claude.com/docs/en/memory (troubleshoot section) +To debug which instruction files load, when and why (path-scoped rules, lazy-loaded subdirectory +files), the audit points to the `InstructionsLoaded` hook. It only observes: it cannot block +loading or change content. -Observability-only: cannot block loading or modify content. +- **Pointer**: [`InstructionsLoaded`](https://code.claude.com/docs/en/hooks#instructionsloaded) and + [Troubleshoot memory issues](https://code.claude.com/docs/en/memory#troubleshoot-memory-issues). ## Specificity -> "Write instructions that are concrete enough to verify." -> -> code.claude.com/docs/en/memory +The audit flags an instruction too vague to verify (format code properly, test your changes, keep +files organized) and proposes a concrete one: an indentation width, a named test command, a named +directory. -Official examples: - -- "Use 2-space indentation" instead of "Format code properly" -- "Run `npm test` before committing" instead of "Test your changes" -- "API handlers live in `src/api/handlers/`" instead of "Keep files organized" +- **Pointer**: [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions). ## Consistency -> "If two rules contradict each other, Claude may pick one arbitrarily. Review your CLAUDE.md files, nested CLAUDE.md files in subdirectories, and `.claude/rules/` periodically to remove outdated or conflicting instructions." -> -> code.claude.com/docs/en/memory +The audit flags two instructions that contradict each other across CLAUDE.md files, nested +CLAUDE.md files and `.claude/rules/`, since the model may pick either. + +- **Pointer**: [Write effective instructions](https://code.claude.com/docs/en/memory#write-effective-instructions) + and [Audit your instruction files](https://code.claude.com/docs/en/memory#audit-your-instruction-files). ## Rules files -> "For larger projects, you can organize instructions into multiple files using the `.claude/rules/` directory. This keeps instructions modular and easier for teams to maintain. Rules can also be scoped to specific file paths, so they only load into context when Claude works with matching files, reducing noise and saving context space." -> -> code.claude.com/docs/en/memory +The audit treats `.claude/rules/` as the place for modular instructions, and a path-scoped rule as +loading only after Claude reads a file its globs match. Features it relies on: -Additional features: +- Symlinks in `.claude/rules/` share rules across projects. +- User-level rules in `~/.claude/rules/` apply to every project and load before project rules. +- A path-specific rule uses `paths:` YAML frontmatter with glob patterns. -- Symlinks supported in `.claude/rules/` to maintain shared rules across projects -- User-level rules in `~/.claude/rules/` apply to every project (loaded before project rules) -- Path-specific rules use `paths:` YAML frontmatter with glob patterns +- **Pointer**: [Set up rules](https://code.claude.com/docs/en/memory#set-up-rules), + [Path-specific rules](https://code.claude.com/docs/en/memory#path-specific-rules), + [User-level rules](https://code.claude.com/docs/en/memory#user-level-rules). -**Path scoping status (verified working 2026-07-24 on Claude Code 2.1.219):** Path scoping defers as documented. A path-scoped rule is not in context at session start and loads when Claude reads a matching file. A first-party repro on 2.1.219 with `paths: ["**/*.tsx"]` found the rule absent at session start, present after reading a matching `.tsx` file, and absent again after reading a non-matching one: deferral works in both directions. No changelog entry or maintainer comment pins the version where this began working, so do not claim a version floor. Recheck trigger: a Claude Code release note or memory-doc change touching rule loading, or any session in which a path-scoped rule is present at session start. +**Path scoping status (verified working 2026-07-24 on Claude Code 2.1.219):** Path scoping defers as +documented. A path-scoped rule is not in context at session start and loads when Claude reads a +matching file. A first-party repro on 2.1.219 with `paths: ["**/*.tsx"]` found the rule absent at +session start, present after reading a matching `.tsx` file, and absent again after reading a +non-matching one: deferral works in both directions. No changelog entry or maintainer comment pins +the version where this began working, so do not claim a version floor. Recheck trigger: a Claude +Code release note or memory-doc change touching rule loading, or any session in which a path-scoped +rule is present at session start. Caveats that do survive, each verified: -- An `@import` **inside** a path-scoped rule defeats the scoping: the imported content inlines at session start whether or not a matching file is ever read. Per code.claude.com/docs/en/memory, "Imported files are expanded and loaded into context at launch". -- Path-scoped content is not inherited by a subagent, and is invisible to teammates and skill-forked contexts. Issue #32906 covers this and is closed as not planned, so it is accepted behavior rather than a pending fix. Basis: `gh api repos/anthropics/claude-code/issues/32906`, which returns `state: closed` and `state_reason: not_planned`. Verified 2026-09-06 against Claude Code 2.1.263. Recheck when that issue reopens or closes as completed, or when the memory page's subagent section changes. -- Non-inheritance is not unreachability: a subagent does receive a path-scoped rule once it reads a covered path itself. Claim: on Claude Code 2.1.268 a non-fork subagent inherits none of its parent's on-demand instruction surfaces, and receives a path-scoped `.claude/rules/` file, or a nested `CLAUDE.md` and the `AGENTS.md` its shim imports, when it reads a path that surface covers; the glob is matched against the requested path, so even a read that finds no file fires it. Basis: first-party probe run inside a dispatched general-purpose subagent on the harness `claude --version` reports as `2.1.268 (Claude Code)`, observing the `Contents of :` block appended to `Read` tool results. As of 2026-09-13. Recheck when the consuming repository's Claude Code minor version moves past 2.1.268, when a release note names subagent context inheritance, memory loading, or path-scoped rule triggering, or when a read of a covered path inside a subagent injects nothing. -- Writing a NEW file does not trigger the rule. The trigger is a read, per code.claude.com/docs/en/memory: "Path-scoped rules trigger when Claude reads files matching the pattern, not on every tool use". -- Excluding `project` from `--setting-sources` also excludes on-demand rules, both path-scoped rules and rules in nested `.claude/rules/` directories (code.claude.com/docs/en/memory). +- An `@import` **inside** a path-scoped rule defeats the scoping: the imported content inlines at + session start whether or not a matching file is ever read, since imports load at launch (pointer: + [Import additional files](https://code.claude.com/docs/en/memory#import-additional-files)). +- Path-scoped content is not inherited by a subagent, and is invisible to teammates and + skill-forked contexts. Issue #32906 covers this and is closed as not planned, so it is accepted + behavior rather than a pending fix. Pointer: `gh api repos/anthropics/claude-code/issues/32906`, + which returns `state: closed` and `state_reason: not_planned`. As of: 2026-09-06, Claude Code + 2.1.263. Recheck trigger: that issue reopens or closes as completed, or the memory page's + subagent section changes. +- Non-inheritance is not unreachability: a subagent does receive a path-scoped rule once it reads + a covered path itself. On Claude Code 2.1.268 a non-fork subagent inherits none of its parent's + on-demand instruction surfaces, and receives a path-scoped `.claude/rules/` file, or a nested + `CLAUDE.md` and the `AGENTS.md` its shim imports, when it reads a path that surface covers; the + glob is matched against the requested path, so even a read that finds no file fires it. + Pointer: our probe run inside a dispatched general-purpose subagent on the harness + `claude --version` reports as `2.1.268 (Claude Code)`, observing the `Contents of :` block + appended to `Read` tool results. As of: 2026-09-13. Recheck trigger: the consuming repository's + Claude Code minor version moves past 2.1.268, a release note names subagent context inheritance, + memory loading, or path-scoped rule triggering, or a read of a covered path inside a subagent + injects nothing. +- Writing a NEW file does not trigger the rule. We treat a read, not any tool use, as the trigger + (pointer: [Path-specific rules](https://code.claude.com/docs/en/memory#path-specific-rules)). +- Excluding `project` from `--setting-sources` also drops the on-demand rules of both kinds: + path-scoped ones, and those kept in a nested `.claude/rules/` (pointer: + [Set up rules](https://code.claude.com/docs/en/memory#set-up-rules)). ## Auto-memory limits -> "The first 200 lines of `MEMORY.md`, or the first 25KB, whichever comes first, are loaded at the start of every conversation. Content beyond that threshold is not loaded at session start." -> -> code.claude.com/docs/en/memory - - - -> "This limit applies only to `MEMORY.md`. CLAUDE.md files are loaded in full regardless of length, though shorter files produce better adherence." -> -> code.claude.com/docs/en/memory +The audit's size gate treats only the start of `MEMORY.md` as loaded at session start: 200 lines, +cut shorter when those lines pass 25KB. Everything after that is unread. The limit applies to +`MEMORY.md` only; a CLAUDE.md loads in full. The gate measures only the content that loads: YAML +frontmatter and block-level HTML comments are stripped first and do not count. - - -> "The check measures only the content that loads: YAML frontmatter and block-level HTML comments are stripped before the index is loaded, so they don't count toward the limits." -> -> code.claude.com/docs/en/memory (limit check on writes to MEMORY.md) +- **Pointer**: [How it works](https://code.claude.com/docs/en/memory#how-it-works) (auto memory). ## Auto-memory storage -> "Each project gets its own memory directory at `~/.claude/projects//memory/`. The `` path is derived from the git repository, so all worktrees and subdirectories within the same repo share one auto memory directory." -> -> code.claude.com/docs/en/memory - -**`autoMemoryDirectory` setting:** Override default location; read from any settings scope: user, project, local, policy, or `--settings`. From a project's `.claude/settings.json` or `.claude/settings.local.json`, the value is honored only after you accept the workspace trust dialog for that folder (the same gate that governs hooks). +The audit resolves the auto-memory directory as `~/.claude/projects//memory/`, with +`` derived from the repository, so every worktree and subdirectory of one repository shares +one directory. -## Subagent persistent memory +**`autoMemoryDirectory` setting:** overrides the default location; read from any settings scope: +user, project, local, policy, or `--settings`. From a project's `.claude/settings.json` or +`.claude/settings.local.json`, the value is honored only after the workspace trust dialog for that +folder is accepted (the same gate that governs hooks). -> "The `memory` field gives the subagent a persistent directory that survives across conversations." -> -> code.claude.com/docs/en/sub-agents +- **Pointer**: [Storage location](https://code.claude.com/docs/en/memory#storage-location). -Three scopes: +## Subagent persistent memory -| Scope | Location | Use when | -|-------|----------|----------| -| `user` | `~/.claude/agent-memory//` | Learnings across all projects | -| `project` | `.claude/agent-memory//` | Project-specific, shareable via version control | -| `local` | `.claude/agent-memory-local//` | Project-specific, not checked in | +The audit reads a subagent's `memory` field as one of the scopes the page defines and resolves +each scope to the directory the page lists for it. It applies the same `MEMORY.md` load limit it +applies to auto memory. It recommends `project` unless the agent's learnings must stay out of +version control. -- Same 200-line/25KB limit on subagent's MEMORY.md -- `project` is recommended default scope -- Read/Write/Edit tools auto-enabled for memory management +- **Pointer**: [Enable persistent memory](https://code.claude.com/docs/en/sub-agents#enable-persistent-memory). +- **As of**: 2026-10-01 +- **Recheck trigger**: the section adds or renames a scope, moves a directory, or changes its load limit. ## HTML comments -> "Block-level HTML comments (``) in CLAUDE.md files are stripped before the content is injected into Claude's context. Use them to leave notes for human maintainers without spending context tokens on them. Comments inside code blocks are preserved." -> -> code.claude.com/docs/en/memory +The audit treats block-level HTML comments in a CLAUDE.md as stripped before injection, so they are +the place for maintainer notes that spend no context; comments inside code blocks are kept. A +direct Read of the file still shows them. -When you open a CLAUDE.md file directly with the Read tool, comments remain visible. +- **Pointer**: [How CLAUDE.md files load](https://code.claude.com/docs/en/memory#how-claude-md-files-load). ## Boris Cherny (CC creator) -> "Anytime we see Claude do something incorrectly we add it to the CLAUDE.md, so Claude knows not to do it next time." -> -> howborisusesclaudecode.com +Practices we take from Boris Cherny's published workflow, in our words: - +- Add a CLAUDE.md line whenever Claude makes a mistake you do not want repeated. +- Keep editing CLAUDE.md until the mistake rate measurably drops. +- End a correction by asking Claude to update CLAUDE.md so the mistake is not repeated. +- **Auto-Dream (memory consolidation):** a subagent that reviews past sessions and consolidates + what matters into cleaner memory. +- **`@.claude` PR tags:** `@.claude` tags in PR comments trigger CLAUDE.md updates during review. +- **Progressive disclosure for skills:** a skill is a folder, with SKILL.md as the hub and spoke + files doing the work. -> "Ruthlessly edit your CLAUDE.md over time. Keep iterating until Claude's mistake rate measurably drops." -> -> howborisusesclaudecode.com - - - -> "End corrections with: 'Update your CLAUDE.md so you don't make that mistake again'" -> -> howborisusesclaudecode.com - -**Auto-Dream (memory consolidation):** Boris describes a subagent that "reviews past sessions, keeps what matters, removes what doesn't, and merges insights into cleaner structured memory." - -**`@.claude` PR tags:** Use `@.claude` tags in PR comments to trigger automatic CLAUDE.md updates during code reviews. - -**Progressive disclosure for skills:** "A skill is a folder, not a file. SKILL.md is the hub, spoke files do the work." +- **Pointer**: [howborisusesclaudecode.com](https://howborisusesclaudecode.com). ## Style enforcement -> "Never send an LLM to do a linter's job. LLMs are comparably expensive and incredibly slow." -> -> humanlayer.dev/blog/writing-a-good-claude-md +The audit flags a CLAUDE.md line that asks the model to do a linter's job and recommends a linter +or hook instead: a model is slower and costlier than a linter at that. + +- **Pointer**: [humanlayer.dev, writing a good CLAUDE.md](https://humanlayer.dev/blog/writing-a-good-claude-md). ## Compaction by steering method (June 2026) -Per [Steering Claude Code](https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more) and [memory docs](https://code.claude.com/docs/en/memory), what survives `/compact` vs what reloads on demand: +The audit models what survives `/compact` and what reloads on demand per destination as follows, +and prices a recommended move with that destination's row: | Method | Session start | After compaction | On-demand trigger | |--------|---------------|------------------|-------------------| @@ -305,34 +292,42 @@ Per [Steering Claude Code](https://claude.com/blog/steering-claude-code-skills-h | Auto-memory MEMORY.md | First 200 lines / 25KB | Persists on disk | None | | Output style | If non-default | Persists for session | `/config` | +- **Pointer**: [Troubleshoot memory issues](https://code.claude.com/docs/en/memory#troubleshoot-memory-issues), + the "Instructions seem lost after `/compact`" entry, and + [Skill content lifecycle](https://code.claude.com/docs/en/skills#skill-content-lifecycle) + (correlate with [Steering Claude Code](https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more)). + `AGENTS.md` is its own row's worth of behavior, and the row depends on the repository. The memory -doc's `AGENTS.md` section used to state "Claude Code reads `CLAUDE.md`, not `AGENTS.md`", and -prescribed an `@AGENTS.md` import or a symlink as the way to make one load; that sentence is gone -from the page as fetched 2026-09-19 and is quoted here only as the superseded basis. The page now -says Claude reads `AGENTS.md` as the project instructions where there is no `CLAUDE.md`, -`.claude/CLAUDE.md` or `CLAUDE.local.md` in the working directory or above it, needs v2.1.277 or -later to do so, and cannot in some sessions (the built-in `agents-md` plugin disabled, in some cases -the first session after an upgrade; before v2.1.281 also sessions on Amazon Bedrock or with -telemetry disabled), where the import is still what carries it. No hooks setting stops it: the -settings and flags that stop installed mods "don't stop built-in mods" -([mods overview](https://code.claude.com/docs/en/plugins/mods/overview), "Mods built into Claude -Code", fetched 2026-10-01). - -- **Claim**: an `AGENTS.md` loads either directly, where no `CLAUDE.md` displaces it and support is - available, or through a `CLAUDE.md` that imports or symlinks it, on that `CLAUDE.md`'s row. - Directly read, it fires no `InstructionsLoaded` hook, and `/memory` lists it from v2.1.280 - (before that, `/memory` and `/context` did not); imported, it behaves as part of its - `CLAUDE.md`. -- **Basis**: code.claude.com/docs/en/memory, "AGENTS.md", "When Claude Code reads AGENTS.md", "When - AGENTS.md support is unavailable", "Where AGENTS.md differs from CLAUDE.md", "My AGENTS.md isn't - loading"; fetched 2026-09-29. -- **As of**: 2026-09-29. -- **Recheck trigger**: that section changes which file names count for the check or which sessions - lack support, the `/memory` listing sentence or the difference table changes, or a release note - names `AGENTS.md`. +page changed its `AGENTS.md` guidance between 2026-06-20 and 2026-09-19, so the row below is +derived from the page as of 2026-09-29, not from the earlier shim-only guidance. + +The audit models an `AGENTS.md` as loading either directly, where no `CLAUDE.md`, +`.claude/CLAUDE.md` or `CLAUDE.local.md` in the working directory or above it displaces it and +`AGENTS.md` support is available, or through a `CLAUDE.md` that imports or symlinks it, on that +`CLAUDE.md`'s row. Read directly, it fires no `InstructionsLoaded` hook, and `/memory` lists it +from v2.1.280 (before that, `/memory` and `/context` did not); imported, it behaves as part of its +`CLAUDE.md`. No hooks setting makes direct support unavailable: the loader is a built-in mod, and +the audit never counts `disableAllHooks` or `allowManagedHooksOnly` against it. + +- **Pointer**: for when an `AGENTS.md` is read, when support is unavailable, and how it differs + from `CLAUDE.md`, see [AGENTS.md](https://code.claude.com/docs/en/memory#agents-md), + [When Claude Code reads AGENTS.md](https://code.claude.com/docs/en/memory#when-claude-code-reads-agents-md), + [When AGENTS.md support is unavailable](https://code.claude.com/docs/en/memory#when-agents-md-support-is-unavailable), + [Where AGENTS.md differs from CLAUDE.md](https://code.claude.com/docs/en/memory#where-agents-md-differs-from-claude-md) + and the "My AGENTS.md isn't loading" entry under + [Troubleshoot memory issues](https://code.claude.com/docs/en/memory#troubleshoot-memory-issues); + for which settings stop built-in mods, the "Mods built into Claude Code" section of + . +- **As of**: 2026-09-29 for the memory page; 2026-10-01 for the mods overview +- **Recheck trigger**: those sections change which file names displace an `AGENTS.md` or which + sessions lack support, the `/memory` listing or the difference table changes, the mods overview + changes which settings stop a built-in mod, or a release note names `AGENTS.md`. ## No official scoring rubric -There is no official scoring rubric for CLAUDE.md quality. The 6-category, 100-point rubric shipped by the `claude-md-improver` skill of the `claude-md-management` plugin (Anthropic's `claude-plugins-official` marketplace, not this one) is invented by the plugin author, not derived from official documentation. +We use no scoring rubric for CLAUDE.md quality, because no official one exists. The 6-category, +100-point rubric shipped by the `claude-md-improver` skill of the `claude-md-management` plugin +(Anthropic's `claude-plugins-official` marketplace, not this one) is invented by the plugin author, +not derived from official documentation. -The official quality measure is the **deletion test**: "Would removing this cause Claude to make mistakes?" +The quality measure the audit uses is the [deletion test](#the-deletion-test). diff --git a/plugins/harness-memory/skills/stateless/SKILL.md b/plugins/harness-memory/skills/stateless/SKILL.md index 75d2202c9e..25743b8ede 100644 --- a/plugins/harness-memory/skills/stateless/SKILL.md +++ b/plugins/harness-memory/skills/stateless/SKILL.md @@ -23,13 +23,14 @@ Inspect and disable Claude Code **auto memory**, the store Claude writes for its directory per repo (`~/.claude/projects//memory/`, relocatable via `autoMemoryDirectory`). Governs auto-memory only. Not in scope: CLAUDE.md / CLAUDE.local.md / `.claude/rules/` (use `/harness-memory:audit`), transcripts, history, or shell snapshots. For the -official full per-project wipe, use `claude project purge`. What it does and does not delete is -quoted verbatim in -[reference/official-guidance.md](reference/official-guidance.md); the deletion plan and flags live -in the [claude-directory doc](https://code.claude.com/docs/en/claude-directory). +official full per-project wipe, use `claude project purge`. +[reference/official-guidance.md](reference/official-guidance.md), "Out of scope for this skill", +records how this skill treats that command's scope; read the deletion plan and flags at +[Clear local data](https://code.claude.com/docs/en/claude-directory#clear-local-data). -Criteria and exact doc quotes live in [reference/official-guidance.md](reference/official-guidance.md); -re-fetch the source pages listed there if a fact is load-bearing before you act. +[reference/official-guidance.md](reference/official-guidance.md) holds the decisions this skill +acts on, each with a pointer to its docs section and none of the docs' text; re-read the pointed-at +section before acting on a load-bearing fact. ## Scope @@ -40,10 +41,10 @@ re-fetch the source pages listed there if a fact is load-bearing before you act. | `CLAUDE_CODE_DISABLE_AUTO_MEMORY` | OS env or settings `env` block | Yes. Reads and writes | | CLAUDE.md / `.claude/rules/` | repo + user | No. Use `/harness-memory:audit` | | CLAUDE.local.md | repo only, no user-scope equivalent | No. Use `/harness-memory:audit` | -| Transcripts | `~/.claude/projects//` | No. Auto-cleaned by `cleanupPeriodDays`; `claude project purge` deletes this project's now | -| Prompt history | `~/.claude/history.jsonl` | No. Persists indefinitely, not swept by `cleanupPeriodDays`; `claude project purge` filters this project's lines | -| Session files | `~/.claude/sessions/` | No. One file per running session, cleared when the session exits rather than age-swept; not in `claude project purge`'s deletion list | -| Shell snapshots / backups | `~/.claude/shell-snapshots/`, `~/.claude/backups/` | No. Swept by `cleanupPeriodDays`, but not project-scoped, so `claude project purge` leaves them untouched | +| Transcripts | `~/.claude/projects//` | No. How we treat its age sweep and `claude project purge`: official-guidance.md, "Out of scope for this skill" | +| Prompt history | `~/.claude/history.jsonl` | No. Same record | +| Session files | `~/.claude/sessions/` | No. Same record | +| Shell snapshots / backups | `~/.claude/shell-snapshots/`, `~/.claude/backups/` | No. Same record | | Claude Desktop / claude.ai memory | server-side account | Direction only. See [context/desktop.md](context/desktop.md) | ## Argument parsing @@ -56,15 +57,14 @@ re-fetch the source pages listed there if a fact is load-bearing before you act. | `purge` | **Destructive.** Delete the auto-memory files. Reads `autoMemoryDirectory` at every scope first, shows a manifest, offers an opt-in pre-delete backup, and deletes only after explicit confirmation. | | `purge all` | **Destructive, machine-wide.** Same flow with every per-project store as the candidate set, one combined manifest, and ONE combined gate stating the total count and every directory. | -## Precedence (documented) +## Precedence -`CLAUDE_CODE_DISABLE_AUTO_MEMORY` **overrides** `autoMemoryEnabled`: per the env-vars doc, `=1` -disables and `=0` forces auto memory *on* even when `autoMemoryEnabled: false` would disable -it. When the env var is unset, `autoMemoryEnabled` (by settings precedence) governs. So a set -env var of `0` alongside `autoMemoryEnabled: false` means auto memory is effectively **on**. -`status` must report the env var as authoritative whenever it is set. `disable` sets the env -var to `1` (the authoritative lever) and `autoMemoryEnabled: false` together. See the -reference file's "Precedence: the env var overrides the setting (VERIFIED)". +This skill treats `CLAUDE_CODE_DISABLE_AUTO_MEMORY` as authoritative whenever it is set: `1` +reports auto memory off, and `0` reports it **on** even against `autoMemoryEnabled: false`. Only +when the env var is unset does `autoMemoryEnabled` (by settings precedence) decide. `status` must +report the env var as authoritative whenever it is set. `disable` sets the env var to `1` and +`autoMemoryEnabled: false` together. The reference file's "Precedence: the env var overrides the +setting (VERIFIED)" holds the pointer, as-of date and recheck trigger. ## Actions @@ -78,17 +78,21 @@ wants to be stateless everywhere, not just in this repo. The `context/` files write each bundled script as `/scripts/.sh`, where `` is this skill's directory: `${CLAUDE_SKILL_DIR}`. Put that path in place of the -placeholder before running a command; a file read through the Read tool is not substituted, and -the Bash tool's environment has no `CLAUDE_PLUGIN_ROOT`. Basis: the plugins reference, "Where each -variable resolves", verified 2026-09-27; recheck when that table adds supporting files. +placeholder before running a command. We never put a `${…}` token in those files: we do not rely +on one being substituted in a file read through the Read tool, or on the Bash tool's environment +carrying `CLAUDE_PLUGIN_ROOT`. + +- **Pointer**: for where each `${…}` reference resolves, see + [Where each variable resolves](https://code.claude.com/docs/en/plugins-reference#where-each-variable-resolves). +- **As of**: 2026-09-27 +- **Recheck trigger**: that table adds supporting files to where a `${…}` reference resolves. ## Boundary, the built-in `/memory` command "Turn off auto memory" and "what has Claude saved" can land on either. -- **`/memory` (built-in command)**: an interactive dialog to edit CLAUDE.md files, turn auto memory - on or off, and view auto memory entries in the running session. It is reserved for the person to - run; the model does not invoke it. +- **`/memory` (built-in command)**: the person's interactive editor for CLAUDE.md files and the + auto-memory toggle and entries in the running session. This skill never runs it. - **This skill (marketplace plugin).** Reports the effective auto-memory state across every settings scope and the env var that overrides them, disables it durably through both levers, and purges the store behind a manifest and a confirmation gate. @@ -108,19 +112,28 @@ which `status` reports. ## Gotchas -- **Precedence**: `CLAUDE_CODE_DISABLE_AUTO_MEMORY` overrides `autoMemoryEnabled` (`=0` forces - on even against `autoMemoryEnabled: false`). A set env var is authoritative in `status`. (See above.) -- **`autoMemoryDirectory` relocates the store** and is read from *any* scope. The snapshot - prints the slug-derived default only. `purge` and `status` must read the override at every - scope or they act on the wrong directory. -- **`CLAUDE_CONFIG_DIR` relocates the whole config root**: when set, the user `settings.json` - *and* the `projects//memory/` tree live under it, not `~/.claude`. All scope and - memory-dir resolution honors `${CLAUDE_CONFIG_DIR:-~/.claude}` (scripts + workflows); the - snapshot reports the resolved root, and `purge`'s relocation check treats it as expected. -- **Windows managed policy** can live in the registry (`HKLM`/`HKCU\SOFTWARE\Policies\ClaudeCode`), - not a file. `scope-report.sh` can't read it. Report managed scope as unread, don't assume empty. -- **`disable` applies next session**, not immediately: the setting and `env` block are read at - startup. Tell the user to restart / start a new session. +- **Precedence**: a set `CLAUDE_CODE_DISABLE_AUTO_MEMORY` is authoritative in `status`, `0` + included. (See above.) +- **`autoMemoryDirectory`**: we treat it as able to relocate the store from *any* scope + (official-guidance.md, "Storage location"). The snapshot prints the slug-derived default only. + `purge` and `status` must read the override at every scope or they act on the wrong directory. +- **`CLAUDE_CONFIG_DIR`**: we treat it as relocating the whole config root, the user + `settings.json` *and* the `projects//memory/` tree included (official-guidance.md, + "CLAUDE_CONFIG_DIR relocates the whole config root"). All scope and memory-dir resolution honors + `${CLAUDE_CONFIG_DIR:-~/.claude}` (scripts + workflows); the snapshot reports the resolved root, + and `purge`'s relocation check treats it as expected. +- **Windows managed policy**: we treat it as possibly held in the registry + (`HKLM`/`HKCU\SOFTWARE\Policies\ClaudeCode`) rather than a file, which `scope-report.sh` can't + read. Report managed scope as unread, don't assume empty. +- **`disable` leaves this session's loaded memory in place.** What auto memory loaded at startup + stays in the current context whether or not the setting reloads mid-session, so tell the user a + new session is the first one that starts without it. + - **Pointer**: for which settings edits reach a running session, see + ; for the toggle, see + . + - **As of**: 2026-10-01 + - **Recheck trigger**: either section changes whether an `autoMemoryEnabled` or `env` edit + reaches a running session. - **Tracked `settings.json`**: a live edit to a dotfile-manager-tracked settings file must be backfilled to the source; never run an `apply` that could revert the edit. - **Desktop / claude.ai memory is server-side**. `purge` cannot delete it; give direction only. diff --git a/plugins/harness-memory/skills/stateless/context/purge.md b/plugins/harness-memory/skills/stateless/context/purge.md index 1df5e020e2..1eaca509a1 100644 --- a/plugins/harness-memory/skills/stateless/context/purge.md +++ b/plugins/harness-memory/skills/stateless/context/purge.md @@ -25,8 +25,9 @@ included") and offer to additionally check any repos the user names. ## Step 1: Resolve EVERY candidate directory -The store may be relocated by `autoMemoryDirectory`, which is read from **any** settings scope -(user, project, local, policy, `--settings`). Miss that and you purge the wrong place. So: +We treat `autoMemoryDirectory` as able to relocate the store from **any** settings scope (user, +project, local, policy, `--settings`); [official-guidance.md](../reference/official-guidance.md), +"Storage location", holds the pointer. Miss that and you purge the wrong place. So: 1. Read `autoMemoryDirectory` from every present settings scope: managed, local, project, and user. The snapshot in SKILL.md lists which files exist; Read each. Expand `~/` to `$HOME`. @@ -66,10 +67,10 @@ Present to the user: it could point at an unrelated directory. - That this deletes auto-memory notes only, **not** CLAUDE.md, rules, transcripts, or history. If the intent is the full per-project wipe, point to `claude project purge` instead, and state - its scope to the user (what it deletes and what it leaves alone) - from the verbatim quotes in - [reference/official-guidance.md](../reference/official-guidance.md) rather than from memory. - owns the deletion plan and flags. + its scope to the user (what it deletes and what it leaves alone) from the record in + [reference/official-guidance.md](../reference/official-guidance.md), "Out of scope for this + skill", rather than from memory. Read the deletion plan and flags at + [Clear local data](https://code.claude.com/docs/en/claude-directory#clear-local-data). - If `$manifest` is empty, report that there is nothing to purge and stop (no-op). ## Step 3: Confirmation gate (with backup offer) @@ -158,7 +159,7 @@ otherwise leaving the empty directory is harmless. stateless, point to `disable` (or run it now if they ask) so Claude doesn't immediately re-accumulate memory. - If the intent was wiping everything Claude holds for this repo, point to - `claude project purge` (Step 2's pointer). Its scope is the full per-project one quoted in + `claude project purge` (Step 2's pointer). Its scope is the full per-project one recorded in [reference/official-guidance.md](../reference/official-guidance.md), not auto memory alone. - If the user wants to be stateless everywhere, summarize the Claude Desktop / claude.ai account store steps in [desktop.md](desktop.md). That store is server-side and cannot be diff --git a/plugins/harness-memory/skills/stateless/reference/official-guidance.md b/plugins/harness-memory/skills/stateless/reference/official-guidance.md index 6c73ff4f86..f603095500 100644 --- a/plugins/harness-memory/skills/stateless/reference/official-guidance.md +++ b/plugins/harness-memory/skills/stateless/reference/official-guidance.md @@ -1,90 +1,73 @@ # Official Claude Code Guidance on Auto Memory State -Last researched: 2026-07-22; code.claude.com/docs/en/claude-directory, -code.claude.com/docs/en/settings, and code.claude.com/docs/en/cli-reference verified 2026-08-10 -(the other sources below were not re-checked on that date) -Sources: [code.claude.com/docs/en/memory](https://code.claude.com/docs/en/memory), -[code.claude.com/docs/en/settings](https://code.claude.com/docs/en/settings), -[code.claude.com/docs/en/env-vars](https://code.claude.com/docs/en/env-vars), -[code.claude.com/docs/en/claude-directory](https://code.claude.com/docs/en/claude-directory), -[code.claude.com/docs/en/cli-reference](https://code.claude.com/docs/en/cli-reference) +Each section states what this skill does, in our words, and points at the section of the official +page that covers the topic. Read the page there for its wording; this file stores none of it. + +Last researched: 2026-07-22; the claude-directory, settings and cli-reference pointers verified +2026-08-10 (the other sources below were not re-checked on that date) +Sources: [memory](https://code.claude.com/docs/en/memory), +[settings](https://code.claude.com/docs/en/settings), +[env-vars](https://code.claude.com/docs/en/env-vars), +[claude-directory](https://code.claude.com/docs/en/claude-directory), +[cli-reference](https://code.claude.com/docs/en/cli-reference) Refresh this file from current official docs before relying on it (re-fetch every source listed above). -**Recheck trigger:** re-derive every claim below when any of these becomes observable: the +**Recheck trigger:** re-derive every decision below when any of these becomes observable: the `/memory` command gains, loses or renames its auto-memory toggle; the `autoMemoryEnabled` setting or the `CLAUDE_CODE_DISABLE_AUTO_MEMORY` environment variable changes name, default or semantics; the per-project memory path under `~/.claude/projects//memory/` moves; or any of the five source pages above changes its auto-memory section. A Claude Code release note touching memory, settings or the CLI reference is the usual way one of these surfaces. The dates above record when -the claims last matched their sources and confer no standing authority on their own. +the decisions last matched their sources and confer no standing authority on their own. --- ## What auto memory is -> "Auto memory lets Claude accumulate knowledge across sessions without you writing -> anything. Claude saves notes for itself as it works: build commands, debugging insights, -> architecture notes, code style preferences, and workflow habits." -> -> code.claude.com/docs/en/memory - -Distinct from CLAUDE.md (which **you** write). This skill governs only the Claude-written -auto-memory store, not CLAUDE.md / CLAUDE.local.md / `.claude/rules/` (the sibling +This skill governs only the notes Claude writes for itself across sessions (auto memory), never +CLAUDE.md / CLAUDE.local.md / `.claude/rules/`, which **you** write (the sibling `/harness-memory:audit` skill owns that instruction layer). -## Enable / disable +- **Pointer**: [Auto memory](https://code.claude.com/docs/en/memory#auto-memory). -> "Auto memory is on by default. To toggle it, open `/memory` in a session and use the auto -> memory toggle, which saves `autoMemoryEnabled` to your user settings at -> `~/.claude/settings.json`. To turn it off for a single project, set `autoMemoryEnabled` in -> that project's settings" -> -> code.claude.com/docs/en/memory +## Enable / disable -> "To disable auto memory via environment variable, set `CLAUDE_CODE_DISABLE_AUTO_MEMORY=1`." -> -> code.claude.com/docs/en/memory +This skill treats auto memory as on by default, toggled by `/memory` (which writes +`autoMemoryEnabled` to user settings), settable per project through `autoMemoryEnabled` in that +project's settings, and disabled by `CLAUDE_CODE_DISABLE_AUTO_MEMORY`, either as an OS environment +variable or in a settings file's `env` block. With `autoMemoryEnabled: false`, it treats the +auto-memory directory as neither read nor written. -> "When `false`, Claude does not read from or write to the auto memory directory. You can -> also toggle this with `/memory` during a session. To disable via environment variable, set -> `CLAUDE_CODE_DISABLE_AUTO_MEMORY` in `env`" -> -> code.claude.com/docs/en/settings (`autoMemoryEnabled` description) +- **Pointer**: [Enable or disable auto memory](https://code.claude.com/docs/en/memory#enable-or-disable-auto-memory) + and [`autoMemoryEnabled`](https://code.claude.com/docs/en/settings-reference#automemoryenabled). ### Precedence: the env var overrides the setting (VERIFIED) -> "`CLAUDE_CODE_DISABLE_AUTO_MEMORY` | Set to `1` to disable auto memory. Set to `0` to force -> auto memory on even when `--bare` mode or `autoMemoryEnabled: false` would otherwise disable -> it. When disabled, Claude does not create or load auto memory files" -> -> code.claude.com/docs/en/env-vars +When the env var is set (to `0` or `1`), this skill treats it as **overriding** +`autoMemoryEnabled`: `=1` disables, `=0` forces auto memory on even against +`autoMemoryEnabled: false` or `--bare` mode. When the env var is unset, `autoMemoryEnabled` +(resolved by settings precedence) governs. `status` reports the env var as authoritative whenever it +is set. A set env var of `0` alongside `autoMemoryEnabled: false` means auto memory is effectively +**on**. `disable` sets the env var to `1` (the strong, authoritative lever) and +`autoMemoryEnabled: false` together, so the state is unambiguous and survives the env var later +being unset. -So when the env var is set (to `0` or `1`), it **overrides** `autoMemoryEnabled`: `=1` -disables, `=0` forces on even against `autoMemoryEnabled: false`. When the env var is unset, -`autoMemoryEnabled` (resolved by settings precedence) governs. `status` reports the env var as -authoritative whenever it is set. A set env var of `0` alongside `autoMemoryEnabled: false` -means auto memory is effectively **on**. `disable` sets the env var to `1` (the strong, -authoritative lever) and `autoMemoryEnabled: false` together, so the state is unambiguous and -survives the env var later being unset. +- **Pointer**: the `CLAUDE_CODE_DISABLE_AUTO_MEMORY` row of + [Environment variables](https://code.claude.com/docs/en/env-vars). ## Storage location -> "Each project gets its own memory directory at `~/.claude/projects//memory/`. The -> `` path is derived from the git repository, so all worktrees and subdirectories -> within the same repo share one auto memory directory. Outside a git repo, the project root -> is used instead." -> -> code.claude.com/docs/en/memory - -> "To store auto memory in a different location, set `autoMemoryDirectory` in your -> `settings.json`. It is read from any settings scope: user, project, local, policy, or -> `--settings`. ... The value must be an absolute path or start with `~/`. When set in a -> project's `.claude/settings.json` or `.claude/settings.local.json`, the value is honored -> only after you accept the workspace trust dialog for that folder" -> -> code.claude.com/docs/en/memory +This skill resolves the default store as `~/.claude/projects//memory/`, with `` +derived from the repository (every worktree and subdirectory of one repository shares one +directory) and from the project root outside a repository. It reads `autoMemoryDirectory` at every +settings scope (user, project, local, policy, `--settings`), accepts only an absolute or `~/` +path, and honors a project or local value only once the folder's workspace trust dialog is +accepted. + +- **Pointer**: [Storage location](https://code.claude.com/docs/en/memory#storage-location) and + [`autoMemoryDirectory`](https://code.claude.com/docs/en/settings-reference#automemorydirectory). **What `purge` depends on:** because `autoMemoryDirectory` is read from *any* scope, the real memory dir may not be the slug-derived default. Purge must read that key at every scope @@ -92,140 +75,84 @@ before it enumerates what to delete, or it can miss (and fail to purge) a reloca ### CLAUDE_CONFIG_DIR relocates the whole config root -> "On Windows, `~/.claude` resolves to `%USERPROFILE%\.claude`. If you set `CLAUDE_CONFIG_DIR`, -> every `~/.claude` path on this page lives under that directory instead." -> -> code.claude.com/docs/en/claude-directory (the page scopes settings AND memory under `~/.claude`) - -So the config root is `${CLAUDE_CONFIG_DIR:-~/.claude}`: when the env var is set, the user -`settings.json` and the `projects//memory/` tree both live under it. Every scope and -memory-dir resolution in this skill (the `scope-report.sh` snapshot, the shared -`resolve-memory-dir.sh`, and the disable/purge workflows) resolves the config root this way, so -a relocated root is honored rather than mistaken for an `autoMemoryDirectory` override. +This skill resolves the config root as `${CLAUDE_CONFIG_DIR:-~/.claude}` (`~/.claude` is +`%USERPROFILE%\.claude` on Windows): when the env var is set, the user `settings.json` and the +`projects//memory/` tree both live under it. Every scope and memory-dir resolution in this +skill (the `scope-report.sh` snapshot, the shared `resolve-memory-dir.sh`, and the disable/purge +workflows) resolves the config root this way, so a relocated root is honored rather than mistaken +for an `autoMemoryDirectory` override. -The directory holds a `MEMORY.md` index plus optional topic files (layout per -code.claude.com/docs/en/memory): +- **Pointer**: the `CLAUDE_CONFIG_DIR` row of + [Environment variables](https://code.claude.com/docs/en/env-vars), and + [claude-directory](https://code.claude.com/docs/en/claude-directory#explore-the-directory). -```text -~/.claude/projects//memory/ -├── MEMORY.md # Concise index, loaded into every session -├── debugging.md # Detailed notes on debugging patterns -├── api-conventions.md # API design decisions -└── ... # Any other topic files Claude creates -``` +The directory holds a `MEMORY.md` index plus topic files; the layout is in the Storage location section. -> "Auto memory files are plain markdown you can edit or delete at any time." -> -> code.claude.com/docs/en/memory +This skill treats every file there as plain markdown a person may edit or delete. There is no +auto-memory-only built-in command, so selective deletion is manual removal of these files. +`claude project purge` deletes the store only as part of the full per-project wipe (see "Out of +scope" below). -There is no auto-memory-only built-in command, so selective deletion is manual removal of these -files. `claude project purge` deletes the store only as part of the full per-project wipe (see -"Out of scope" below). +- **Pointer**: [Storage location](https://code.claude.com/docs/en/memory#storage-location) and + [Audit and edit your memory](https://code.claude.com/docs/en/memory#audit-and-edit-your-memory). ## Settings scopes and precedence -> "Settings apply in order of precedence. From highest to lowest: -> -> 1. **Managed settings** (server-managed, MDM/OS-level policies, or managed settings) -> 2. **Command line arguments** -> 3. **Local project settings** (`.claude/settings.local.json`) -> 4. **Shared project settings** (`.claude/settings.json`) -> 5. **User settings** (`~/.claude/settings.json`)" -> -> code.claude.com/docs/en/settings (verified 2026-08-10; each item's nested detail bullets are -> omitted, and item 1's three parenthetical links are flattened to their labels) - -> "Cannot be overridden by any other level, including command line arguments, apart from the -> exceptions in the bullets below" -> -> code.claude.com/docs/en/settings (a nested bullet under item 1, verified 2026-08-10) - -Item 1's exception bullets are longer and more varied than is useful to enumerate here. Read -them on the page. What matters here is a negative: none of them names `autoMemoryEnabled`, -`CLAUDE_CODE_DISABLE_AUTO_MEMORY`, or auto memory at all (verified 2026-08-10), so no lower -settings scope overrides a managed `autoMemoryEnabled` value. That negative governs settings -scopes only: the `CLAUDE_CODE_DISABLE_AUTO_MEMORY` environment variable sits outside settings -precedence and, when set, still overrides the effective value, managed or not (see "Precedence: -the env var overrides the setting" above). +This skill resolves settings highest first: managed settings, command-line arguments, local project +settings (`.claude/settings.local.json`), shared project settings (`.claude/settings.json`), user +settings (`~/.claude/settings.json`). It treats a managed value as overriding every lower scope. +The managed-precedence exceptions are many and are read on the page; none of them names +`autoMemoryEnabled`, `CLAUDE_CODE_DISABLE_AUTO_MEMORY`, or auto memory at all (checked +2026-08-10), so no lower settings scope overrides a managed `autoMemoryEnabled` value. That +negative governs settings scopes only: the `CLAUDE_CODE_DISABLE_AUTO_MEMORY` environment variable +sits outside settings precedence and, when set, still overrides the effective value, managed or not +(see "Precedence: the env var overrides the setting" above). Managed settings live outside the repo (macOS `/Library/Application Support/ClaudeCode/`, Linux/WSL `/etc/claude-code/`, Windows registry `HKLM`/`HKCU\SOFTWARE\Policies\ClaudeCode`). -> "Environment variables applied to every session and to subprocesses Claude Code spawns from -> it." -> -> code.claude.com/docs/en/settings (the `env` setting's description, first sentence; verified -> 2026-08-10) +So `CLAUDE_CODE_DISABLE_AUTO_MEMORY` can be set as a real OS environment variable **or** inside a +settings file's `env` block, which applies to every session and the subprocesses it spawns. -So `CLAUDE_CODE_DISABLE_AUTO_MEMORY` can be set as a real OS environment variable **or** -inside a settings file's `env` block; the docs bless the `env`-block form explicitly. +- **Pointer**: [Settings precedence](https://code.claude.com/docs/en/settings#settings-precedence), + [Exceptions to managed settings precedence](https://code.claude.com/docs/en/settings#exceptions-to-managed-settings-precedence), + and [`env`](https://code.claude.com/docs/en/settings-reference#env). +- **As of**: 2026-08-10 ## Out of scope for this skill (verified, deliberate) -- **Transcripts / history / shell snapshots / sessions.** Transcripts and shell snapshots are - auto-cleaned at startup by `cleanupPeriodDays` (default 30, minimum 1). The other two are - not: `history.jsonl` persists until deleted, and `sessions/` is cleared per session rather +- **Transcripts / history / shell snapshots / sessions.** This skill treats transcripts and shell + snapshots as cleaned at startup by `cleanupPeriodDays` (default 30, minimum 1), and the other two + as not: `history.jsonl` persists until deleted, and `sessions/` is cleared per session rather than by age. Purging any of them is a different concern. The official per-project wipe is - `claude project purge`, quoted in full below (the deletion plan and flags live in the doc, - not here). - > "**Default**: `30` days, minimum `1`. Claude Code deletes session files and other - > application data older than this period at startup." - > - > code.claude.com/docs/en/settings (the `cleanupPeriodDays` setting's description, first two - > sentences; verified 2026-08-10) - - Read "session files" there as per-session data files, not the `sessions/` directory: the page - links that phrase to claude-directory's "Cleaned up automatically" table, whose rows are the - transcript, `shell-snapshots/`, `debug/`, `tasks/`, `file-history/`, and similar per-session - artifacts. `sessions/` is not a row in that table, and the same page says so directly two - quotes down. - - > "The following paths are not covered by automatic cleanup and persist indefinitely." - > - > code.claude.com/docs/en/claude-directory, heading the table whose first row is - > `history.jsonl` (verified 2026-08-10) - - > "`sessions/` holds one small file per running session, used to detect concurrent sessions - > and crashes. It isn't part of the age-based sweep: Claude Code removes each file when its - > session exits and clears crash leftovers on the next launch." - > - > code.claude.com/docs/en/claude-directory (verified 2026-08-10) - - > "Run `claude project purge` to delete the state Claude Code holds for one project. It - > deletes: - > - > - Transcripts and auto memory under `projects/` - > - Per-session `tasks/`, `debug/`, and `file-history/` entries - > - Matching prompt lines in `history.jsonl` - > - The project's entry in `~/.claude.json`" - > - > code.claude.com/docs/en/claude-directory (verified 2026-08-10) - - code.claude.com/docs/en/claude-directory and code.claude.com/docs/en/cli-reference document - `claude project purge` with no version requirement (verified 2026-08-10). Do not state a - version floor for the command. - - What it leaves alone, from the same page: - - > "The command leaves `shell-snapshots/` and `backups/` alone because those are not - > project-scoped, and warns about them in the plan output." - > - > code.claude.com/docs/en/claude-directory (verified 2026-08-10) - - `sessions/` appears nowhere in the deletion list above. That is this plugin's reading of that - list, not a separate upstream statement. - - It also does not delete unprompted: - - > "The command prints the full deletion plan and asks for confirmation before removing - > anything." - > - > code.claude.com/docs/en/claude-directory (verified 2026-08-10) - - `CLAUDE_CODE_SKIP_PROMPT_HISTORY` skips "writing transcripts and prompt history in any mode" - (code.claude.com/docs/en/claude-directory). It is the true "no session persistence" lever, and - the complement to deleting the files after the fact. Recorded for that contrast; this skill acts - on neither. + `claude project purge`; read its deletion plan and flags on the page. + + We read the age-based sweep as covering per-session data files (transcripts, `shell-snapshots/`, + `debug/`, `tasks/`, `file-history/` and similar), not the `sessions/` directory, which holds one + file per running session, removed when that session exits, with crash leftovers cleared on the + next launch. `history.jsonl` sits among the paths kept until deleted. + + - **Pointer**: [`cleanupPeriodDays`](https://code.claude.com/docs/en/settings-reference#cleanupperioddays), + [Cleaned up automatically](https://code.claude.com/docs/en/claude-directory#cleaned-up-automatically), + [Kept until you delete them](https://code.claude.com/docs/en/claude-directory#kept-until-you-delete-them). + - **As of**: 2026-08-10 + + This skill treats `claude project purge` as deleting, for one project, the transcripts and auto + memory under `projects/`, its per-session `tasks/`, `debug/` and `file-history/` entries, its + prompt lines in `history.jsonl`, and its entry in `~/.claude.json`; as leaving `shell-snapshots/` + and `backups/` alone, with a warning, since they are not project-scoped; and as printing its full + deletion plan and asking for confirmation before removing anything. The docs give the command no + version requirement, so do not state a version floor for it. `sessions/` appears nowhere in the + deletion list. That is this plugin's reading of that list, not a separate upstream statement. + + - **Pointer**: [Clear local data](https://code.claude.com/docs/en/claude-directory#clear-local-data) + and [CLI commands](https://code.claude.com/docs/en/cli-reference#cli-commands). + - **As of**: 2026-08-10 + + We record `CLAUDE_CODE_SKIP_PROMPT_HISTORY` as the true "no session persistence" lever, since it + stops transcripts and prompt history being written, and the complement to deleting the files + after the fact. Recorded for that contrast; this skill acts on neither. Pointer: + [Plaintext storage](https://code.claude.com/docs/en/claude-directory#plaintext-storage). - **Claude Desktop / claude.ai account memory.** That is a server-side account store, not local files, so this skill cannot delete it and only gives direction (see diff --git a/plugins/harness-ops/.claude-plugin/plugin.json b/plugins/harness-ops/.claude-plugin/plugin.json index 3b19cb4cdf..bd0fd3e424 100644 --- a/plugins/harness-ops/.claude-plugin/plugin.json +++ b/plugins/harness-ops/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "harness-ops", - "version": "2.0.1", + "version": "2.0.2", "description": "Claude Code operations toolkit. Fifteen skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used: a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which sheds descriptions lowest-score-first so an unused skill loses the keywords that would let it be matched, from skills genuinely not wanted, from skills the run cannot observe at all; computes whether the listing overflows from documented settings, and withholds every cold verdict the data cannot support rather than reporting absence of data as absence of use), inventory (read-only enumeration of the complete invocable surface: every built-in CLI command with aliases and hidden/gated status, every bundled skill, every built-in subagent and tool, every built-in plugin with its components, and every component of every installed plugin across all marketplaces; reads the shipped binary because upstream publishes no built-in command list, and carries an integrity verdict so a drifted build reports counts as floors rather than silently short totals), audit-install-state (read-only audit of the machine-scope ~/.claude installation directory and ~/.claude.json: full inventory split into an authored surface and rolled-up bulk trees, product-managed retention vs genuinely unmanaged state, filename-scheme resolution before any process-liveness check, and deliberate/mid-experiment detection; reports, never deletes), audit-performance (read-only slowness-diagnostic capture run at the moment the machine or a session feels slow: CLI version, retention-sweep health including the unparsable-settings pause, which warns in /status, a timed census walk of the install tree as a sweep-cost proxy, active-session and plugin-fleet counts, a process census, and the fan-out layer, which covers a load-labeled no-op spawn baseline, every hook that will fire bucketed per-tool-call versus per-turn with its invocation shape, the configured statusline, subagent concurrency and spawn-depth ceilings against documented defaults, whether running sessions predate the settings file they are judged by, and orphan attribution by parent liveness rather than age, plus on Windows a kernel-object census (Token objects against uptime, paged pool) that names a host-level leak beneath all four suspects; read against a bundled known-performance-issues reference that also records the causes tested and cleared; separates the four documented suspects of accumulated state, version regression, component bloat, and per-spawn fan-out cost, and routes remediation out; reports, never mutates, and never executes a discovered hook or statusline command), audit-native-overlap (map native Claude Code surfaces, namely built-in CLI commands, bundled skills, plugin-backed built-ins, and session-provided skills, against the current repo's plugin skills and agents, so a custom component never silently duplicates what Claude Code itself ships; bare invocation is a read-only overlap report carrying the extraction's integrity floors and a shared-listing-budget exposure section, verdicts are human-gated in a committed store rendered into a generated registry whose every row carries an observable recheck trigger, and only an explicit apply step bakes presence-gated native references into descriptions and Boundary sections), observability (read locally captured telemetry from the OTEL store, the collector, the per-session hook event log and hook-event JSONL, and ccusage, with trend reports, a per-session report of what fired, what was blocked and the event timeline, and store pruning), known-issues (search known Claude product GitHub bugs, check service health, maintain a persistent tracked-issue registry), changelog (ingest Claude Code changelog entries and turn them into decisions: apply executes those in scope one PR per owner plugin and hands larger ones off as work items, then re-extract the native surface and file its drift as work items), prerequisites (read-only table of external binaries declared by enabled plugins; never installs), check (read-only check that node and jq resolve for the harness-ops hooks; never installs), machine-profile (discover this machine's facts and per-tree identity domains, store them as a re-runnable profile with the observation behind every value, and diff the stored profile against the host now; read-only unless the operator confirms a write, never installs and never reapplies a stored value on its own), plugins (bring a machine's plugin fleet current on demand: marketplace refresh, effective-scope updates including in-repo project/local installs, new-plugin install per policy, scope-divergence detection and explicit convergence), morning-brief (read-only gh-based operator morning view: queue-label counts, merge-ready PRs, parked decisions with their RECOMMENDED lines, and loop-lane telemetry freshness), lanes (start/restart/stop/status loop lanes as named background Claude Code sessions seeded from canonical prompt files, with per-lane model/effort, a repo-pull + marketplace-refresh launch step, and a consume-restarts action, an OS-schedulable reader that relaunches stopped lanes whose telemetry carries a restart_request), and a re-runnable setup action that settles where the known-issues registry, the skill-usage log and the hook log root live, places the root's self-ignoring guard, and detects retired conventions. Plus an opt-in, default-off per-session hook event log (one JSON line per hook event on every event the generated registry marks observable, written to /sessions/.jsonl, with SessionEnd retention by session count or age and an optional detached pre-prune command), a family of eight advisory *-audit hooks (API errors, config changes, instruction loads, permission denials, pre-compaction, skill usage, tool failures, and unsurfaced hook failures. The last also warns the user via systemMessage, since a hook that fails to launch enforces nothing and Claude Code surfaces the failure to nobody) that emit the shared hook-telemetry envelope, and a reference sink that routes envelopes under the same root: per session when the envelope carries a session id, else into the shared hook-events.jsonl the observability skill reads.", "author": { "name": "Melodic Software", diff --git a/plugins/harness-ops/CHANGELOG.md b/plugins/harness-ops/CHANGELOG.md index 35fc84b56b..e2464c6532 100644 --- a/plugins/harness-ops/CHANGELOG.md +++ b/plugins/harness-ops/CHANGELOG.md @@ -3,6 +3,27 @@ All notable changes to the `harness-ops` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [2.0.2] - 2026-10-02 + +### Changed + +- **The `audit-native-overlap` bake step writes the links-only record shape.** A reference it bakes + into a skill holds our decision in our words, a pointer to the exact upstream section, the as-of + date and the recheck trigger, with no upstream text. A table form uses the header + `| Decision | Pointer | As of | Recheck when |` in place of `| Claim | Basis | Recheck trigger | + Verified |`, and the skill's own two upstream dependencies are restated in that form. +- **The `known-issues`, `observability` and `plugins` records follow the same shape.** The model + fallback and quality-tracker notes, the hook-latency event record and the plugin scope + semantics each state our decision and point at the docs section or at our own probe, instead of + restating the page. +- **The `known-issues` model-fallback note was re-read after its trigger fired.** It now covers a + refusal when a flagged category has no fallback target, with a 2026-10-01 as-of date. +- `observability` points its latency record at the documented `hook_execution_complete` event and + keeps only the string-typed duration as our probe. `plugins` scope-semantics treats `--force` as + the answer to the MCP or LSP reload warning only, and records that the reference now offers a + fixed-options list for `userConfig` and when we adopt it. Two cloud-sessions links in `audit- + native-overlap` follow the docs site's new heading ids. + ## [2.0.1] - 2026-10-02 ### Fixed diff --git a/plugins/harness-ops/skills/audit-native-overlap/SKILL.md b/plugins/harness-ops/skills/audit-native-overlap/SKILL.md index 9994175f1c..ac31f2f7e3 100644 --- a/plugins/harness-ops/skills/audit-native-overlap/SKILL.md +++ b/plugins/harness-ops/skills/audit-native-overlap/SKILL.md @@ -297,8 +297,10 @@ Two baked surfaces exist, and they are gated differently because they cost diffe **The Boundary section lands with the row.** A store row whose verdict is not `defer` and whose observation is extraction-evidence is written together with a `## Boundary` section in the component's body, in the same change: the surfaces by provenance class, the routing split, and -the mutation gate, with the four-part detail (basis, as-of, recheck trigger, evidence) in a -reference file inside the same skill that the section links. A body loads only on invocation, so +the mutation gate, with the records behind them in a reference file inside the same skill that +the section links. Each record holds our decision in our words, a pointer to the exact upstream +section, the as-of date and the recheck trigger, and no upstream text; a table form uses the +header `| Decision | Pointer | As of | Recheck when |`. A body loads only on invocation, so the section spends no listing budget and moves no routing; it is what makes the verdict real for the model, and a row without it fails the self-check. Boundary-only baking may cover several plugins in one change. @@ -375,30 +377,29 @@ Any claim about what Claude Code itself ships must come from the raw markdown en returns a small model's answer *about* the page, so absence from that answer is not evidence of absence. A `200` is also not proof you got the page you asked for: retired slugs are silently aliased, so confirm the slug against `https://code.claude.com/docs/llms.txt` and read the body's -own first heading before quoting it. +own first heading before citing it. -Two upstream facts this skill depends on, each with the trigger that obliges re-deriving it: +Two upstream dependencies of this skill, each with the trigger that obliges re-deriving it: -| Claim | Basis | Recheck trigger | Verified | +| Decision | Pointer | As of | Recheck when | |---|---|---|---| -| Descriptions load into context by default, truncated at 1,536 chars per entry, listing capped at 1% of the context window, with name-only degradation on overflow. The docs state that degradation goes least-invoked-first; the shipped binary instead ranks by a decay-weighted score and grants first-fit, so use the mechanism recorded in [`audit-skill-visibility/reference/listing-scorer.md`](../audit-skill-visibility/reference/listing-scorer.md), not the documented order | `docs/en/skills.md` (Frontmatter reference; Troubleshooting), `docs/en/settings-reference.md`, plus the binary for the order | Either default moves, or the binary's scorer or grant loop diverges from that reference | 2026-09-01 | -| Native availability varies on settings/env, plan, platform/provider, and host surface, so no static availability claim holds | `docs/en/settings-reference.md`, `docs/en/env-vars.md`, `docs/en/commands.md`, `docs/en/cloud-environments.md` | A release or docs change adds, removes, or renames a gating axis | 2026-08-23 | +| We treat the description as the routing surface and measure a baked description against a 1,536-character per-entry cap and a listing budget of 1% of the context window. For which entries overflow drops, we use the order our own binary reading records in [`audit-skill-visibility/reference/listing-scorer.md`](../audit-skill-visibility/reference/listing-scorer.md); the docs page and that reading disagree on the drop order | [Frontmatter reference](https://code.claude.com/docs/en/skills#frontmatter-reference), [Skill descriptions are cut short](https://code.claude.com/docs/en/skills#skill-descriptions-are-cut-short), [`skillListingMaxDescChars`](https://code.claude.com/docs/en/settings-reference#skilllistingmaxdescchars), [`skillListingBudgetFraction`](https://code.claude.com/docs/en/settings-reference#skilllistingbudgetfraction); the drop order is our binary reading | 2026-09-01 | Either default moves, or the binary's scorer or grant loop diverges from that reference | +| We make no static availability claim for a native surface, because we treat settings and environment, plan, platform or provider, and host surface as each able to remove one | [`disableBundledSkills`](https://code.claude.com/docs/en/settings-reference#disablebundledskills), [environment variables](https://code.claude.com/docs/en/env-vars), [Commands](https://code.claude.com/docs/en/commands), [What's available in cloud sessions](https://code.claude.com/docs/en/cloud-environments#what%E2%80%99s-available-in-cloud-sessions) | 2026-08-23 | A release or docs change adds, removes, or renames a gating axis | ## Gotchas - **A plugin skill never shadows a native one.** Ours are namespaced, so both resolve and the model chooses. That is why the routing lives in descriptions rather than in a name. - **`plugin_backed` is its own lane.** `security-review` is reported there, not under - `builtin_commands`. Read the wrong key and the row looks absent. Verified 2026-09-29 against - Claude Code 2.1.284, by reading the `plugin_backed` key of an `inventory.py --binary-only` - extraction on this machine, which holds `security-review` and nothing else. Recheck when the - extractor's provenance lanes change or a release note moves a bundled surface between them. -- **A bundled skill can carry aliases.** `code-review` answers to `review`; treating an alias as a - separate surface produces a duplicate row for one capability. Basis: - gives `/code-review` the line "Alias: `/review`". - Verified 2026-09-29 against Claude Code 2.1.284 (the extraction lists `review` as the alias) and - that page as fetched that day. Recheck when - the commands page drops the alias line or a release note renames a bundled skill. + `builtin_commands`. Read the wrong key and the row looks absent. Pointer: our probe, the + `plugin_backed` key of an `inventory.py --binary-only` extraction on this machine, which held + `security-review` and nothing else. As of: 2026-09-29, Claude Code 2.1.284. Recheck trigger: + the extractor's provenance lanes change or a release note moves a bundled surface between them. +- **A bundled skill can carry aliases.** We treat `review` as an alias of `code-review`, never a + separate surface, since a separate row would duplicate one capability. Pointer: + [All commands](https://code.claude.com/docs/en/commands#all-commands), the `/code-review` row, + and our extraction, which lists `review` as the alias. As of: 2026-09-29, Claude Code 2.1.284. + Recheck trigger: the commands page drops the alias or a release note renames a bundled skill. - **Absent from the binary is not absent from the product.** Session-provided skills exist only in a live roster. "Not in the extraction" is a statement about the extraction. - **A verdict is not permanent.** The trigger is the load-bearing part of the row; a date alone diff --git a/plugins/harness-ops/skills/known-issues/context/action-quality.md b/plugins/harness-ops/skills/known-issues/context/action-quality.md index dfc69f03ac..d5bfc969e7 100644 --- a/plugins/harness-ops/skills/known-issues/context/action-quality.md +++ b/plugins/harness-ops/skills/known-issues/context/action-quality.md @@ -6,7 +6,7 @@ Check current Claude model quality and service health from multiple sources. Use ## Sources (checked in order) -**Source 1: [Marginlab Performance Tracker](https://marginlab.ai/trackers/claude-code/).** Independent daily benchmarks on SWE-Bench-Pro. Updated daily, 50 evals/day. Statistical significance testing (p < 0.05). +**Source 1: [Marginlab Performance Tracker](https://marginlab.ai/trackers/claude-code/).** Our model-performance signal: an independent daily benchmark tracker. Its benchmark, sample size and significance method are stated on the page; read them there. Fetch via WebFetch or curl and extract: @@ -14,7 +14,7 @@ Fetch via WebFetch or curl and extract: - Today's pass rate vs 7-day and 30-day averages - Statistical significance of any delta -**Source 2: [status.claude.com](https://status.claude.com/).** Official Anthropic status page. Covers claude.ai, API, Claude Code, platform. +**Source 2: [status.claude.com](https://status.claude.com/).** Our service-health signal: the official status page, read per component. Fetch and extract: @@ -57,20 +57,21 @@ gh search issues "degraded OR degradation OR quality OR nerfed OR slower" --repo ## Before blaming the model: check for a flag fallback -A sudden quality change mid-session can be a model switch, not a regression. When a safety -classifier flags a request (most often cybersecurity or biology content), Claude Code re-runs it -on an older fallback model, shows a notice naming that model in the transcript, and the session -stays on it. Check the transcript for that notice first. To recover: +A sudden quality change mid-session can be a model switch, not a regression. Before blaming the +model, check the transcript for a notice that the session moved to a fallback model after a +flagged message, or for a request that ended in a refusal because the flagged category had no +fallback for the model in use. To recover, we use: -- `/model` switches back to the original model. +- `/model` to switch back to the original model. - `/config` > **Switch models when a message is flagged** off (or `switchModelsOnFlag: false`) - asks each time instead of switching. -- If the flag looks wrong, report it with `/feedback` (vendor-reported advice, from the - [Opus 5.5 usage guide](https://claude.dev/blog/getting-the-most-out-of-opus-5-5/)). - -Verification: claim, the fallback behavior and the two recovery settings above; basis, the -[model-config page, "Automatic model fallback"](https://code.claude.com/docs/en/model-config); -as of 2026-09-23; recheck when that section changes or a new model gains or loses a fallback. + to be asked each time instead of switched. +- `/feedback` to report a flag that looks wrong. + +Pointer: for which models fall back, to what, and the recovery settings, see +[Automatic model fallback](https://code.claude.com/docs/en/model-config#automatic-model-fallback) +(correlate with the [Opus 5.5 usage guide](https://claude.dev/blog/getting-the-most-out-of-opus-5-5/)). +As of: 2026-10-01. Recheck trigger: that section changes, or a new model gains or loses a +fallback. ## Fragility note diff --git a/plugins/harness-ops/skills/observability/SKILL.md b/plugins/harness-ops/skills/observability/SKILL.md index b6ae99cfd0..804cfd5d76 100644 --- a/plugins/harness-ops/skills/observability/SKILL.md +++ b/plugins/harness-ops/skills/observability/SKILL.md @@ -135,10 +135,11 @@ remains the durable record). `latency` telemetry record: -- **Claim**: Claude Code emits one `hook_execution_complete` log event per hook event firing. It carries `hook_event` (the event name), `total_duration_ms` (wall time for every hook that firing ran, with `num_hooks` counting them) and `session.id`. `claude.lane` is not a Claude Code attribute; it is a custom attribute set through `OTEL_RESOURCE_ATTRIBUTES`, which Claude Code includes on all events, and rows without it report as lane `unknown`. -- **Basis**: the event is not documented. , fetched 2026-09-24, lists no hook execution event (only the `claude_code.hook` trace span); it does document `OTEL_RESOURCE_ATTRIBUTES` custom attributes "included in all metrics and events". The event shape is observed in the live OTEL store on 2026-09-24 under Claude Code 2.1.281: `total_duration_ms` arrives as a `stringValue` (read with an `intValue` fallback in case that changes). -- **As of**: 2026-09-24, Claude Code 2.1.281. -- **Recheck trigger**: the monitoring page documents a hook execution event, a release note names `hook_execution_complete` or its attributes, or `latency` exits 2 with no rows on a store that has recent sessions. +`latency` reads one `hook_execution_complete` log event per hook event firing, using its `hook_event`, `total_duration_ms`, `num_hooks` and `session.id`. Our probe of the live OTEL store under Claude Code 2.1.281 observed `total_duration_ms` arriving as a `stringValue`, so the script reads it with an `intValue` fallback in case that changes. `claude.lane` is our own custom attribute, set through `OTEL_RESOURCE_ATTRIBUTES`; rows without it report as lane `unknown`. + +- **Pointer**: for the event and its attributes, see ; the string-typed duration is our probe of the live OTEL store on 2026-09-24; for custom resource attributes, see . +- **As of**: 2026-10-01 for the docs sections; 2026-09-24, Claude Code 2.1.281, for the probe +- **Recheck trigger**: that section renames the event or an attribute `latency` reads, a release note changes the duration's type, or `latency` exits 2 with no rows on a store that has recent sessions. Action invocation: `/harness-ops:observability clean [flags]`. @@ -208,8 +209,8 @@ retention in effect" section, the six probe lines verbatim. One native surface also answers "where did my tokens go", and the two get conflated whenever a session feels expensive: -- **`explain-usage` (bundled skill)**: explains where the current session's tokens went, with one - simple chart in plain language. +- **`explain-usage` (bundled skill)**: we route a plain-language breakdown of the current + session's tokens to it. The model and the person can both invoke it where it resolves. - **This skill (marketplace plugin).** Reads locally captured telemetry (the OTEL store, the hook event log, ccusage) across sessions: trends, cost, hook latency, which hooks fired, and a @@ -224,8 +225,8 @@ is read-only apart from `--write` reports and the explicit `clean` action. Never the other's behalf. **Availability is never assumed.** The skill is gated, and bundled skills vary by settings, plan, -and host; this section states what to do when it resolves, never that it is present. The four-part -records live in [reference/native-explain-usage.md](reference/native-explain-usage.md). +and host; this section states what to do when it resolves, never that it is present. The records +behind it live in [reference/native-explain-usage.md](reference/native-explain-usage.md). ## Spoke paths @@ -233,19 +234,20 @@ The `context/` files write this skill's directory as ``, which is `${ Put that path in place of the placeholder before running a command or writing it into a brief. Those files arrive through the Read tool as plain bytes, so a `${…}` token in them would reach the Bash tool unsubstituted, and the Bash tool's environment has no `CLAUDE_SKILL_DIR` to expand it from. -Basis: the plugins reference, -, verified -2026-09-30; recheck when that table adds supporting files to where a `${…}` reference resolves. +Pointer: for where a `${…}` reference resolves, see +. As of: +2026-09-30. Recheck trigger: that table adds supporting files to where a `${…}` reference +resolves. ## Gotchas - Empty stores are normal on first run. Degrade gracefully -- **No `${user_config.*}` inside a pre-compute command.** A `${user_config.*}` value renders in plain skill content only; shell-executing content rejects it because the shell would re-parse whatever the value holds, and a placeholder left unrendered on a shell line is a bash `bad substitution` that aborts the whole invocation, since one failed pre-compute line aborts every line. The options render as plain content above the probe lines; the probe lines pass none and print no option tier (`--observed`), so a manifest default never appears where an effective value belongs; the model hands the rendered values to the probe through its own Bash call, the one place the options render. Basis: "Fields that run in a shell reject `${user_config.*}`" under "User configuration" at , and the substitution list under "Dynamic context injection" at . Verified 2026-09-09 against Claude Code 2.1.263 and both pages as fetched that day; recheck when either page names `user_config` for pre-compute lines +- **No `${user_config.*}` inside a pre-compute command.** We render `${user_config.*}` values only in plain skill content, never on a shell line: a placeholder left unrendered on a shell line is a bash `bad substitution` that aborts the whole invocation, since one failed pre-compute line aborts every line. The options render as plain content above the probe lines; the probe lines pass none and print no option tier (`--observed`), so a manifest default never appears where an effective value belongs; the model hands the rendered values to the probe through its own Bash call, the one place the options render. Pointer: for where `${user_config.*}` renders, see ; for the pre-compute substitutions, see . As of: 2026-09-09, Claude Code 2.1.263. Recheck trigger: either page names `user_config` for pre-compute lines - **The pipeline line names two tiers.** `envelope:` counts rows the telemetry sink wrote for the audit hooks, across `sessions/*.jsonl` (rows marked `source: "envelope"`) and the whole shared `hook-events.jsonl` plus its rotated `hook-events.jsonl.1` (the legacy shape for a hook payload with no session id); those follow the per-hook audit toggles and never the event-log switch. `event log:` is the switch. `event log: off` beside a populated root is the normal state, not a contradiction -- **`session_id` joins only per-session files**. Rows in `sessions/.jsonl` carry the id; rows in the shared `hook-events.jsonl` do not, and are never attributed to a session (say "legacy rows, shared file, time proximity only"). OTEL rows join on `session_id` as before; `cwd` + `branch` + time proximity is the fallback for a producer that sends none. Hook input carries `session_id` on every event (the common input fields at ), so a row without one comes from a producer that dropped it, never from the harness. Verified 2026-09-06 against Claude Code 2.1.263 and that page as fetched that day; recheck when the common input fields drop `session_id` +- **`session_id` joins only per-session files**. Rows in `sessions/.jsonl` carry the id; rows in the shared `hook-events.jsonl` do not, and are never attributed to a session (say "legacy rows, shared file, time proximity only"). OTEL rows join on `session_id` as before; `cwd` + `branch` + time proximity is the fallback for a producer that sends none. We treat `session_id` as present on every hook input, so a row without one comes from a producer that dropped it, never from the harness. Pointer: . As of: 2026-09-06, Claude Code 2.1.263. Recheck trigger: the common input fields drop `session_id` - **Per-hook duration per session covers producers that emit `data.session_id`** (the nine harness-ops audit hooks). Other hooks appear in the whole-root tables only - **Hooks run in parallel**. Row order within one second is write order, not fire order; group by `prompt_id` or `tool_use_id`, not by adjacency -- **Stop is a per-turn event, not a session boundary**. It fires "When Claude finishes responding", so a session with many turns emits many Stop rows; `SessionEnd` is the row that fires "When a session terminates". Aggregate per session on `SessionEnd`, never on Stop. Basis: the hook lifecycle table at . Verified 2026-09-06 against Claude Code 2.1.263 and that page as fetched that day. Recheck when the lifecycle table changes either row, or a release note names Stop or `SessionEnd` +- **Stop is a per-turn event, not a session boundary**. We treat Stop as firing once per turn, so a session with many turns emits many Stop rows, and `SessionEnd` as the session boundary. Aggregate per session on `SessionEnd`, never on Stop. Pointer: . As of: 2026-09-06, Claude Code 2.1.263. Recheck trigger: the lifecycle table changes either row, or a release note names Stop or `SessionEnd` - **`cc_spans` / `cc_traces`**. Views skip bind until `cc-traces.json` has content ## What this skill does NOT do diff --git a/plugins/harness-ops/skills/plugins/SKILL.md b/plugins/harness-ops/skills/plugins/SKILL.md index 68682c1fa0..f5073ff971 100644 --- a/plugins/harness-ops/skills/plugins/SKILL.md +++ b/plugins/harness-ops/skills/plugins/SKILL.md @@ -22,11 +22,11 @@ you're standing in with its own project/local-scope installs, are the latest pub everything the marketplace offers, and surfaces any state where something older or unintended is what really runs. -Distinct from what Claude Code's own background `autoUpdate` does (see -[context/scope-semantics.md](context/scope-semantics.md)): `autoUpdate` silently refreshes marketplace -data and bumps already-installed plugins post-startup. It never installs a new catalog plugin, never -checks `enabledPlugins` completeness, never detects or reports scope divergence, and only runs once -per session start on its own schedule, not on demand. This skill covers exactly that gap. +Distinct from Claude Code's own background `autoUpdate`, which we treat as a background refresh of +already-installed plugins, never a substitute for this skill +([context/scope-semantics.md](context/scope-semantics.md) holds that reading and its pointer). This +skill covers what we do not rely on `autoUpdate` for: installing new catalog plugins, checking +`enabledPlugins` completeness, detecting and reporting scope divergence, and running on demand. Distinct from `harness-config`'s `audit` skill's plugin-drift check: that check compares a project's committed `enabledPlugins` against a marketplace's *upstream* `marketplace.json` (orphan/new/rename @@ -76,9 +76,10 @@ Each Description names the territory an action covers, never its algorithm, the ordering, and their failure handling live only in the linked file. One block below is a deliberate exception to that index-only rule and has to live in the hub -rather than in a spoke: the **`install_new` render**, because Claude Code substitutes -`${user_config.*}` when it renders the *skill*; a spoke opened later as a file read is plain bytes, -so the same token in a spoke would arrive as a literal placeholder with no error to warn anyone. +rather than in a spoke: the **`install_new` render**, because we rely on `${user_config.*}` being +substituted only when Claude Code renders the *skill*; a spoke opened later as a file read is plain +bytes, so the same token in a spoke would arrive as a literal placeholder with no error to warn +anyone. See [context/gotchas.md](context/gotchas.md). The Report section below is a pointer: the report itself is rendered by the script. @@ -98,8 +99,8 @@ it prints the JSON digest, then the report. The step sequence, and the reason be stay in [context/sync.md](context/sync.md); the script is bound to that file. 1. **Run it.** Substitute the marketplace target, the policy, and the flags into this command. The - journal root is written here because `${CLAUDE_PLUGIN_DATA}` resolves in skill content and - **not** in a `context/*.md` spoke, which is read raw: + journal root is written here because we rely on `${CLAUDE_PLUGIN_DATA}` resolving in skill + content and **not** in a `context/*.md` spoke, which is read raw (see "Spoke paths" below): ```bash "${CLAUDE_PLUGIN_ROOT}"/skills/plugins/scripts/sync-run.sh \ @@ -238,25 +239,22 @@ carries, and do not add rows: a row the render omitted is a row whose digest fie marketplace in `all` mode the whole block repeats, because Steps 2 through 5 run once per marketplace; the trailing `Run journal:` and `Timing:` lines cover the invocation. -**The one line the model owns is the reload guidance**, appended after the render and stated as -the docs' own two-step rather than as a prediction about which case will trigger it: recommend bare -`/reload-plugins`; if it warns that the reload would re-read the conversation, rerun it as -`/reload-plugins --force`. The general condition `--force` exists for is prompt-cache -invalidation; a plugin shipping an MCP server whose tools aren't deferred is the common cause, not -the only one, so do not present it as the sole trigger and do not tell the user `--force` would be -wrong when the bare command has already warned them. Never recommend `--force` pre-emptively -alongside every reload: it opts into a real token cost the bare command declines to pay on its own -(see [context/scope-semantics.md](context/scope-semantics.md)). Monitors are already covered: the -render's `Action needed` names each updated plugin whose installed build declares one and -attributes "monitors require a session restart" to the plugins reference, so the reload line does -not repeat it. After the line, answer follow-up questions from the digest and the run directory it -names; load [context/sync.md](context/sync.md) when a question is about why a step behaved the -way it did. - -(A plugin updated mid-session keeps resolving to the previous version's path, which is what the -self-update note reports. `plugins-reference`, re-fetched 2026-09-05 and unchanged; behavior -observed on Claude Code 2.1.240 and not re-run on 2.1.261, because it needs an interactive session. -See [context/gotchas.md](context/gotchas.md).) +**The one line the model owns is the reload guidance**, appended after the render as a two-step +rather than as a prediction about which case will trigger it: recommend bare `/reload-plugins`, +and `/reload-plugins --force` only when the bare command warns. Do not present any one cause as the +sole trigger of that warning, and do not tell the user `--force` would be wrong when the bare +command has already warned them. Never recommend `--force` pre-emptively alongside every reload: we +treat it as opting into a real token cost the bare command declines to pay on its own. +[context/scope-semantics.md](context/scope-semantics.md) "`/reload-plugins`: bare by default, +`--force` for the MCP-cache-invalidation case" holds the pointers and dates. Monitors are already +covered: the render's `Action needed` names each updated plugin whose installed build declares one, +so the reload line does not repeat it. After the line, answer follow-up questions from the digest +and the run directory it names; load [context/sync.md](context/sync.md) when a question is about +why a step behaved the way it did. + +(We observed a plugin updated mid-session keep resolving to the previous version's path, which is +what the self-update note reports: Claude Code 2.1.240, not re-run on 2.1.261 because it needs an +interactive session. The record is in [context/gotchas.md](context/gotchas.md).) ## Stale project records and cache content @@ -270,8 +268,12 @@ either section reads as it does. ## userConfig: `install_new` -Controls new-catalog-plugin install policy during `sync`. Ships as a plain `string` (the manifest -schema has no `enum` type. Verified against the published schema), default `"ask"`: +Controls new-catalog-plugin install policy during `sync`. Ships as a plain `string`, default +`"ask"`, with its three values validated by this skill rather than by the manifest +([context/scope-semantics.md](context/scope-semantics.md) holds the option-schema record; for +option types and fixed options, see +[User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration) and +[Limit a field to fixed options](https://code.claude.com/docs/en/plugins-reference#limit-a-field-to-fixed-options)): - `ask` (default). Offer every not-yet-installed catalog plugin in one batched multi-select prompt - `all`. Install every not-yet-installed catalog plugin automatically @@ -293,28 +295,28 @@ proceed unattended:** resolve the total install gap first (one `audit all`, or ` configured value as written with the operator's own marketplace in mind, and downgrade to `ask` when no human is present to receive the count. -**Configured value: `${user_config.install_new}`**. Claude Code text-substitutes a `userConfig` -value into this skill's content before the model sees the rendered skill, but **only when the key is -explicitly set** in user settings (`~/.claude/settings.json`), `--settings`, or managed settings: precedence managed → `--settings` → user. It is **not** "some `pluginConfigs` scope": a project's -`.claude/settings.json` or `.claude/settings.local.json` entry is ignored, and setting `install_new` -there does nothing at all. Declaring the option in `plugin.json` alone does not make its value -readable here either. See [context/scope-semantics.md](context/scope-semantics.md) for the read path -and why it differs from `enabledPlugins`, which this same skill reads from project and local scope. - -The manifest's `"default": "ask"` is **not** substituted for an unset key: the render leaves the -placeholder token unchanged while a sibling key set in `~/.claude/settings.json` and -`${CLAUDE_PLUGIN_ROOT}` both substitute in the same render. That is a probed claim, not a documented -one, and it stays pointed at its dated record: the verification (as-of date, CLI version, basis, -recheck trigger), the `pluginConfigs` payload shape, and the probe recipe are in +**Configured value: `${user_config.install_new}`**. We treat this line as rendering the configured +word into the skill content before the model sees it, but **only when the key is explicitly set** +in one of the three sources this skill reads `pluginConfigs` from: user settings +(`~/.claude/settings.json`), `--settings`, or managed settings, with precedence managed → +`--settings` → user. A project's `.claude/settings.json` or `.claude/settings.local.json` entry does +nothing at all, and declaring the option in `plugin.json` alone does not make its value readable +here either. See [context/scope-semantics.md](context/scope-semantics.md) for the read path, its +pointer, and why it differs from `enabledPlugins`, which this same skill reads from project and +local scope. + +The manifest's `"default": "ask"` is **not** substituted for an unset key: our probe saw the render +leave the placeholder token unchanged while a sibling key set in `~/.claude/settings.json` and +`${CLAUDE_PLUGIN_ROOT}` both substituted in the same render. That is a probed claim, not a +documented one, and it stays pointed at its dated record: the probe, its as-of date and CLI version, +the docs pointer for the `default` field and the substitution surfaces, the recheck trigger, the +`pluginConfigs` payload shape, and the probe recipe are in [context/scope-semantics.md](context/scope-semantics.md) "`userConfig`: an unset key renders the literal placeholder". Read it when the rendered value looks wrong or before re-running the probe. -The `plugins-reference` page describes `default` as "Value used when the user provides nothing" and -states the substitution surface as "Each value is available for substitution as `${user_config.KEY}` -in MCP and LSP server configs and hook commands. Non-sensitive values can also be substituted in -skill and agent content." Substitution into content happens in what Claude Code renders, and never in -a file a spoke read returns, which is why the **Configured value** line lives here and cannot move to -a spoke. So for the common default-config user, with no `pluginConfigs` set anywhere, the -**Configured value** line above still shows that literal placeholder token, not `ask`. +Substitution into content happens in what Claude Code renders, and never in a file a spoke read +returns, which is why the **Configured value** line lives here and cannot move to a spoke. So for +the common default-config user, with no `pluginConfigs` set anywhere, the **Configured value** line +above still shows that literal placeholder token, not `ask`. Read that literal placeholder token as the **expected unset state → use the default `ask`**, and do NOT report it as an invalid value. Only a rendered value that is a real word other than @@ -325,12 +327,14 @@ default when that render is still the placeholder token, not on the option's nam ## Spoke paths The `context/` files write this skill's directory as ``, which is `${CLAUDE_SKILL_DIR}`. -Put that path in place of the placeholder before running a command or writing it into a brief. Those -files arrive through the Read tool as plain bytes, so a `${…}` token in them would reach the Bash -tool unsubstituted, and the Bash tool's environment has no `CLAUDE_SKILL_DIR` to expand it from. -Basis: the plugins reference, -, verified -2026-09-30; recheck when that table adds supporting files to where a `${…}` reference resolves. +Put that path in place of the placeholder before running a command or writing it into a brief. We +never put a `${…}` token in those files: we do not rely on one being substituted in a file read +through the Read tool, or on the Bash tool's environment carrying `CLAUDE_SKILL_DIR`. + +- **Pointer**: for where each `${…}` reference resolves, see + [Where each variable resolves](https://code.claude.com/docs/en/plugins-reference#where-each-variable-resolves). +- **As of**: 2026-09-30 +- **Recheck trigger**: that table adds supporting files to where a `${…}` reference resolves. ## Next diff --git a/plugins/harness-ops/skills/plugins/context/scope-semantics.md b/plugins/harness-ops/skills/plugins/context/scope-semantics.md index 74c2babb52..f3f23be38d 100644 --- a/plugins/harness-ops/skills/plugins/context/scope-semantics.md +++ b/plugins/harness-ops/skills/plugins/context/scope-semantics.md @@ -19,8 +19,9 @@ - [`marketplace remove` leaves the cache tree, marked for the orphan sweep](#marketplace-remove-leaves-the-cache-tree-marked-for-the-orphan-sweep) - [`autoUpdate` is a background complement, not a substitute](#autoupdate-is-a-background-complement-not-a-substitute) -Every claim below was verified against a fetched official-docs page or an empirical test on a real -machine, not assumed from training data. Last re-verified 2026-09-05 against +Every record below is our decision, resting on a pointer to an official-docs section or on our own +probe on a real machine, never on training data, and stores no upstream text. Last re-verified +2026-09-05 against [plugins-reference](https://code.claude.com/docs/en/plugins-reference), [discover-plugins](https://code.claude.com/docs/en/discover-plugins), [plugin-marketplaces](https://code.claude.com/docs/en/plugin-marketplaces), and the published @@ -164,8 +165,8 @@ and one tracked `.claude/settings.json`, hold independent records and pin indepe Removing the directory a project/local install was made from leaves the install record in place, still naming the path. **Re-verified on Claude Code 2.1.261**: `claude plugin --help` lists no verb that -removes an install record by path, and `claude plugin prune --help` reports "Remove auto-installed -dependencies that are no longer needed", a *dependency* axis, whose own `-s project` has the same +removes an install record by path, and `claude plugin prune --help` describes a cleanup of unneeded +dependencies, a *dependency* axis, whose own `-s project` has the same no-path-flag behavior documented above, so it acts on the cwd and cannot reach a record belonging to a directory that is gone. @@ -189,24 +190,22 @@ do. ## Where project-scope records come from, and why the skill cannot reap them The section above says the records cannot be reaped. This one says where they come from. Every claim -below is doc-sourced or verified by a probe, each with its source or CLI version named. **Recheck +below rests on a pointer or on a probe, each with its source or CLI version named. **Recheck trigger:** any change to how a repo's committed `enabledPlugins` block is applied at session start, or any `claude plugin` release note adding a verb that removes an install record by path. -**A repo's committed `.claude/settings.json` `enabledPlugins` block is the documented cloud install -mechanism.** Per -[cloud-environments](https://code.claude.com/docs/en/cloud-environments) ("What carries over from -your setup", fetched 2026-09-05), plugins declared in that committed block are "Installed at session -start from the marketplace you declared." Plugins enabled only in a user's own settings do not carry -over to a cloud session at all. So the block exists to make a team's plugin set reproducible -somewhere the user's `~/.claude` is not. - -**Locally, session start writes the records.** Per -[discover-plugins](https://code.claude.com/docs/en/discover-plugins) ("Configure team marketplaces", -fetched 2026-09-05), as of v2.1.195 a plugin that only project settings enable, coming from an -external source, "doesn't load until the team member installs it." That sentence covers the case -where the user has never installed the plugin. When the user already holds it at user scope, the -session start does the install itself. **Verified 2026-09-06 on Claude Code 2.1.263**: a scratch git +**We treat a repo's committed `.claude/settings.json` `enabledPlugins` block as the cloud install +mechanism**, and a plugin enabled only in a user's own settings as absent from a cloud session. So +the block exists to make a team's plugin set reproducible somewhere the user's `~/.claude` is not. +Pointer: for what a cloud session carries over, see +[What carries over from your setup](https://code.claude.com/docs/en/cloud-environments#what-carries-over-from-your-setup). +As of: 2026-09-05. + +**Locally, session start writes the records.** A plugin that only project settings enable and that +the user has never installed is a separate case, for which see +[Enabled in project settings but not installed](https://code.claude.com/docs/en/plugins/loading#enabled-in-project-settings-but-not-installed) +(as of 2026-09-05). When the user already holds the plugin at user scope, the session start does the +install itself. **Verified 2026-09-06 on Claude Code 2.1.263**: a scratch git repo under the temp directory with a committed `.claude/settings.json` declaring `extraKnownMarketplaces` for an already-registered marketplace and two `enabledPlugins: true` ids already installed at user scope; one headless `claude -p` session run from that directory; then @@ -219,13 +218,12 @@ repo path sharing one `installedAt` second) has the same shape: one session-star per `true` entry per checkout path. The write happens even though nothing new was fetched; the record is a pin, not a download. -**Precedence explains why a user-scope duplicate does not prevent the project record.** Per -[settings-reference](https://code.claude.com/docs/en/settings-reference#enabledplugins) (fetched -2026-09-05), `enabledPlugins` resolves managed > `--settings` > local > project > user, and -"Project settings take precedence over user settings, so setting a plugin to false in -~/.claude/settings.json doesn't disable a plugin that the project's .claude/settings.json enables. -To opt out of a project-enabled plugin on your machine, set it to false in .claude/settings.local.json -instead." Precedence settles which `enabledPlugins` value is effective, and the probe above +**Precedence explains why a user-scope duplicate does not prevent the project record.** We resolve +`enabledPlugins` managed > `--settings` > local > project > user, so a user-scope `false` never +overrides a project-scope `true`, and the opt-out we advise on one machine is a `false` in +`.claude/settings.local.json`. Pointer: for the precedence and the local opt-out, see +[`enabledPlugins`](https://code.claude.com/docs/en/settings-reference#enabledplugins). As of: +2026-09-05. Precedence settles which `enabledPlugins` value is effective, and the probe above establishes that an effective project-scope `true` writes its own record regardless of the user scope. So every project-scope `true` duplicating a user-scope install produces one version-pinned project record per plugin per checkout. @@ -247,85 +245,93 @@ section above records as unreapable. **Nothing on either side of the boundary reaps the result.** `git worktree remove` deletes the directory and does not touch `~/.claude`, and the section above records that no CLI verb removes a -record by path (`-s project` acts on the cwd only, `prune` is dependency-only). The product's own -retention sweep does not cover them either: the "Cleaned up automatically" list at -[claude-directory](https://code.claude.com/docs/en/claude-directory) (fetched 2026-09-05) names -nothing under `~/.claude/plugins/`. That the per-project records are a live, maintained mechanism -rather than vestigial state is visible in the Claude Code changelog for 2.1.224, "Fixed plugin -install records being silently corrupted when the same plugin is installed in multiple projects". -Nothing between 2.1.200 and 2.1.261 adds a prune-by-path verb. - -**Synced plugins are the contrast case, not a source of these records.** Per -[plugins-reference](https://code.claude.com/docs/en/plugins-reference) ("Synced plugins", fetched -2026-09-05), plugins enabled on a claude.ai account load as `@synced` in Cowork and cloud -sessions "with no marketplace and no install record", and "Claude Code doesn't load them in sessions -you start in your own terminal." That is enablement without any record at all, so a synced plugin -never explains a project-scope row. Whether a custom GitHub marketplace can be enabled at account -level on a personal account is undocumented. +record by path (`-s project` acts on the cwd only, `prune` is dependency-only). We treat the +product's own retention sweep as not covering them either: as of 2026-09-05 its list named nothing +under `~/.claude/plugins/` (pointer: +[Cleaned up automatically](https://code.claude.com/docs/en/claude-directory#cleaned-up-automatically)). +We treat the per-project records as a live, maintained mechanism rather than vestigial state, +because the [2.1.224 changelog entry](https://code.claude.com/docs/en/changelog#2-1-224) fixes a +defect in them. Nothing between 2.1.200 and 2.1.261 adds a prune-by-path verb. + +**Synced plugins are the contrast case, not a source of these records.** We treat a plugin enabled +on a claude.ai account (`@synced`) as enablement with no install record at all, so a synced +plugin never explains a project-scope row. Pointer: for synced plugins, see +[Plugins synced from claude.ai](https://code.claude.com/docs/en/plugins/loading#synced-plugins). As +of: 2026-09-05. Whether a custom GitHub marketplace can be enabled at account level on a personal +account is undocumented. ## `/reload-plugins`: bare by default, `--force` for the MCP-cache-invalidation case -Everything in this section is doc-sourced and was re-fetched 2026-09-05 with the quoted text -unchanged. The *behavior*, what a bare reload actually warns about in a live session, is **not -re-run on 2.1.261**: it needs an interactive session, which a non-interactive probe pass cannot -drive. The `≥ 2.1.163` gate for `--force` is likewise **not re-verified on 2.1.261**, because the -current docs page states the flag without naming the version that introduced it. - -**Headless sessions can run it.** `/reload-plugins` also runs in the desktop app, the Agent SDK, -and non-interactive `-p` when it is typed into the session directly, from Claude Code 2.1.260. -Plugin MCP server changes in those sessions wait until the next session. -**Claim, basis, as of, recheck:** that sentence, -[prompt caching](https://code.claude.com/docs/en/prompt-caching), 2026-09-28, and a re-fetch of -that paragraph that drops those sessions. Whether a loop whose skill body is already in context +Every decision in this section rests on a docs pointer last re-read 2026-09-05. The *behavior*, +what a bare reload actually warns about in a live session, is **not re-run on 2.1.261**: it needs +an interactive session, which a non-interactive probe pass cannot drive. The `≥ 2.1.163` gate for +`--force` is likewise **not re-verified on 2.1.261**, because the current docs page states the flag +without naming the version that introduced it. + +**Headless sessions can run it.** We treat `/reload-plugins` as available, from Claude Code +2.1.260, in `-p` runs, Agent SDK sessions and the desktop app, only on input typed +into the session (a copy relayed over Remote Control or a message is refused), and with plugin MCP +server changes waiting for the next session. Whether a loop whose skill body is already in context can reach the command is unprobed; `lanes` `context/refresh.md` owns that limit. -**Verified against `code.claude.com/docs/en/discover-plugins`**: `/reload-plugins` refreshes skills, -agents, hooks, MCP, and LSP servers in-process. It does **not** cover monitors. Per -`code.claude.com/docs/en/plugins-reference`, "monitors require a session restart". Recommend bare `/reload-plugins` by default; call out the restart requirement -only when an updated plugin ships a monitor. - -**An install can now activate itself, but not the installs this skill issues.** As of Claude Code -2.1.221, an install started from the in-session `/plugin` interface reports its own activation state: -per `code.claude.com/docs/en/discover-plugins` (re-fetched 2026-09-05, unchanged), the summary says either -`Plugin is now active.`, meaning "Claude Code activated the plugin as part of the install", or -`Run /reload-plugins to activate.`, which happens "because activating it would invalidate the prompt -cache or because the activation attempt failed". Before 2.1.221, "no install took effect in the -current session until you ran `/reload-plugins` or restarted". - -This does **not** relax the reload guidance below, because `sync` installs through the shell command, -not the interface: "The `claude plugin install` shell command doesn't run in a session, so Claude Code -loads the plugins it installs the next time you start Claude Code, or when you run `/reload-plugins` -in a session that's already open." So a `sync` report still ends with reload guidance for everything -it installed. The activation line matters only for reading a user's own `/plugin` install summary. +- **Pointer**: for reloads in sessions without an interactive terminal, see + [Sessions without an interactive terminal](https://code.claude.com/docs/en/plugins/cli-reference#sessions-without-an-interactive-terminal), + the `/reload-plugins` row of [Commands](https://code.claude.com/docs/en/commands#all-commands), + and [Enabling or disabling a plugin](https://code.claude.com/docs/en/prompt-caching#enabling-or-disabling-a-plugin). +- **As of**: 2026-09-28 +- **Recheck trigger**: those sections stop listing headless sessions, or change the typed-input + and MCP limits. + +**Monitors need a restart.** We treat `/reload-plugins` as refreshing skills, agents, hooks, MCP +and LSP servers in-process, but not monitors. Recommend bare `/reload-plugins` by default; call out +a session restart only when an updated plugin ships a monitor. Pointer: +[`/reload-plugins`](https://code.claude.com/docs/en/plugins/cli-reference#reload-plugins) and +[`monitors`](https://code.claude.com/docs/en/plugins-reference#monitors). As of: 2026-09-05. + +**An install can now activate itself, but not the installs this skill issues.** From Claude Code +2.1.221, an install started from the in-session `/plugin` interface reports its activation state, +and this skill reads two lines of that summary: `Plugin is now active.` means the plugin is live +with no reload, and `Run /reload-plugins to activate.` means it is not (the prompt-cache case or a +failed activation). Before 2.1.221, no install took effect in the session without a reload or +restart. + +This does **not** relax the reload guidance below, because `sync` installs through the shell +command, not the interface, and we treat a shell install as loading only at the next start or the +next `/reload-plugins`. So a `sync` report still ends with reload guidance for everything it +installed. The activation line matters only for reading a user's own `/plugin` install summary. When they say a plugin is already active, believe the summary rather than telling them to reload again; and when the summary named the prompt-cache case, that is the same condition `--force` exists for below. -`--force` is real (Claude Code ≥ 2.1.163). **The general condition it exists for is prompt-cache -invalidation**, per `code.claude.com/docs/en/discover-plugins`: "When the reload would invalidate -the prompt cache, the command warns and skips until you rerun it with `--force`." +- **Pointer**: for the install summary, see + [Install a plugin](https://code.claude.com/docs/en/discover-plugins#install-a-plugin); for shell + installs, see + [Install from your shell](https://code.claude.com/docs/en/discover-plugins#install-from-your-shell). +- **As of**: 2026-09-05 +- **Recheck trigger**: either summary line changes, or a shell install starts activating in an open + session. -The MCP case is the docs' worked example of that condition, not the condition itself: a plugin -providing an MCP server whose tools aren't deferred by tool search "costs more when its tools aren't -deferred by tool search", so it is **the common cause** of the warning, but treating it as the sole -trigger tells a reader that a warning arising any other way is not a `--force` case, when it is. +`--force` is real (Claude Code ≥ 2.1.163). We treat it as the answer to one warning only: the +reload declining a change to the session's MCP or LSP tools because of the prompt cache. Which +changes raise that warning is read at the pointer below, not predicted here. -So follow the docs' own two-step rather than predicting the cause: +So follow this two-step rather than predicting the cause: -> Check the install summary: if it reports `Run /reload-plugins to activate.`, run `/reload-plugins`, -> and if that warns that the reload will re-read the conversation, rerun it as `/reload-plugins --force`. +1. Check the install summary: if it reports `Run /reload-plugins to activate.`, run + `/reload-plugins`. +2. If that warns that the reload will re-read the conversation, rerun it as + `/reload-plugins --force`. Never recommend `--force` pre-emptively alongside every reload. It exists specifically to opt into a real token cost the bare command declines to pay automatically. Recommend bare; escalate on the warning. -**Headless sessions can run `/reload-plugins`.** Claim: the command is available in non-interactive -`-p` sessions, the Agent SDK, and the desktop app, from Claude Code 2.1.260. In those sessions it -runs only on input typed into the session, it does not apply plugin MCP server changes, and a copy -that arrives over Remote Control or a relayed message is refused. Basis: - and the `/reload-plugins` -row of . As of: 2026-09-28. Recheck: that section stops -listing headless sessions, or changes the typed-input and MCP limits. +- **Pointer**: for the cache warning and `--force`, see + [Reloads that change MCP tools](https://code.claude.com/docs/en/plugins/cli-reference#reloads-that-change-mcp-tools) + and [Enabling or disabling a plugin](https://code.claude.com/docs/en/prompt-caching#enabling-or-disabling-a-plugin). +- **As of**: 2026-10-01 +- **Recheck trigger**: a release note changes when `/reload-plugins` warns, or what `--force` + applies. ## `pluginConfigs` and `enabledPlugins` have OPPOSITE scope rules @@ -333,20 +339,20 @@ This skill reads both surfaces, and they do not agree on which scopes count. Get is silent in both directions, so the asymmetry is stated here once and pointed at from everywhere else. -**`pluginConfigs`: three sources only.** Re-fetched 2026-09-05, wording unchanged. Per -`code.claude.com/docs/en/plugins-reference`: "Claude -Code reads all `pluginConfigs` values from only three settings sources". Those are user settings +**`pluginConfigs`: three sources only.** This skill reads `pluginConfigs` from user settings (`~/.claude/settings.json`), `--settings`, and managed settings, with precedence -managed → `--settings` → user. In every one of those sources the value nests under `options`: +managed → `--settings` → user, and ignores entries in a project's `.claude/settings.json` or +`.claude/settings.local.json` (Claude Code ignores them from v2.1.207, since a cloned repository +could otherwise supply values). In every one of those sources the value nests under `options`: `{"pluginConfigs":{"@":{"options":{"":""}}}}`. A key placed directly -under the plugin id is silently ignored and the render shows the literal placeholder (verified -2026-09-06 on **Claude Code 2.1.263**). And explicitly: +under the plugin id is silently ignored and the render shows the literal placeholder (our probe, +2026-09-06 on **Claude Code 2.1.263**). The restriction is specific to `pluginConfigs`. -> Entries in a project's `.claude/settings.json` or `.claude/settings.local.json` are ignored. Both -> files live in the workspace, so a cloned repository could supply values there, and those values -> would flow into plugin hook commands, MCP server configs, LSP commands, and monitor commands. -> Before v2.1.207, these entries were read. The restriction is specific to `pluginConfigs`: -> `enabledPlugins` still honors project and local settings. +- **Pointer**: for where `pluginConfigs` values are read, see + [Where values are stored](https://code.claude.com/docs/en/plugins-reference#where-values-are-stored). +- **As of**: 2026-09-05 +- **Recheck trigger**: the honored source set changes, or project or local settings become a + source again. **`enabledPlugins`: user, project, and local all count**, merged local > project > user. That is why `fleet-state.sh` reads all three settings maps for enablement, and why doing the same for @@ -363,37 +369,40 @@ Two consequences this skill must not get wrong: ## `userConfig` has no `enum` type -**Re-verified 2026-09-05 against the published plugin-manifest JSON Schema**: allowed `type` values -are `string`, `number`, `boolean`, `directory`, `file`. There is no `enum` *type*. The schema does -use an `enum` keyword, but only to constrain `type` itself to that list; an option cannot declare its -own allowed values. The schema's `required` array for -a `userConfig` option is `type`, `title`, `description`, and `claude plugin validate` on 2.1.261 -rejects an option that omits `title`. `install_new` ships as `type: string` with its -valid values (`ask`/`all`/`none`) documented in `description` and validated in prose by this skill, -not by the manifest schema. +This skill declares `userConfig` options only with the `type` values `string`, `number`, `boolean`, +`directory` and `file`, declares no fixed-options list, and gives every option `type`, `title` and +`description` (`claude plugin validate` on 2.1.261 rejected an option that omitted `title`, our +probe). `install_new` ships as `type: string` with its valid values (`ask`/`all`/`none`) documented +in `description` and validated in prose by this skill. The reference offers a fixed-options list; +this skill does not declare one, because doing so would raise the CLI version a consumer needs to +load the plugin, and the published JSON Schema does not carry it yet. + +- **Pointer**: for the option schema, see the published plugin-manifest JSON Schema + (), + [User configuration](https://code.claude.com/docs/en/plugins-reference#user-configuration) and + [Limit a field to fixed options](https://code.claude.com/docs/en/plugins-reference#limit-a-field-to-fixed-options). +- **As of**: 2026-10-01 +- **Recheck trigger**: the schema's `type` list or `required` array changes, the schema gains the + fixed-options key, or the marketplace raises its CLI floor to the version that section names. ## `userConfig`: an unset key renders the literal placeholder -**Claim.** When a `userConfig` key is set in none of the three `pluginConfigs` sources, the skill -render leaves that key's placeholder token unchanged; the manifest's `default` is not substituted -into skill content. A sibling key that is set, and `${CLAUDE_PLUGIN_ROOT}`, substitute in the same -render, so the unchanged token is the unset signal and not a substitution failure. `SKILL.md`'s -**Configured value** line reads that token as "unset, use the default `ask`". - -**Basis.** Empirical probe on a throwaway plugin from a local marketplace: first observed 2026-07-23 -on Claude Code 2.1.218, re-run 2026-09-06 on **Claude Code 2.1.263** with the same result. The -plugins reference (fetched 2026-09-11) describes `default` only as "Value used when the user -provides nothing" and states the substitution surface as every value being available for -substitution, through its `user_config` placeholder, in MCP and LSP server configs and hook -commands, and "Non-sensitive values can also be substituted in skill and agent content." It does not -say the default substitutes into skill content, so the page and the probe do not conflict; the -probe settles what the page leaves open. The page's own sentence, placeholder and all, is quoted in -`SKILL.md`, the one file where the token may appear. - -**As of.** 2026-09-06, Claude Code 2.1.263. - -**Recheck trigger.** Any Claude Code minor-version bump that touches plugin `userConfig` -substitution, or a change to the plugins reference's `default` row or its substitution sentence. +When a `userConfig` key is set in none of the three `pluginConfigs` sources, we treat the skill +render as leaving that key's placeholder token unchanged, with the manifest's `default` not +substituted into skill content. A sibling key that is set, and `${CLAUDE_PLUGIN_ROOT}`, substitute +in the same render, so the unchanged token is the unset signal and not a substitution failure. +`SKILL.md`'s **Configured value** line reads that token as "unset, use the default `ask`". Our +probe settles a point the plugins reference leaves open: the reference does not say whether the +default substitutes into skill content, so the two do not conflict. + +- **Pointer**: our probe on a throwaway plugin from a local marketplace, first run 2026-07-23 on + Claude Code 2.1.218 and re-run 2026-09-06 on **Claude Code 2.1.263** with the same result; for + the `default` field and the substitution surfaces, see + [Reference a saved value](https://code.claude.com/docs/en/plugins-reference#reference-a-saved-value) + (read 2026-09-11). +- **As of**: 2026-09-06, Claude Code 2.1.263 +- **Recheck trigger**: any Claude Code minor-version bump that touches plugin `userConfig` + substitution, or a change to the reference's `default` field or its substitution surfaces. **Probe recipe.** The `pluginConfigs` payload must nest the key under `options`, in the shape "`pluginConfigs` and `enabledPlugins` have OPPOSITE scope rules" above gives; a key placed directly @@ -408,13 +417,14 @@ receives `userConfig` substitution". ## Renames are CC-native (≥ v2.1.193) -Claude Code rewrites a marketplace's `renames` map into installed/enabled state automatically at -session start (old id → new id; `null` means removal). This skill hard-codes no rename knowledge. -Its only rename-adjacent behavior is that anything present in the current catalog but absent from -`installed_plugins.json` shows up as `missing_from_install`, which naturally covers a renamed -plugin's new id. Renames mapping requires ≥ v2.1.193, re-confirmed 2026-09-05 against -`code.claude.com/docs/en/plugin-marketplaces`, which still says "Automatic migration requires Claude -Code v2.1.193 or later." The `claude plugin prune` ≥ v2.1.121 gate is **not re-verified on 2.1.261**: +We rely on Claude Code to apply a marketplace's `renames` map to installed and enabled state at +session start (old id → new id; `null` means removal), from v2.1.193. This skill hard-codes no +rename knowledge. Its only rename-adjacent behavior is that anything present in the current catalog +but absent from `installed_plugins.json` shows up as `missing_from_install`, which naturally covers +a renamed plugin's new id. Pointer: +[Migrate users with a renames map](https://code.claude.com/docs/en/plugins/host-marketplace#migrate-users-with-a-renames-map). +As of: 2026-09-05. Recheck trigger: the version floor or the map's semantics change. The +`claude plugin prune` ≥ v2.1.121 gate is **not re-verified on 2.1.261**: the current docs describe `prune` without naming an introducing version, so the gate stands on its original source and nothing this pass found contradicts it. @@ -465,22 +475,25 @@ marketplace was added, so the registry side is clean. `~/.claude/plugins/cache/< stayed on disk with the plugin's version directory intact. The tree is not a permanent orphan. The `uninstall` step wrote a `.orphaned_at` marker file (epoch -milliseconds) into the version directory, and the marker survived the marketplace removal. Per -[plugins-reference](https://code.claude.com/docs/en/plugins-reference) -("Plugin cache", fetched 2026-09-06), a marked directory is removed by the background sweep roughly -14 days later, the sweep runs only while at least one plugin is installed, and a cache folder is -removed only once it holds no directory or symlink. So a removed marketplace's tree is swept on the -same clock as any other orphaned version, marketplace folder included, provided the machine keeps -any plugin installed. The page says nothing about marketplace removal itself; the marker is the -observation that connects the two. The sweep does visit a removed marketplace's cache folders: the -Claude Code 2.1.270 debug log (2026-09-14, recorded on -[#3835](https://github.com/melodic-software/claude-code-plugins/issues/3835)) shows its -folder-retention pass logging `Keeping /: it still holds a directory, a -symlink, a versioned archive or an entry of unknown type` for two removed marketplaces. That the -sweep removes a marked version directory under a removed marketplace at 14 days is inferred from -the documented rule, not yet observed; both probes that would show it were lost before a reading. -**Recheck trigger:** any Claude Code release note or `plugins-reference` change touching -marketplace removal, the orphan sweep, or the cache layout. +milliseconds) into the version directory, and the marker survived the marketplace removal. We model +the background sweep as removing a marked directory about 14 days later, running only while at +least one plugin is installed, and removing a cache folder only once it holds no directory or +symlink. So a removed marketplace's tree is swept on the same clock as any other orphaned version, +marketplace folder included, provided the machine keeps any plugin installed. The docs say nothing +about marketplace removal itself; the marker is our observation that connects the two. The sweep +does visit a removed marketplace's cache folders: the Claude Code 2.1.270 debug log (2026-09-14, +recorded on [#3835](https://github.com/melodic-software/claude-code-plugins/issues/3835)) shows its +folder-retention pass keeping two removed marketplaces' folders because each still held content. +That the sweep removes a marked version directory under a removed marketplace at 14 days is +inferred from the documented rule, not yet observed; both probes that would show it were lost +before a reading. + +- **Pointer**: for the orphan sweep, see + [Cleanup of previous versions](https://code.claude.com/docs/en/plugins/loading#cleanup-of-previous-versions); + the marker and the debug log are our probes. +- **As of**: 2026-09-06 +- **Recheck trigger**: any Claude Code release note or docs change touching marketplace removal, + the orphan sweep, or the cache layout. **That observation was taken after an `uninstall`, with no install record left for the removal to find.** With plugins from the marketplace still installed, the same command deletes their records at @@ -498,10 +511,17 @@ installed, because the sweep never runs there; a version directory under it that ## `autoUpdate` is a background complement, not a substitute -Official-Anthropic marketplaces default `autoUpdate: true`; third-party and local-dev marketplaces -default it off (absent from `known_marketplaces.json`, not `false`). When on, Claude Code refreshes -marketplace data and bumps already-installed plugins once per session start, after a random delay of -up to ten minutes. This skill never mutates the setting. It only reports the marketplace's current -`autoUpdate` state and suggests enabling it when off, since it never overlaps with what this skill -covers (new-plugin install, `enabledPlugins` completeness, divergence detection/convergence, -deterministic on-demand execution). +This skill reads a marketplace's `autoUpdate` as off when the key is absent from +`known_marketplaces.json` (not only when it is `false`), and treats auto-update as a background +refresh of already-installed plugins after session start, never as a substitute for `sync`. This +skill never mutates the setting. It only reports the marketplace's current `autoUpdate` state and +suggests enabling it when off, since it never overlaps with what this skill covers (new-plugin +install, `enabledPlugins` completeness, divergence detection/convergence, deterministic on-demand +execution). + +- **Pointer**: for the per-marketplace defaults and when auto-update runs, see + [Keep plugins updated](https://code.claude.com/docs/en/discover-plugins#keep-plugins-updated) and + [When auto-update runs](https://code.claude.com/docs/en/plugins/loading#when-auto-update-runs). +- **As of**: 2026-09-05 +- **Recheck trigger**: a marketplace kind's default changes, or auto-update starts installing new + plugins. diff --git a/plugins/instruction-placement/.claude-plugin/plugin.json b/plugins/instruction-placement/.claude-plugin/plugin.json index f9381ce59e..fdfa3d2f70 100644 --- a/plugins/instruction-placement/.claude-plugin/plugin.json +++ b/plugins/instruction-placement/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "instruction-placement", - "version": "0.18.2", + "version": "0.18.3", "description": "Routes agent-instruction content to the surface that loads it at the right moment. The audit skill sweeps a repository's instruction layer and its ordinary markdown for content whose scope is narrower than the surface carrying it, meaning conventions keyed to one file type or one subtree sitting in an always-loaded CLAUDE.md or AGENTS.md, and for normative conventions stranded in documentation Claude never loads at all, then classifies each against a routing rubric and proposes a destination whose `paths:` glob is machine-validated before it is ever offered. Safety-class content (irreversible actions, secrets, data integrity, external publication, compliance, agent authority) is hard-denied from demotion and reported as held back rather than proposed, because demotion trades guaranteed presence for conditional presence: a deferred surface is absent until a read matches it, absent after a compaction until that trigger recurs, and never inherited by a subagent, which re-acquires it only by reading a covered path itself. Every accepted move regenerates an always-loaded index of deferred surfaces, which is what keeps a demoted rule discoverable from any context that has not happened to touch a path it covers. The audit is read-only and emits a diffable findings artifact; realignment is a separate skill gated per item with no blanket-approve path; a deterministic check skill gates that every rule glob still resolves and the index is current; a migrate skill moves a repository to AGENTS.md as the content home and keeps a CLAUDE.md shim while one is needed; and a setup skill verifies the one thing no other gate can see: that nothing in the repository stops Claude Code reading the index target, since a CLAUDE.md in the working directory or above it is read instead of the AGENTS.md beside it.", "author": { "name": "Melodic Software", diff --git a/plugins/instruction-placement/CHANGELOG.md b/plugins/instruction-placement/CHANGELOG.md index 7c21d6028d..409edd6da5 100644 --- a/plugins/instruction-placement/CHANGELOG.md +++ b/plugins/instruction-placement/CHANGELOG.md @@ -3,6 +3,22 @@ All notable changes to the `instruction-placement` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.18.3] - 2026-10-02 + +### Changed + +- **`remove-shims.sh` prints the shim-cost record by its new labels.** The price of removal is the + decision paragraph directly above the record's `Pointer` line plus the pointer itself, stopping + before the as-of line, in place of the old claim-to-basis span. A new test fixture pins that + output. +- **The `migrate` sources and verification references, the `check` and `setup` skills, the README + and `verified-mechanics.md` hold decisions plus pointers.** Each record states our decision in our + words and points at the exact documentation section, an as-of date and a recheck trigger, with + no upstream text. `cutover-check.sh` names the same pointer-record shape in its comments and + usage. +- `migrate`'s prose-shim gotcha states the plan's action in our words and points at the memory + page's workaround section. + ## [0.18.2] - 2026-10-02 ### Changed diff --git a/plugins/instruction-placement/README.md b/plugins/instruction-placement/README.md index 8de432b453..026351c58d 100644 --- a/plugins/instruction-placement/README.md +++ b/plugins/instruction-placement/README.md @@ -60,9 +60,10 @@ Four facts make the naive version of this migration actively harmful, and each o The full evidence, including a first-party repro, is in [`context/verified-mechanics.md`](context/verified-mechanics.md). -**A rule without `paths:` costs exactly what `CLAUDE.md` costs.** Unscoped rules load at launch with -the same priority as `.claude/CLAUDE.md`. Moving a section into `.claude/rules/` without a glob is -bookkeeping, not a saving. The glob is the product. +**A rule without `paths:` costs exactly what `CLAUDE.md` costs.** The plugin prices an unscoped +rule as always-loaded, the same as `.claude/CLAUDE.md` (the pointer is in the evidence file above). +Moving a section into `.claude/rules/` without a glob is bookkeeping, not a saving. The glob is the +product. **Nothing that defers is inherited, and no deferred surface says it exists.** Measured on Claude Code 2.1.268: a subagent dispatched *after* its parent had loaded a nested `CLAUDE.md`, a nested @@ -87,16 +88,17 @@ it load, though it stays the cover for the sessions that cannot read `AGENTS.md` Claude Code's own AGENTS.md support is version- and session-dependent, which is why the plugin's posture is to write the shim while a root `CLAUDE.md` exists in a repository. -- **Claim**: Claude Code reads `AGENTS.md` as the project instructions only where there is no - `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` in the working directory or above it, and - attaches a subdirectory's `AGENTS.md` on a Read there under the same condition; reading - `AGENTS.md` directly depends on a CLI version floor and on the session, both in - [`skills/migrate/reference/sources.md`](skills/migrate/reference/sources.md), "The minimum CLI - version". -- **Basis**: [memory](https://code.claude.com/docs/en/memory), "AGENTS.md", "When Claude Code reads - AGENTS.md" and "When AGENTS.md support is unavailable"; confirmed by canary runs on 2.1.278. -- **As of**: 2026-09-29. -- **Recheck trigger**: that section changes which file names count for the check, or a release note +The plugin treats an `AGENTS.md`, at the root or in a subdirectory, as shadowed wherever a +`CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` sits in the working directory or above it, +and treats direct `AGENTS.md` reading as dependent on the CLI floor and session recorded in +[`skills/migrate/reference/sources.md`](skills/migrate/reference/sources.md), "The minimum CLI +version". Canary runs on 2.1.278 observed the shadowing. + +- **Pointer**: for when Claude Code reads `AGENTS.md` and when that support is unavailable, see + and + . +- **As of**: 2026-09-29 +- **Recheck trigger**: those sections change which file names shadow `AGENTS.md`, or a release note names `AGENTS.md` or instruction-file loading. ## What this plugin does NOT buy you @@ -107,9 +109,9 @@ always-loaded file. **That claim was measured and not supported**, so it has bee softened. Across 32 trials at two bloat levels, using a realistic 251-line always-loaded file and an extreme -1,927-line one nearly ten times the official 200-line guidance, a clear convention was followed -**100% of the time in both arms**. Full method, caveats, and the ceiling effect the run hit: -[`evals/adherence-results.md`](evals/adherence-results.md). +1,927-line one nearly ten times the 200-line size the memory page recommends, a clear convention +was followed **100% of the time in both arms**. Full method, caveats, and the ceiling effect the +run hit: [`evals/adherence-results.md`](evals/adherence-results.md). Weigh a migration on context cost and on the promote lane. Do not expect your instructions to be obeyed better afterwards. @@ -136,10 +138,12 @@ posture keeps shared content portable: subtree conventions go in a nested `AGENT `CLAUDE.md` shim beside it wherever a `CLAUDE.md` on that path would otherwise be read instead, and the generated index lives in the root `AGENTS.md` when one exists. -One semantic difference is deliberately not papered over: other agents resolve `AGENTS.md` -nearest-wins, while Claude concatenates the whole ancestor chain. Subtree content is therefore -written as additive and self-contained, and a candidate that only makes sense as an override is -reported rather than moved. +One semantic difference is deliberately not papered over: the plugin assumes other agents resolve +`AGENTS.md` nearest-wins while Claude loads the whole ancestor chain (for Claude's order, see +, as of 2026-10-01; recheck when +that section changes how ancestor files combine). Subtree content is therefore written as additive +and self-contained, and a candidate that only makes sense as an override is reported rather than +moved. ## Scope boundary: what this plugin does not own diff --git a/plugins/instruction-placement/context/verified-mechanics.md b/plugins/instruction-placement/context/verified-mechanics.md index c7b19e8dcf..63c38f7212 100644 --- a/plugins/instruction-placement/context/verified-mechanics.md +++ b/plugins/instruction-placement/context/verified-mechanics.md @@ -4,11 +4,12 @@ The evidence spine behind every routing decision this plugin makes. Read it befo candidate whose destination turns on *when* content loads, *whether it survives compaction*, or *whether a subagent can see it*. -**Citation posture.** Claims are marked *(doc)* when an official Anthropic page states them, -*(measured)* when this plugin's own first-party repro established them, and *(inferred)* when -neither. An inference is never presented as either of the other two. A `measured` claim names the -Claude Code version it was taken on, because these mechanics have moved between releases and a -version-less measurement cannot be re-verified or aged out. +**Citation posture.** A claim marked *(doc)* is one this plugin relies on an official Anthropic +page for; the page section is named in the pointer record beside it and is read there, not +restated here. *(measured)* marks what this plugin's own first-party repro established, and +*(inferred)* marks neither. An inference is never presented as either of the other two. A +`measured` claim names the Claude Code version it was taken on, because these mechanics have moved +between releases and a version-less measurement cannot be re-verified or aged out. ## Contents @@ -28,7 +29,7 @@ whether a move is safe. |---|---|---|---| | Root `CLAUDE.md` (cwd + ancestors) | Session start, in full *(doc)* | Re-read from disk and re-injected *(doc)* | **Yes** *(measured)* | | `@import` from root `CLAUDE.md` | Session start, inlined *(doc)* | With its parent *(inferred)* | **Yes** *(measured)* | -| Unscoped `.claude/rules/*.md` | Session start, "same priority as `.claude/CLAUDE.md`" *(doc)* | Re-injected *(doc)* | Unmeasured at 2.1.268; was **No** *(measured 2.1.238)* | +| Unscoped `.claude/rules/*.md` | Session start, priced as `.claude/CLAUDE.md` *(doc)* | Re-injected *(doc)* | Unmeasured at 2.1.268; was **No** *(measured 2.1.238)* | | Path-scoped rule (`paths:`) | On **read** of a matching file *(doc, measured)* | Re-injected when a match recurs *(doc)* | **Yes, on a matching read inside the subagent itself** *(measured 2.1.268)* | | Nested `CLAUDE.md` | On read of a file in that subtree *(doc, measured)* | Reloads when the subtree is touched again *(doc)* | **Yes, on a matching read inside the subagent itself** *(measured 2.1.268)* | | `@import` from a **nested** `CLAUDE.md` | With its parent, deferred *(measured)* | With its parent *(inferred)* | **Yes, with its parent, inside the subagent itself** *(measured 2.1.268)* | @@ -36,6 +37,25 @@ whether a move is safe. | Bare nested `AGENTS.md` (no shim), nothing on its path | On read of a file in that subtree, where AGENTS.md support is available *(doc, measured 2.1.278)* | Reloads when the subtree is touched again *(doc)* | Yes, on a matching read inside the subagent itself *(measured 2.1.278)* | | Skill body | On invocation *(doc)* | Listing re-injected; body on re-invoke *(doc)* | Discovered via the Skill tool *(doc)* | +The plugin prices each instruction-file destination by the *(doc)* cells above and re-derives them +from the memory page, never from this table. + +- **Pointer**: for when each instruction file loads, see + , + , + and + ; for what reloads after + compaction, . +- **As of**: 2026-10-01 +- **Recheck trigger**: any of those sections changes when a surface loads or reloads, or a release + note names instruction-file loading, rules, or compaction. + +The skill-body row predates this record and has fired its trigger: a 2026-10-01 read of + no longer matches its +compaction cell. It awaits re-derivation from that section and +; until then treat that cell as +unverified. + Three facts from that table carry the whole design: - **An unscoped rule costs exactly what `CLAUDE.md` costs.** Moving a section from `CLAUDE.md` into @@ -84,15 +104,16 @@ Four findings follow, each of which a rubric rule depends on: rather than exclusive: a nested `CLAUDE.md` and a nested `AGENTS.md` both attach on a Read in that directory, and what the root `CLAUDE.md` does is make Claude Code read `CLAUDE.md` files *instead of* `AGENTS.md`. - - **Claim**: Claude reads `AGENTS.md` only where no `CLAUDE.md`, `.claude/CLAUDE.md` or - `CLAUDE.local.md` sits in the working directory or above it, and attaches a subdirectory's - `AGENTS.md` on a Read there under the same condition; reading it directly - depends on a CLI version floor and on the session, both in - `skills/migrate/reference/sources.md`, "The minimum CLI version". - - **Basis**: [memory](https://code.claude.com/docs/en/memory), "AGENTS.md", "When Claude Code - reads AGENTS.md", "When AGENTS.md support is unavailable"; canary runs on 2.1.278. - - **As of**: 2026-09-29. - - **Recheck trigger**: that section changes which file names count for the check, or a release + The rubric treats an `AGENTS.md`, at the root or in a subdirectory, as shadowed wherever a + `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` sits in the working directory or above + it, and treats direct `AGENTS.md` reading as dependent on the CLI floor and session recorded in + `skills/migrate/reference/sources.md`, "The minimum CLI version". Canary runs on 2.1.278 + observed the shadowing. + - **Pointer**: for when Claude Code reads `AGENTS.md` and when that support is unavailable, see + and + . + - **As of**: 2026-09-29 + - **Recheck trigger**: those sections change which file names shadow `AGENTS.md`, or a release note names `AGENTS.md` or instruction-file loading. 4. **A subagent inherits none of the parent's on-demand loads.** Dispatched *after* the parent had already loaded all five surfaces, a general-purpose subagent reported exactly @@ -149,16 +170,19 @@ Two boundaries this measurement does **not** cross, stated so nothing generalize ### Verification record -- **Claim.** On Claude Code 2.1.268, a path-scoped `.claude/rules/` file and a nested - `CLAUDE.md`/`AGENTS.md` pair are injected inside a general-purpose subagent when that subagent - reads a path the surface covers, and the glob matches the requested path whether or not the file - exists. A subagent still inherits none of its parent's deferred loads. -- **Basis.** First-party probe run inside a subagent dispatched into this repository on the harness - reported by `claude --version` as `2.1.268 (Claude Code)`, observing the `Contents of :` - blocks appended to `Read` results for one nonexistent `**/*.py` path and one existing - `plugins/autonomy/CLAUDE.md`, against the rules and shims tracked at commit `49912c63`. -- **As of.** 2026-09-13. -- **Recheck trigger.** The consuming repository's Claude Code minor version moves past 2.1.268; or +The rubric relies on what this probe observed on Claude Code 2.1.268: a path-scoped +`.claude/rules/` file and a nested `CLAUDE.md`/`AGENTS.md` pair are injected inside a +general-purpose subagent when that subagent reads a path the surface covers, the glob matches the +requested path whether or not the file exists, and a subagent still inherits none of its parent's +deferred loads. + +- **Pointer**: the first-party probe above, run inside a subagent dispatched into this repository + on the harness reported by `claude --version` as `2.1.268 (Claude Code)`, observing the + `Contents of :` blocks appended to `Read` results for one nonexistent `**/*.py` path and + one existing `plugins/autonomy/CLAUDE.md`, against the rules and shims tracked at commit + `49912c63`. +- **As of**: 2026-09-13 +- **Recheck trigger**: the consuming repository's Claude Code minor version moves past 2.1.268; or a Claude Code release note touches subagent context inheritance, memory loading, or path-scoped rule triggering; or a real session observes a covered `Read` inside a subagent that injects nothing. Any of these obliges re-running the three steps above and refreshing this record with @@ -181,38 +205,46 @@ context. The index guarantees **availability**, not attention: injection is auto is discretionary. It therefore mitigates rather than erases, which is why the hard-deny class below is not also delegated to it. -**The write-trigger gap.** "Path-scoped rules trigger when Claude reads files matching the pattern, -not on every tool use" *(doc)*. Editing an existing file implies reading it, so the common case -holds; **creating a new file does not**. Content that governs the *creation* of files, such as +**The write-trigger gap.** The rubric treats a path-scoped rule as firing on a read of a matching +file and on nothing else *(doc; pointer in the surface-table record)*. Editing an existing file +implies reading it, so the common case holds; **creating a new file does not**. Content that governs the *creation* of files, such as scaffolding templates, "every new component must…", and file-header requirements, is therefore served badly by a path-scoped rule no matter how clean its glob looks. *Closed by:* routing creation-governing content to a directory-nested surface or leaving it always-loaded, never to `paths:`. -**The compaction gap.** Root `CLAUDE.md` is re-read from disk after `/compact`; deferred surfaces -return only when their trigger recurs *(doc)*. A long session that compacts mid-task and then works +**The compaction gap.** The rubric prices compaction as bringing back the root `CLAUDE.md` and +leaving each deferred surface to return only when its trigger recurs *(doc; pointer in the +surface-table record)*. A long session that compacts mid-task and then works in a different subtree never re-loads what it demoted. *Closed by:* pricing this into every recommendation, and by the hard-deny class for content whose absence is unrecoverable. ## Glob semantics and their budgets -All *(doc)* unless marked. The `check` skill enforces each mechanically. - -- Patterns are globs over repo-relative paths: `**/*.ts`, `src/**/*`, `*.md` (root only), - `src/components/*.tsx`. -- Brace expansion is supported and multiplies: `src/*.{ts,tsx}` is two patterns, - `{a,b}/{c,d}/*.{ts,tsx}` is eight. A rule's whole `paths:` list shares one budget of **1,000 - expanded patterns and 4 MiB**. A pattern exceeding the budget is used **unexpanded**, so its - literal braces match nothing, a silent no-op rather than an error. -- `[` opens a bracket expression. A `[` that cannot be read as one, as in `photos [2024/**`, makes - that pattern match nothing while the rule's other patterns keep working. Escape a literal one as - `photos \[2024/**`. -- Symlinked paths into the project directory match as of v2.1.198. -- Rules are discovered recursively under `.claude/rules/`, so subdirectories are organizational. -- User-level `~/.claude/rules/` load before project rules, giving project rules higher priority. - -A glob that matches **zero** tracked files is not an error to Claude Code. The rule simply never +The `check` skill (through `scripts/glob-tools.sh`) holds these as its own settings and enforces +each mechanically: + +- Every `paths:` entry is a glob over repo-relative paths, matched against tracked files. +- Brace groups multiply, and the check counts a rule's whole `paths:` list against a shared cap: + **1,000 expanded patterns, 4 MiB in total**. A pattern over the budget is reported `over-budget`, + because Claude Code would use it unexpanded and its literal braces would match nothing, a silent + no-op rather than an error. +- A `[` the check cannot read as a bracket expression is reported `bad-bracket`, because that one + pattern would match nothing while the rule's other patterns keep working. The fix is an escaped + literal, `\[`. +- The check walks `.claude/rules/` recursively, so subdirectories are organizational. + +A glob that matches **zero** tracked files raises no error in Claude Code. The rule simply never fires. That silence is exactly why `check` treats it as a failure. +- **Pointer**: for glob syntax, the brace-expansion budget, bracket expressions, recursive + discovery and user-level rule order, see + , + and + . +- **As of**: 2026-10-01 +- **Recheck trigger**: that section moves the budget, changes how an unreadable `[` or an + over-budget pattern behaves, or a release note names rule globs. + ## Re-verification These mechanics are version-sensitive and have changed repeatedly across minor releases. Re-run the @@ -225,5 +257,7 @@ claim's confidence: The repro is cheap: a temp git repo with canary tokens on each surface, an `InstructionsLoaded` hook appending each payload to a log, one headless run that reads a file in the subtree, and a read of -the log. `InstructionsLoaded` is observability-only and cannot block or modify a load, so the -measurement never perturbs what it measures. +the log. The repro relies on `InstructionsLoaded` being an observe-only event, so the measurement +does not perturb what it measures (see +, as of 2026-10-01; recheck when that +section gives the event a decision control). diff --git a/plugins/instruction-placement/skills/check/SKILL.md b/plugins/instruction-placement/skills/check/SKILL.md index a456637f30..5b5976a981 100644 --- a/plugins/instruction-placement/skills/check/SKILL.md +++ b/plugins/instruction-placement/skills/check/SKILL.md @@ -58,14 +58,17 @@ same silence repeats one level down: a nested `AGENTS.md` is indexed as a surfac Claude reads its directory, and a `CLAUDE.md` on that path takes its place unless one of them imports or symlinks it. Sync, reachability, and wiring are independent questions; ask all three. -- **Claim**: Claude Code reads `AGENTS.md` only where no `CLAUDE.md`, `.claude/CLAUDE.md` or - `CLAUDE.local.md` sits in the working directory or above it, at the root and in a subdirectory - alike; reading it directly depends on a CLI version floor and on the session, both in - `skills/migrate/reference/sources.md`, "The minimum CLI version". -- **Basis**: [memory](https://code.claude.com/docs/en/memory), "AGENTS.md", "When Claude Code reads - AGENTS.md", "When AGENTS.md support is unavailable"; canary runs on Claude Code 2.1.278. -- **As of**: 2026-09-29. -- **Recheck trigger**: that section changes which file names count for the check, or a release note +The reachability and wiring gates treat an `AGENTS.md`, at the root and in a subdirectory alike, +as shadowed wherever a `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` sits in the working +directory or above it, and treat direct `AGENTS.md` reading as dependent on the CLI floor and +session recorded in `skills/migrate/reference/sources.md`, "The minimum CLI version". Canary runs +on Claude Code 2.1.278 observed the shadowing. + +- **Pointer**: for when Claude Code reads `AGENTS.md` and when that support is unavailable, see + and + . +- **As of**: 2026-09-29 +- **Recheck trigger**: those sections change which file names shadow `AGENTS.md`, or a release note names `AGENTS.md` or instruction-file loading. Over-broad is the one **warning** rather than a failure: breadth is a judgment about whether a diff --git a/plugins/instruction-placement/skills/migrate/SKILL.md b/plugins/instruction-placement/skills/migrate/SKILL.md index 610dc89a37..ba680baebc 100644 --- a/plugins/instruction-placement/skills/migrate/SKILL.md +++ b/plugins/instruction-placement/skills/migrate/SKILL.md @@ -323,8 +323,8 @@ the run removed. The restore writes the one import line back and verifies it byt than asking git for an index copy a staged edit may have replaced; a restore that did not land says so and the run still exits non-zero. -Removal is priced, and the price is printed before the confirmation: a directly read `AGENTS.md` -fires no `InstructionsLoaded` hook, and `/memory` lists it only from v2.1.280 +Removal is priced, and the price is printed before the confirmation: the hook-visible load goes, +and `/memory` becomes the only place to see the file ([`reference/sources.md`](reference/sources.md), "What shim removal costs"). That is a decision to make, not tidying. @@ -348,29 +348,30 @@ The shim is what carries `AGENTS.md` in two cases a repository cannot talk itsel `CLAUDE.md` above the file being read instead of it, and a session that cannot read `AGENTS.md` directly at all. -- **Claim**: Claude Code reads `AGENTS.md` as the project instructions only where there is no - `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` in the working directory or above it, and - attaches a subdirectory's `AGENTS.md` when a Read opens a file there and that subdirectory has - none of those three names of its own. Reading it directly needs v2.1.277 or later; the sessions - that cannot are listed in `reference/shim-droppable.md`, condition D. A `CLAUDE.md` containing `@AGENTS.md` never makes - Claude read the file twice. -- **Basis**: [memory](https://code.claude.com/docs/en/memory), fetched 2026-09-29 (49,601 bytes; - slug in `llms.txt`; first heading "How Claude remembers your project"), sections "AGENTS.md", - "When Claude Code reads AGENTS.md", "When AGENTS.md support is unavailable", and "Remove an - earlier AGENTS.md workaround". Canary runs on Claude Code 2.1.278 confirmed the displacement - rule; this 2026-09-29 pass did not re-run the displacement canary. -- **As of**: 2026-09-29. -- **Recheck trigger**: that page changes which file names count for the check or which sessions lack - support, or a release note names `AGENTS.md` or instruction-file loading. - -It is priced, not free. From v2.1.280, `/memory` lists a directly read `AGENTS.md`; before that, -`/memory` and `/context` did not. `InstructionsLoaded` still does not fire for an `AGENTS.md` read -through the Project instructions setting, and does fire when a `CLAUDE.md` imports it or is a -symlink to it. One reached through a shim behaves like part of its `CLAUDE.md` and keeps the hook. -That is a reason the shim is worth its ~55 tokens, and a reason removing it later is a decision -rather than tidying. The dated quotes are in `reference/sources.md`, "What shim removal costs". - -Every upstream fact the cutover turns on lives as a four-part dated record in +This skill keeps the shim wherever a `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` can +shadow the `AGENTS.md` it carries, at the root or in a subdirectory, and wherever a session may +lack direct `AGENTS.md` support (the CLI floor is in `reference/sources.md`, "The minimum CLI +version"; the sessions that cannot read `AGENTS.md` directly are in `reference/shim-droppable.md`, +condition D). It treats the import as safe to keep: the shim never makes Claude read the file +twice. Canary runs on Claude Code 2.1.278 observed the shadowing; the 2026-09-29 pass did not +re-run that canary. + +- **Pointer**: for when Claude Code reads `AGENTS.md`, when that support is unavailable, and what + removing a shim involves, see + , + and + . +- **As of**: 2026-09-29 +- **Recheck trigger**: those sections change which file names shadow `AGENTS.md` or which sessions + lack support, or a release note names `AGENTS.md` or instruction-file loading. + +It is priced, not free. A load through a shim keeps the `InstructionsLoaded` hook, which this +plugin's load verification depends on; a direct read loses it, and only recent versions list a +directly read `AGENTS.md` in `/memory`. That is a reason the shim is worth its ~55 tokens, and a +reason removing it later is a decision rather than tidying. The price, with its pointers, is in +`reference/sources.md`, "What shim removal costs". + +Every upstream fact the cutover turns on lives as a dated pointer record in [`reference/sources.md`](reference/sources.md): the remote flag and how its code default is read, the documented feature-flag dependency, the CLI floor, the `claude-code-action` release to CLI map, the CI canary result, the current fleet grade, and what shim removal costs. `cutover-check.sh` @@ -380,17 +381,17 @@ skip checking. Read it before arguing about the shim from memory. One record the detects a load through the `InstructionsLoaded` hook, so it measures a **shimmed** surface and cannot see an `AGENTS.md` that Claude reads directly. -**One setting changes the reading, and no repository can ship it.** Under `instructionFiles: -claude-md-and-agents-md`, Claude Code reads "Your `CLAUDE.md` and `AGENTS.md` files together, each -directory's `CLAUDE.md` files first and its `AGENTS.md` after them". The page does not say when a -subdirectory's `AGENTS.md` loads under that value, so an `UNWIRED` row stays a finding there too -and the nested shim stays. The import is harmless there: "Claude Code skips an -`AGENTS.md` it has already loaded, so one that your `CLAUDE.md` imports or symlinks to isn't read -twice". The value is a user, `--settings` or managed setting, ignored in project and local settings, -so a repository cannot rely on it and the gates keep the default's answer -([memory](https://code.claude.com/docs/en/memory), "Choose which instruction files load"; fetched -2026-10-01; recheck when that table changes, the page comes to state when a subdirectory's -`AGENTS.md` loads under that value, or a release note names the setting). +**One setting changes the reading, and no repository can ship it.** For an operator who sets +`instructionFiles: claude-md-and-agents-md`, the memory page does not say when a subdirectory's +`AGENTS.md` loads, so an `UNWIRED` row stays a finding there too and the nested shim stays; the +import stays harmless. The setting is not one a repository's project or local settings can carry, +so the gates keep the default's answer. + +- **Pointer**: for the `instructionFiles` values and where the setting is read, see + . +- **As of**: 2026-10-01 +- **Recheck trigger**: that table changes, the page comes to state when a subdirectory's + `AGENTS.md` loads under that value, or a release note names the setting. ## Boundary, the built-in `cc-plugin-agents-md` plugin @@ -442,9 +443,9 @@ The `reference/` files write the plugin's root directory as ``, whi `${CLAUDE_PLUGIN_ROOT}`. Put that path in place of the placeholder before running a command or writing it into a brief. Those files arrive through the Read tool as plain bytes, so a `${…}` token in them would reach the Bash tool unsubstituted, and the Bash tool's environment has no -`CLAUDE_PLUGIN_ROOT` to expand it from. Basis: the plugins reference, -, verified -2026-09-30; recheck when that table adds supporting files to where a `${…}` reference resolves. +`CLAUDE_PLUGIN_ROOT` to expand it from. Pointer: for where each `${…}` variable resolves, see +. As of: +2026-09-30. Recheck trigger: that table adds supporting files to where a `${…}` reference resolves. ## Next @@ -458,8 +459,11 @@ in sync after the move. that review as correct and load nothing. - **`/init` writes `CLAUDE.md`.** Running it after a migration re-creates the content the migration moved out. Say so in the PR body of a repository whose contributors run it. -- **A `CLAUDE.md` that tells Claude in prose to read `AGENTS.md` does not load it.** Only an - `@AGENTS.md` import does. Replace the sentence, never keep it as a belt. +- **A prose shim is a finding, not a shim.** The plan converts a `CLAUDE.md` sentence asking for + `AGENTS.md` into the one-line `@AGENTS.md` import, or drops the file where the session reads + `AGENTS.md` natively; it never keeps the sentence beside the import. Pointer: + [Remove an earlier AGENTS.md workaround](https://code.claude.com/docs/en/memory#remove-an-earlier-agents-md-workaround). + As of: 2026-10-01. Recheck trigger: that section changes what a prose instruction does. - **A committed symlink is not portable.** On a Windows checkout it materializes as a plain text file holding the link target, so the import is the form that works everywhere. - **A `CLAUDE.local.md` one developer keeps silently turns `AGENTS.md` off for them.** It counts for diff --git a/plugins/instruction-placement/skills/migrate/reference/sources.md b/plugins/instruction-placement/skills/migrate/reference/sources.md index a654744e46..7317dbf8c5 100644 --- a/plugins/instruction-placement/skills/migrate/reference/sources.md +++ b/plugins/instruction-placement/skills/migrate/reference/sources.md @@ -1,18 +1,18 @@ # Upstream sources the cutover check reads -Every upstream fact the AGENTS.md cutover turns on, as a four-part record: claim, basis, as-of date, -recheck trigger, per the -[upstream-drift convention](../../../../../docs/conventions/upstream-drift/README.md). The records -are here so a reader can judge the check's verdict without re-deriving the research, and so a -firing trigger has one place to land. +Every upstream fact the AGENTS.md cutover turns on, as a pointer record per the +[upstream-drift convention](../../../../../docs/conventions/upstream-drift/README.md): our decision +or our own probe result, a pointer to where the specific lives, an as-of date, and a recheck +trigger. The records are here so a reader can judge the check's verdict without re-deriving the +research, and so a firing trigger has one place to land. Per-run output (per-condition tables, graded commit SHAs) is posted as a comment on the tracker issue. This file holds only the current grade of each fact: per the convention's "When a trigger fires", refreshing a date with no verdict change is no entry and no version bump. Every page below was fetched by the convention's rung-1 route (`curl` the `.md` to a file, search -the file locally), slug confirmed against `https://code.claude.com/docs/llms.txt`, and quoted from -the bytes rather than paraphrased. +the file locally), slug confirmed against `https://code.claude.com/docs/llms.txt`, and read from +the bytes. No page text is stored here. Contents: [The remote flag](#the-remote-flag-and-how-its-code-default-is-read) · [Feature-flag dependency](#the-documented-feature-flag-dependency) · @@ -28,59 +28,57 @@ Contents: [The remote flag](#the-remote-flag-and-how-its-code-default-is-read) ## The remote flag, and how its code default is read -- **Claim**: reading `AGENTS.md` directly is gated on the GrowthBook flag `tengu_agents_md_mod`. In - the shipped bundle the built-in plugin exports `isOnByDefault`, whose value is a minifier-assigned - identifier declared once nearby as `!0` (true) or `!1` (false). In Claude Code 2.1.282 that - identifier is `W` and the window around the second flag-string occurrence reads - `isOnByDefault:()=>W` and `var W=!0;var B=()=>oi("tengu_agents_md_mod",W)`, so the **code - default is true**. The first occurrence is a string in another window and does not tie to an - `isOnByDefault` export. -- **Basis**: the installed bundle at `node_modules/@anthropic-ai/claude-code/bin/claude.exe` - (`claude` on PATH), 2.1.282, sha256 - `3afe8535c0cc33f0e24f7b25dab7a1727b8b592196f8496a8bc302ba2161eed3`, read as bytes. The flag - string occurs twice, at offsets 103001528 and 225456771. Only 225456771 is the code site. - Offsets and the identifier are per build and per host, so the check resolves both at run time and - hardcodes neither. `cutover-check.sh` on 2026-09-28 printed that same offset, identifier, and - `var W=!0`. -- **As of**: 2026-09-28. +Condition 1 treats direct `AGENTS.md` reading as gated on the GrowthBook flag +`tengu_agents_md_mod` and reads the flag's code default out of the shipped bundle at run time. Our +probe of Claude Code 2.1.282 found the flag string twice. Only the second occurrence sits in a +window where the built-in plugin's `isOnByDefault` export resolves to a minifier-assigned +identifier (`W` on that build) declared `!0`, so the **code default is true**. The first +occurrence is a string in another window and ties to no `isOnByDefault` export. + +- **Pointer**: our byte read of the installed bundle at + `node_modules/@anthropic-ai/claude-code/bin/claude.exe` (`claude` on PATH), 2.1.282, sha256 + `3afe8535c0cc33f0e24f7b25dab7a1727b8b592196f8496a8bc302ba2161eed3`. The flag string occurs at + offsets 103001528 and 225456771; only 225456771 is the code site. Offsets and the identifier are + per build and per host, so the check resolves both at run time and hardcodes neither. + `cutover-check.sh` on 2026-09-28 printed that same offset and identifier and a `!0` declaration. +- **As of**: 2026-09-28 - **Recheck trigger**: any Claude Code version bump, a bundle where no window around the flag string carries `isOnByDefault`, or a window where the captured identifier resolves ambiguously. Each of those is `[UNREACH]` for the check, never `[MET]`. ## The documented feature-flag dependency -- **Claim**: `env-vars` still carries the section `## Features that need feature-flag fetching`. - Fetching is still skipped for a session setting `DISABLE_GROWTHBOOK`, `DISABLE_TELEMETRY`, - `DO_NOT_TRACK` or `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC`, a session on a third-party provider, - and a Claude apps gateway session. The same page's subsection "First session after an install or - upgrade" still states a flag-gated feature can be missing in that first session. The "With - fetching off, you can't" list under that heading does not mention `AGENTS.md`. The page does not - mention `AGENTS.md` at all. -- **Basis**: `https://code.claude.com/docs/en/env-vars.md`, fetched 2026-09-28, 507,134 bytes. The - heading is at line 512, the list is lines 520-534, and "First session after an install or - upgrade" is at line 536. A case-insensitive search of the file for `AGENTS.md` returned no line. - The slug appears in `llms.txt` and the body's first heading is "Environment variables", so the - page is the one requested. The same day's `cutover-check.sh` fetch reported the heading found - and the AGENTS.md bullet absent. -- **As of**: 2026-09-28. +Condition 1's second leg fetches the env-vars page and grades the dependency gone when the section +headed `## Features that need feature-flag fetching` no longer names `AGENTS.md`. It requires the +later subsection heading "First session after an install or upgrade" as its marker that the page +arrived whole. Our 2026-09-28 fetch found the heading and the marker, and a case-insensitive search +of the whole page for `AGENTS.md` found no line. + +- **Pointer**: for which features need feature-flag fetching and which sessions skip it, see + and + . Our fetch of + `env-vars.md` on 2026-09-28 was 507,134 bytes, with the heading at line 512, its list at lines + 520-534 and the marker at line 536; the slug is in `llms.txt` and the body's first heading + confirmed the page. The same day's `cutover-check.sh` fetch reported the heading found and the + AGENTS.md bullet absent. +- **As of**: 2026-09-28 - **Recheck trigger**: the heading is renamed or removed, the page names `AGENTS.md` beside a flag again, or the AGENTS.md bullet returns to the list. A missing heading is `[UNREACH]` for the check, because the absence of a heading cannot be read as the absence of the dependency. ## The minimum CLI version -- **Claim**: Claude Code reads `AGENTS.md` as project instructions from **v2.1.277**, verbatim: - "Reading `AGENTS.md` directly requires Claude Code v2.1.277 or later." The same page lists - "You're on a Claude Code version before v2.1.277" among the cases where support is unavailable, - and its removal procedure step 2 is "Run `claude --version` and confirm v2.1.277 or later." - That step continues: "Before v2.1.281, some sessions, such as those on Amazon Bedrock or with - telemetry disabled, couldn't load `AGENTS.md` either, so on those versions update to v2.1.281 - or later." Below v2.1.277 no session reads it, whatever the flag says. The Bedrock and - telemetry-disabled gap is stated for versions before v2.1.281, not as a limit of v2.1.282. -- **Basis**: `https://code.claude.com/docs/en/memory.md`, fetched 2026-09-29 by the rung-1 route, - 49,601 bytes; the floor sentence is at line 352, the unavailable bullet at 402, and step 2 at - 578. The slug is in `llms.txt` and the first heading is "How Claude remembers your project". -- **As of**: 2026-09-29. +`cutover-check.sh` grades every pin against the CLI floor **v2.1.277**: below it the check counts no +session as reading `AGENTS.md` directly, whatever the flag says. Some session kinds need a later +version than the floor; the check does not grade that, so an operator whose fleet runs those +session kinds reads it at the pointer. + +- **Pointer**: for the version that reads `AGENTS.md` directly and the sessions that lack support, + see and + . Our rung-1 fetch + of `memory.md` on 2026-09-29 was 49,601 bytes; the slug is in `llms.txt` and the first heading + confirmed the page. +- **As of**: 2026-09-29 - **Recheck trigger**: the memory page states a different floor, or a release note moves it. ## `claude-code-action` release to installed CLI version @@ -99,46 +97,48 @@ Each row was re-derived on 2026-09-28 by resolving the tag to its commit and rea | `v1.0.231` | `cfc3eb22bfed5c26ef66e3223c982af27e4524de` | `2.1.278` | yes | | `v1.0.235` | `756cc22e19660d20e8cc9496b4f242475a7f7790` | `2.1.283` | yes | -- **Claim**: the table above, and the general rule that the value is assigned in - `base-action/action.yml` as a shell line `CLAUDE_CODE_VERSION=""` immediately before the - `Installing Claude Code v${CLAUDE_CODE_VERSION}...` echo (line 150 at all five commits). -- **Basis**: `gh api repos/anthropics/claude-code-action/commits/` for the commit, then - `gh api repos/anthropics/claude-code-action/contents/base-action/action.yml?ref=`, +The check maps each pin to the CLI version that `base-action/action.yml` assigns to +`CLAUDE_CODE_VERSION` at the pinned commit, in the shell step that installs Claude Code (line 150 +at all five commits), and reads the table above as that mapping. + +- **Pointer**: our derivation, `gh api repos/anthropics/claude-code-action/commits/` for the + commit, then `gh api repos/anthropics/claude-code-action/contents/base-action/action.yml?ref=`, re-read 2026-09-28 for every row in the table. `v1.0.235` is an annotated tag whose tag object `f33305702e43b9f71a532e6f80aed9a399df8288` points at commit `756cc22e19660d20e8cc9496b4f242475a7f7790`. That commit is the `uses:` pin in `melodic-software/ci-workflows` `b570d97203c7973b25c14e3de91c5ff3a4aa0e82`, `.github/workflows/claude-review.yml:167` and `.github/workflows/claude-security-review.yml:160`, both commented `# v1.0.235`. -- **As of**: 2026-09-28. +- **As of**: 2026-09-28 - **Recheck trigger**: a new pin appears in any in-scope repository, or the action stops assigning `CLAUDE_CODE_VERSION` in `base-action/action.yml`. A pin whose file carries no such assignment is `[UNREACH]`, never a pass: an unreadable map is not a satisfied floor. ## The CI canary -- **Claim**: on a GitHub-hosted `ubuntu-24.04` runner with a genuinely fresh install (no - `~/.claude` before the first step), at action `v1.0.235` installing CLI 2.1.283, a workspace - holding a lone non-empty `AGENTS.md` and no `CLAUDE.md` at any level returned the `AGENTS.md` - canary token `CI-AGENTS-51C2` with zero tool calls, on the first session after the install and - again on a second session in the same job. The arrange step deleted the repository's own - `CLAUDE.md` shim from the ephemeral workspace, and the job's instruction-file listing showed only - `AGENTS.md`. The 2026-09-19 run at `v1.0.231` / CLI 2.1.278 returned the same for a lone - `AGENTS.md`, and there a `CLAUDE.md` carrying its own token suppressed the `AGENTS.md` one, the - documented precedence. Separately, `claude-code-action` **rejects the `push` event** - (`Unsupported event type: push`); runs are started by REST dispatch - (`gh api -X POST repos///actions/workflows//dispatches -f ref=`). -- **Basis**: `melodic-software/knowledge-corpus` run `36666844023`, event `workflow_dispatch`, head +Condition 2's canary half rests on this observation of ours. On a GitHub-hosted `ubuntu-24.04` +runner with a genuinely fresh install (no `~/.claude` before the first step), at action `v1.0.235` +installing CLI 2.1.283, a workspace holding a lone non-empty `AGENTS.md` and no `CLAUDE.md` at any +level returned the `AGENTS.md` canary token `CI-AGENTS-51C2` with zero tool calls, on the first +session after the install and again on a second session in the same job. The arrange step deleted +the repository's own `CLAUDE.md` shim from the ephemeral workspace, and the job's instruction-file +listing showed only `AGENTS.md`. The 2026-09-19 canary at `v1.0.231` / CLI 2.1.278 returned the +same for a lone `AGENTS.md`, and there a `CLAUDE.md` carrying its own token suppressed the +`AGENTS.md` one. Separately, `claude-code-action` **refuses a `push` event**, so canaries are +started by REST dispatch +(`gh api -X POST repos///actions/workflows//dispatches -f ref=`). + +- **Pointer**: `melodic-software/knowledge-corpus` run `36666844023`, event `workflow_dispatch`, head SHA `24041cf1613312c6aff31ff637d9ea4beb753302` on the throwaway branch `test/agents-md-ci-canary`, deleted after the run. The workflow is the 2026-09-19 canary workflow (knowledge-corpus `91f0285b1ba2cc7329dff0f89dbbae020171a61b`) cut to case A, `workflow_dispatch` only, with both action uses pinned to `756cc22e19660d20e8cc9496b4f242475a7f7790 # v1.0.235`. Log lines: `2.1.283 (Claude Code)`, reply `CI-AGENTS-51C2`, tools used `[]`, in both report steps. The earlier run `35475056935` - (head SHA `91f0285b1ba2cc7329dff0f89dbbae020171a61b`) is quoted in the migration slice's - `PROOF-ci-canary-knowledge-corpus-2026-09-19.md`. The event list is the action's own + (head SHA `91f0285b1ba2cc7329dff0f89dbbae020171a61b`) is recorded in the migration slice's + `PROOF-ci-canary-knowledge-corpus-2026-09-19.md`. For the events the action accepts, see its own `src/github/context.ts` `parseGitHubContext` switch at the pinned commit. -- **As of**: 2026-09-30. +- **As of**: 2026-09-30 - **Recheck trigger**: a new action release, a new CLI floor, or a runner image change. A repository `CLAUDE.md` shim left in the workspace turns the run into a test of the shim, so the arrange step must remove it. The canary does not show **why** the flag-gated feature was @@ -148,40 +148,40 @@ Each row was re-derived on 2026-09-28 by resolving the tag to its commit and rea ## Canary host (#4282) -- **Claim**: the CI-canary component of cutover condition 2 rests on the knowledge-corpus run - recorded in [The CI canary](#the-ci-canary), at `v1.0.235`. The owner chose one case-A run at - that pin on knowledge-corpus over accepting the 2026-09-19 run, and the run passed. -- **Basis**: the owner decision comments on #4282 of 2026-09-29 ("Option B, run by the agent"). +The CI-canary component of cutover condition 2 rests on the knowledge-corpus run recorded in +[The CI canary](#the-ci-canary), at `v1.0.235`. The owner chose one case-A run at that pin on +knowledge-corpus over accepting the 2026-09-19 run, and the run passed. + +- **Pointer**: the owner decision comments on #4282 of 2026-09-29 ("Option B, run by the agent"). `melodic-software/ci-workflows#599`, the pin half, closed COMPLETED 2026-09-20. `gh repo view melodic-software/claude-lane-sandbox --json isArchived` returned `{"isArchived":true,"name":"claude-lane-sandbox"}` on 2026-09-29, and `git ls-remote --heads origin` in the knowledge-corpus tree returned only `refs/heads/main` on 2026-09-30, after the run's branch was deleted. -- **As of**: 2026-09-30. +- **As of**: 2026-09-30 - **Recheck trigger**: a run id newer than the one in [The CI canary](#the-ci-canary), or a `claude-code-action` release newer than the pinned one. ## Current fleet grade -- **Claim**: `cutover-check.sh` graded every condition `[MET]` and printed - `remove-shims may run`, over ten repositories with none unreadable: claude-code-plugins, medley, - songwriting, claude-code-proxy, knowledge-corpus, codex-plugins, ci-runner, agent-plugins, - cursor-plugins, and provisioning. Condition 1 is `[MET]` because the bundle code default for - `tengu_agents_md_mod` is true. Condition 2 is `[MET]` on pin arithmetic: the - only pin in the ten is medley's `v1.0.231` (CLI 2.1.278), and the two ci-workflows pins, outside - the ten, are `v1.0.235` (CLI 2.1.283), all at or above 2.1.277. This grade printed the - 2026-09-19 canary run at `v1.0.231` as its CI-canary evidence. The record now names run - `36666844023` at `v1.0.235` / CLI 2.1.283, the ci-workflows pin (see - [The CI canary](#the-ci-canary) and [Canary host](#canary-host-4282)); the grade was not re-run. - Condition 3 is `[MET]`: - both `claude -p` legs returned the canary line from a lone non-empty `AGENTS.md`, and it grades - only on a logged-in host. Condition 4 is `[MET]`: 70 path-detection rows, every one - acknowledged with a reviewed reason. -- **Basis**: `cutover-check.sh`, no `--skip-canary`, over the ten repositories on 2026-09-29, +The fleet's current grade is our own run: `cutover-check.sh` graded every condition `[MET]` and +printed `remove-shims may run`, over ten repositories with none unreadable: claude-code-plugins, +medley, songwriting, claude-code-proxy, knowledge-corpus, codex-plugins, ci-runner, agent-plugins, +cursor-plugins, and provisioning. Condition 1 is `[MET]` because the bundle code default for +`tengu_agents_md_mod` is true. Condition 2 is `[MET]` on pin arithmetic: the only pin in the ten is +medley's `v1.0.231` (CLI 2.1.278), and the two ci-workflows pins, outside the ten, are `v1.0.235` +(CLI 2.1.283), all at or above 2.1.277. This grade printed the 2026-09-19 canary run at `v1.0.231` +as its CI-canary evidence. The record now names run `36666844023` at `v1.0.235` / CLI 2.1.283, the +ci-workflows pin (see [The CI canary](#the-ci-canary) and [Canary host](#canary-host-4282)); the +grade was not re-run. Condition 3 is `[MET]`: both `claude -p` legs returned the canary line from a +lone non-empty `AGENTS.md`, and it grades only on a logged-in host. Condition 4 is `[MET]`: 70 +path-detection rows, every one acknowledged with a reviewed reason. + +- **Pointer**: `cutover-check.sh`, no `--skip-canary`, over the ten repositories on 2026-09-29, Claude Code 2.1.284, exit 0. The trees were read as they stood and not fetched, so the grade is - for those local commits, not for the current default branch. The per-repository commit table and the - full per-condition output are in the comment on tracker issue #4281, not copied here. -- **As of**: 2026-09-29. + for those local commits, not for the current default branch. The per-repository commit table and + the full per-condition output are in the comment on tracker issue #4281, not copied here. +- **As of**: 2026-09-29 - **Recheck trigger**: a pin move in any in-scope repository, a Claude Code release whose changelog touches `AGENTS.md` or instruction-file loading, the monthly due date of `agents-md-cutover-check` in `.github/recurring-schedule.json`, or an in-scope repository @@ -189,55 +189,57 @@ Each row was re-derived on 2026-09-28 by resolving the tag to its commit and rea ## Loader behavior of Cursor, Grok Build, and Muse Code -- **Claim**: Each tool was run headless against one recipe tree, and these results are - observed, not docs or source grade. Cursor CLI loads `AGENTS.md`, `CLAUDE.md` and - `CLAUDE.local.md` together at session start, follows symlinks, applies no size cap through - 262,156 bytes, and does not expand `@path` imports. It reads no other name (`Agents.md`, - `AGENT.md`, `.claude/CLAUDE.md` are absent). `.cursor/rules/*.md` never loads; `.mdc` loads - only with frontmatter. Started at the git root, nested files attach when a file under them is - read; started in the nested directory, the ancestor chain loads (12 levels seen). In the - non-git copy the attach on read did not occur. Grok Build loads all eight names of its - documented list per directory, and the two `.claude/` names are gated by - `GROK_CLAUDE_AGENTS_ENABLED` (set to `false`, they vanish). No switch stops it - reading a plain `CLAUDE.md`. It does not expand `@path` imports, follows symlinks, and shows no - size cap through 262,156 bytes. In a trusted folder outside a git repository it loads the - working directory only, so "nothing outside a git repository loads" is contradicted; with - folder trust off nothing project-level loads. Muse Code loads one file per directory with - `AGENTS.md` first, and `CLAUDE.md` alone loads when no `AGENTS.md` exists. It does not expand - `@path` imports and follows symlinks. It skips an `AGENTS.md` over 256,000 bytes ("over the - 256000 byte load limit"), loads one of 244,676 bytes with only the head reaching the model - (65,536-byte delegation startup limit), and skips project files unless the workspace is - trusted (`--trust-workspace`). In a git repository it loads the chain from root to working - directory (12 levels seen); from the root, a read under a nested directory attaches nothing; - outside git it loads the working directory only. - Still open, each with its reason: - - Cursor Team, Project, User precedence: needs a Team plan and the editor rules UI, not - observable headless. - - Cursor editor against CLI, and the editor's "always applied" wording: the editor was not run. - - Cursor `~/.cursor/rules` as a synced file: the path does not exist on the host, and sync - cannot be observed headless. - - Grok path-only reminder text for out-of-chain files: contents stay unloaded and the model - later read the nested files itself, but the streamed transcript carries no reminder text. - - Grok `MAX_WALK_DEPTH` of 10: a working directory at depth 12 loaded all 12 levels, so the - recipe does not show what the constant bounds. - - Muse "sibling `CLAUDE.md` never opened": Muse names the shadowed file on stderr ("is ignored - this session because AGENTS.md takes precedence in that directory"), but whether the file - is opened needs `strace`, which is not installed on the host. - - Muse user-rules path and its Windows resolution: no Muse-native user-rules file exists on - the host, and the Windows path needs a Windows host. -- **Basis**: headless runs on one Linux host with cursor-agent `2026.09.28-64d2043` +The migration plans for other tools on these observations of ours, not on docs or source. Each +tool was run headless against one recipe tree. Cursor CLI loads `AGENTS.md`, `CLAUDE.md` and +`CLAUDE.local.md` together at session start, follows symlinks, applies no size cap through 262,156 +bytes, and does not expand `@path` imports. It reads no other name (`Agents.md`, `AGENT.md`, +`.claude/CLAUDE.md` are absent). `.cursor/rules/*.md` never loads; `.mdc` loads only with +frontmatter. Started at the git root, nested files attach when a file under them is read; started +in the nested directory, the ancestor chain loads (12 levels seen). In the non-git copy the attach +on read did not occur. Grok Build loads eight file names per directory, and the two `.claude/` +names are gated by `GROK_CLAUDE_AGENTS_ENABLED` (set to `false`, they vanish). No switch stops it +reading a plain `CLAUDE.md`. It does not expand `@path` imports, follows symlinks, and shows no size +cap through 262,156 bytes. In a trusted folder outside a git repository it loads the working +directory only, so the expectation that nothing loads outside a git repository does not hold; with +folder trust off nothing project-level loads. Muse Code loads one file per directory with +`AGENTS.md` first, and `CLAUDE.md` alone loads when no `AGENTS.md` exists. It does not expand +`@path` imports and follows symlinks. It skips an `AGENTS.md` over 256,000 bytes and says so on +stderr, loads one of 244,676 bytes with only the head reaching the model (65,536-byte delegation +startup limit), and skips project files unless the workspace is trusted (`--trust-workspace`). In +a git repository it loads the chain from root to working directory (12 levels seen); from the +root, a read under a nested directory attaches nothing; outside git it loads the working directory +only. + +Still open, each with its reason: + +- Cursor Team, Project, User precedence: needs a Team plan and the editor rules UI, not observable + headless. +- Cursor editor against CLI, and how the editor applies an always-on rule: the editor was not run. +- Cursor `~/.cursor/rules` as a synced file: the path does not exist on the host, and sync cannot + be observed headless. +- Grok path-only reminder text for out-of-chain files: contents stay unloaded and the model later + read the nested files itself, but the streamed transcript carries no reminder text. +- Grok `MAX_WALK_DEPTH` of 10: a working directory at depth 12 loaded all 12 levels, so the recipe + does not show what the constant bounds. +- Whether Muse ever opens a shadowed sibling `CLAUDE.md`: Muse names the shadowed file on stderr as + ignored in favor of `AGENTS.md`, but whether the file is opened needs `strace`, which is not + installed on the host. +- Muse user-rules path and its Windows resolution: no Muse-native user-rules file exists on the + host, and the Windows path needs a Windows host. + +- **Pointer**: headless runs on one Linux host with cursor-agent `2026.09.28-64d2043` (`cursor-agent -p --mode ask --trust`), grok `1.0.41 (4220f3b224a6)` (`grok -p --tools ""` with `GROK_FOLDER_TRUST=0`, and `grok inspect --json`), and Muse Code `1.4.1 (1.4.1-R4503.1)` (`muse exec --trust-workspace --disable-shell --disable-write`). The tree, prompt, invocations and expected results are in [`reference/verification.md`](verification.md#the-loader-recipe-for-other-tools); the raw - transcripts are not committed. Each result is one model sample except where repeats agreed. -- **Upstream pointers**: Cursor rules, `https://cursor.com/docs/rules` (fetched 2026-09-30, HTTP - 200, mentions `AGENTS.md`); Grok Build, `https://docs.x.ai/build/overview` (fetched - 2026-09-30, HTTP 200, mentions `AGENTS.md`). Neither page states the loader semantics above, - which is why they are recorded as observed. Muse Code: no public documentation or issue - tracker was found, so its results rest on the recipe alone. -- **As of**: 2026-09-30. + transcripts are not committed. Each result is one model sample except where repeats agreed. For + each tool's own documentation, see Cursor rules, , and Grok + Build, (both fetched 2026-09-30, HTTP 200, both naming + `AGENTS.md`); neither states the loader semantics above, which is why they are recorded as + observed. Muse Code: no public documentation or issue tracker was found, so its results rest on + the recipe alone. +- **As of**: 2026-09-30 - **Recheck trigger**: a new release of any of the three tools, a host that can run the editor, a Team plan, a Windows host, or `strace`, which would settle the open bullets, either upstream page coming to state loader behavior, or any change to the recipe in @@ -253,41 +255,39 @@ Codex leg, and the rollout check are in This is the price of the cutover, and `remove-shims` prints it before it asks. -- **Claim**: `InstructionsLoaded` hooks **do not fire** for an `AGENTS.md` Claude reads directly - through the Project instructions setting. hooks.md, verbatim: "This event doesn't fire when - Claude reads `AGENTS.md` directly through the **Project instructions** setting. It does fire when - a `CLAUDE.md` imports your `AGENTS.md`, with `load_reason` set to `include` as for any other - imported file, and when `CLAUDE.md` is a symlink to it, as a normal `CLAUDE.md` load." The memory - page's difference table says the same hook row as "Don't fire" for `AGENTS.md` read through the - setting. Two further rows of that table: a directory added with `--add-dir` under - `CLAUDE_CODE_ADDITIONAL_DIRECTORIES_CLAUDE_MD` loads its `CLAUDE.md` and not its `AGENTS.md`, and - an external `@path` import loads with no prompt only where external imports were already approved - for that project. `/memory` **does** list a directly read `AGENTS.md` on the current page, - verbatim: "To check whether Claude read your `AGENTS.md`, run `/memory` and look for its path in - the list." The same page says "Before v2.1.280, `/memory` and `/context` didn't list an - `AGENTS.md` that Claude read directly." -- **Basis**: `https://code.claude.com/docs/en/hooks.md`, "InstructionsLoaded" (246,601 bytes, the - quoted paragraph at line 1290) and `https://code.claude.com/docs/en/memory.md`, "Where AGENTS.md - differs from CLAUDE.md" (49,601 bytes, table at line 412) and "My AGENTS.md isn't loading" (lines - 581 and 583). Both fetched by the rung-1 route on 2026-09-29, both slugs present in `llms.txt`. -- **As of**: 2026-09-29. -- **Recheck trigger**: either page changes that table, that paragraph, or the `/memory` listing - sentence. +Removing the shims costs the hook-visible load. Without a shim, the `InstructionsLoaded` hook no +longer reports the `AGENTS.md` load, so `verify-load.sh` and any hook-based audit stop seeing it +(our 2026-09-20 measurement below observed exactly that); with a shim, the load arrives as an +import and the hook reports it. The operator's way to confirm the file loaded moves to `/memory`. +We count two further differences between a direct read and a shimmed one in the price: directories +added with `--add-dir`, and external `@path` imports inside the file. Each is read live at the +pointer, never restated here. + +- **Pointer**: for the hook, see ; for + the differences between a direct read and a `CLAUDE.md` load, and the `/memory` listing, see + and + . Both pages fetched by + the rung-1 route on 2026-09-29, both slugs present in `llms.txt`. `scripts/remove-shims.sh` + prints the paragraph above this line and this pointer, and stops at the As of line. +- **As of**: 2026-09-29 +- **Recheck trigger**: either page changes that table, the hook's firing rule for `AGENTS.md`, or + the `/memory` listing. ## What the loss means for measuring the cutover -- **Claim**: `scripts/verify-load.sh` detects a load through an `InstructionsLoaded` hook, so it - **cannot measure a directly read `AGENTS.md`**. Measured this session in a scratch directory under - home holding only a non-empty `AGENTS.md` and a `README.md`: `verify-load.sh --trigger README.md - --expect AGENTS.md` printed one `LOADED` row for the user-scope `~/.claude/CLAUDE.md` and then - `EXPECTED AGENTS.md MISSING` / `VERDICT FAIL`, exit 1, while a headless `claude -p` run from the - same directory with `--allowedTools Read`, reading `README.md` and asked to quote back the lines - of its project instructions carrying the token, returned the token. So the file loaded and the - instrument could not see it. `verify-load.sh` stays the right instrument for a **shimmed** - surface, where the load arrives as an import and the hook fires with `load_reason` `include`. -- **Basis**: the run above on Claude Code 2.1.278, Windows, 2026-09-20; corroborated by the - hooks.md quotation in the previous record. -- **As of**: 2026-09-20. +`scripts/verify-load.sh` detects a load through an `InstructionsLoaded` hook, so it **cannot +measure a directly read `AGENTS.md`**. Our measurement, in a scratch directory under home holding +only a non-empty `AGENTS.md` and a `README.md`: `verify-load.sh --trigger README.md --expect +AGENTS.md` printed one `LOADED` row for the user-scope `~/.claude/CLAUDE.md` and then +`EXPECTED AGENTS.md MISSING` / `VERDICT FAIL`, exit 1, while a headless `claude -p` run from the +same directory with `--allowedTools Read`, reading `README.md` and asked to quote back the lines of +its project instructions carrying the token, returned the token. So the file loaded and the +instrument could not see it. `verify-load.sh` stays the right instrument for a **shimmed** surface, +where the load arrives as an import and the hook fires with `load_reason` `include`. + +- **Pointer**: the run above on Claude Code 2.1.278, Windows, 2026-09-20; for the hook's firing + rule, see . +- **As of**: 2026-09-20 - **Recheck trigger**: `InstructionsLoaded` starts firing for a directly read `AGENTS.md`, or `verify-load.sh` gains a detection path that does not depend on that hook. Either one puts the two instruments back together. @@ -296,14 +296,13 @@ This is the price of the cutover, and `remove-shims` prints it before it asks. The 2026-09-20 measurement above was not repeated on this pass. -- **Claim**: hooks.md still says `InstructionsLoaded` does not fire for a directly read - `AGENTS.md`, so `verify-load.sh` still cannot see that load. What changed is the `/memory` - listing, recorded under [What shim removal costs](#what-shim-removal-costs): the memory page - says that before v2.1.280 `/memory` and `/context` did not list a directly read `AGENTS.md`, - and that the check now is to look for its path in `/memory`. -- **Basis**: the hooks.md and memory.md fetches in that cost record. This checkout's +A re-read of the hooks page found the hook gap unchanged, so this plugin still treats +`verify-load.sh` as blind to a direct load. What moved is the `/memory` listing, recorded under +[What shim removal costs](#what-shim-removal-costs). + +- **Pointer**: the hooks and memory fetches in that cost record. This checkout's `claude auth status` reported `loggedIn` false, so no new `claude -p` measurement was possible. -- **As of**: 2026-09-28. +- **As of**: 2026-09-28 - **Recheck trigger**: the same as the measurement record above. ## The built-in agents-md plugin diff --git a/plugins/instruction-placement/skills/migrate/reference/verification.md b/plugins/instruction-placement/skills/migrate/reference/verification.md index 74977fb9c7..edee0e7370 100644 --- a/plugins/instruction-placement/skills/migrate/reference/verification.md +++ b/plugins/instruction-placement/skills/migrate/reference/verification.md @@ -99,13 +99,14 @@ Then confirm in the session rollout that **no shell command went looking for the that greps its way to the answer proves the file is on disk, which was never in doubt, and not that Codex loaded it. -- **Claim**: Codex writes a session rollout to - `~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl`. A shell call appears as a JSONL record whose - `payload.type` is `custom_tool_call`, with the command inside `payload.input`; one run observed - it as `exec_command` instead, so match either. Two rollout files can land a second apart, so - pick the one whose records contain the prompt rather than the newest by mtime. -- **Basis**: observed on codex-cli 0.155.1 across the migration runs that produced this file. -- **As of**: 2026-09-19. +The check reads the session rollout at `~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl`. It matches a +shell call as a JSONL record whose `payload.type` is `custom_tool_call`, with the command inside +`payload.input`, or `exec_command`, which one run used instead. Two rollout files can land a second +apart, so it picks the one whose records contain the prompt rather than the newest by mtime. + +- **Pointer**: our observation on codex-cli 0.155.1 across the migration runs that produced this + file. +- **As of**: 2026-09-19 - **Recheck trigger**: a codex-cli release that changes the session-file layout, the record type, or where rollouts are written. @@ -167,7 +168,7 @@ Muse Code `1.4.1`: | `@path` import in an instruction file | not expanded | not expanded | not expanded | | Symlinked instruction file | followed | followed | followed | | `Agents.md`, `AGENT.md`, `.claude/CLAUDE.md` | not read | all eight names read | not read | -| 262,156-byte `AGENTS.md` | loads | loads | skipped, "over the 256000 byte load limit" | +| 262,156-byte `AGENTS.md` | loads | loads | skipped, with a load-limit message on stderr | | Nested files, cwd at the git root | attach when a file under them is read | absent | absent, and a read attaches nothing | | Nested files, cwd in the nested directory | ancestor chain loads, 12 levels | chain loads, 12 levels | chain loads, 12 levels | | Non-git copy, cwd nested | ancestors load | cwd directory only | cwd directory only | @@ -176,8 +177,8 @@ Muse Code `1.4.1`: | `.cursor/rules/x.mdc` | loads with frontmatter only | not loaded | not tested | Grok's two `.claude/` names disappear with `GROK_CLAUDE_AGENTS_ENABLED=false`; the six top-level -names stay. Muse prints "is ignored this session because AGENTS.md takes precedence in that -directory" on stderr for a shadowed sibling. +names stay. For a shadowed sibling, Muse names the file on stderr as ignored in favor of +`AGENTS.md`. ## Progressive disclosure, with a caveat diff --git a/plugins/instruction-placement/skills/migrate/scripts/cutover-check.sh b/plugins/instruction-placement/skills/migrate/scripts/cutover-check.sh index b658f70023..18ebfd4e3a 100755 --- a/plugins/instruction-placement/skills/migrate/scripts/cutover-check.sh +++ b/plugins/instruction-placement/skills/migrate/scripts/cutover-check.sh @@ -24,8 +24,9 @@ # # Every upstream number it compares against (the CLI floor, the action release # to CLI map, the CI canary run) is parsed from `reference/sources.md`, which -# carries the four-part dated record for each. Parsing is fail-hard: a record -# this script cannot read is a fact it must not silently skip checking. +# carries a pointer record (decision, pointer, as-of date, recheck trigger) for +# each. Parsing is fail-hard: a record this script cannot read is a fact it +# must not silently skip checking. # # Usage: # cutover-check.sh --repo [--repo ...] [options] @@ -75,7 +76,8 @@ Usage: cutover-check.sh --repo [--repo ...] [options] default from (default: `claude` on PATH) --env-vars-file read the env-vars page from this file instead of fetching it - --sources the four-part records to compare against + --sources the pointer records (decision, pointer, as-of + date, recheck trigger) to compare against (default: ../reference/sources.md) --claude-bin the CLI the condition-3 canary runs (default: claude) --canary-home-root where the home-cwd canary makes its scratch diff --git a/plugins/instruction-placement/skills/migrate/scripts/remove-shims.sh b/plugins/instruction-placement/skills/migrate/scripts/remove-shims.sh index 0481bbd007..d16f1bc003 100755 --- a/plugins/instruction-placement/skills/migrate/scripts/remove-shims.sh +++ b/plugins/instruction-placement/skills/migrate/scripts/remove-shims.sh @@ -4,9 +4,9 @@ # The only mutating script in this skill, and it refuses far more often than it # acts. In order, and every gate fails closed: # -# 1. It prints what removal costs, from reference/sources.md ("What shim removal -# costs"). As of the 2026-09-28 recheck, /memory lists a directly read -# AGENTS.md and InstructionsLoaded still does not fire for that read. +# 1. It prints what removal costs: the decision and pointer of the pointer +# record (decision, pointer, as-of date, recheck trigger) headed "What shim +# removal costs" in reference/sources.md. # 2. It refuses without --confirm. There is no blanket-yes and no all-repos # mode: one repository per run, one confirmation per run. # 3. It refuses unless the INSTALLED harness-memory and instruction-placement @@ -161,11 +161,16 @@ echo "=== remove-shims: $REPO ===" echo echo "What removing the shims costs (reference/sources.md, 'What shim removal costs'):" if [[ -f "$SOURCES_MD" ]]; then + # The record's decision is the paragraph directly above its Pointer line, so + # the price is that paragraph and the pointer, up to the As of line. tr -d '\r' <"$SOURCES_MD" | awk '/^## What shim removal costs/ { f = 1; next } - f && /^- \*\*Claim\*\*/ { p = 1 } - p && /^- \*\*Basis\*\*/ { exit } - p { print " " $0 }' + !f { next } + /^## / || /^- \*\*As of\*\*/ { exit } + /^- \*\*Pointer\*\*/ { for (i = 1; i <= n; i++) print " " para[i]; print ""; p = 1 } + p { print " " $0; next } + /^[[:space:]]*$/ { gap = 1; next } + { if (gap) { n = 0; gap = 0 } para[++n] = $0 }' else echo " (records file not found at $SOURCES_MD)" fi diff --git a/plugins/instruction-placement/skills/migrate/scripts/remove-shims.test.sh b/plugins/instruction-placement/skills/migrate/scripts/remove-shims.test.sh index 94984e7c29..53ad8ce01d 100755 --- a/plugins/instruction-placement/skills/migrate/scripts/remove-shims.test.sh +++ b/plugins/instruction-placement/skills/migrate/scripts/remove-shims.test.sh @@ -199,9 +199,50 @@ OUT=$(bash "$SCRIPT" --root "$READY" --installed-plugins "$TMP/installed-current --claude-bin "$TMP/bin/claude-met" "${CHECK_ARGS[@]}") || rc=$? assert_eq "no --confirm exits 1" 1 "$rc" assert_contains "and prints what removal costs" "$OUT" "InstructionsLoaded" +assert_contains "and prints the record's pointer" "$OUT" " - **Pointer**: " +assert_not_contains "and stops before the as-of date" "$OUT" "**As of**" assert_contains "and says it is not confirmed" "$OUT" "Not confirmed" assert_eq "and the repository is untouched" "" "$(cd "$READY" && git status --porcelain)" +# The price is the record's decision, the paragraph directly above its +# pointer, and the pointer itself: not the section's lead-in, not the as-of +# date or the trigger, and nothing from the next section. +PRICE_STUB="$TMP/price-stub" +mkdir -p "$PRICE_STUB/scripts" "$PRICE_STUB/reference" +cp "$SCRIPT" "$PRICE_STUB/scripts/remove-shims.sh" +cat >"$PRICE_STUB/reference/sources.md" <<'EOF' +# Sources + +## What shim removal costs + +A lead-in the price leaves out. + +The decision, first line, +and its second line. + +- **Pointer**: for the topic, see ; + a continuation line. +- **As of**: 2026-01-01 +- **Recheck trigger**: an event. + +## The next section + +Text the price never reaches. +EOF +EXPECTED_PRICE="$(printf '%s\n' \ + "What removing the shims costs (reference/sources.md, 'What shim removal costs'):" \ + " The decision, first line," \ + " and its second line." \ + "" \ + " - **Pointer**: for the topic, see ;" \ + " a continuation line.")" +OUT=$(bash "$PRICE_STUB/scripts/remove-shims.sh" --root "$READY" --installed-plugins \ + "$TMP/installed-current.json" --claude-bin "$TMP/bin/claude-met") +# The command substitution drops the blank line the script prints after it. +GOT_PRICE="$(printf '%s\n' "$OUT" | + awk '/^What removing the shims costs/ { f = 1 } /^Not confirmed/ { exit } f')" +assert_eq "the price is exactly the decision and its pointer" "$EXPECTED_PRICE" "$GOT_PRICE" + # --- Case 3: a stale installed plugin refuses before anything is removed -- rc=0 diff --git a/plugins/instruction-placement/skills/setup/SKILL.md b/plugins/instruction-placement/skills/setup/SKILL.md index 885e125ba7..6bee49557a 100644 --- a/plugins/instruction-placement/skills/setup/SKILL.md +++ b/plugins/instruction-placement/skills/setup/SKILL.md @@ -23,14 +23,16 @@ carrying both with no import between them gets a perfectly generated, perfectly never enters context, with every other gate green. Verifying that before a first audit is this skill's job; `/instruction-placement:check` asks the same question again on every gate run. -- **Claim**: Claude Code reads `AGENTS.md` as the project instructions only where there is no - `CLAUDE.md`, `.claude/CLAUDE.md` or `CLAUDE.local.md` in the working directory or above it; - reading it directly depends on a CLI version floor and on the session, both in - `skills/migrate/reference/sources.md`, "The minimum CLI version". -- **Basis**: [memory](https://code.claude.com/docs/en/memory), "AGENTS.md" and "When AGENTS.md - support is unavailable"; canary runs on Claude Code 2.1.278. -- **As of**: 2026-09-29. -- **Recheck trigger**: that section changes which file names count for the check, or a release note +This skill treats the root `AGENTS.md` as shadowed wherever a `CLAUDE.md`, `.claude/CLAUDE.md` or +`CLAUDE.local.md` sits in the working directory or above it, and treats direct `AGENTS.md` reading +as dependent on the CLI floor and session recorded in `skills/migrate/reference/sources.md`, "The +minimum CLI version". Canary runs on Claude Code 2.1.278 observed the shadowing. + +- **Pointer**: for when Claude Code reads `AGENTS.md` and when that support is unavailable, see + and + . +- **As of**: 2026-09-29 +- **Recheck trigger**: those sections change which file names shadow `AGENTS.md`, or a release note names `AGENTS.md` or instruction-file loading. Secondary warrants: `git` backs tracked-file discovery for nested instruction files, `node` launches diff --git a/plugins/knowledge/.claude-plugin/plugin.json b/plugins/knowledge/.claude-plugin/plugin.json index 1e2608f9e3..087185ba69 100644 --- a/plugins/knowledge/.claude-plugin/plugin.json +++ b/plugins/knowledge/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "knowledge", - "version": "0.14.17", + "version": "0.15.0", "description": "Ingest external knowledge into durable, synthesized artifacts. Ships a book-distillation pipeline (PDF/EPUB into concept-organized, author-attributed skill reference files), a video-digest pipeline (watch a single public video from YouTube or X, formerly Twitter: transcript, link harvest, and repo-applicability synthesis), a course-digest pipeline (extract and synthesize online video courses from Dometrain and Teachable into repo-applicable recommendations), a docpage-digest pipeline (single online documentation page into a verified knowledge slice with dual verification including one cross-vendor verifier, and an interview handoff), and a map-corpus pipeline (multi-resource corpus into a classified link map, deterministic node manifests, gate-verified relevance inventory, and an approved queue of docpage-digest runs), plus a re-runnable setup action; a configurable library directory governs where synthesized artifacts land in the consuming repo.", "author": { "name": "Melodic Software", @@ -38,7 +38,7 @@ "library_dir": { "type": "directory", "title": "Knowledge library directory", - "description": "Directory where synthesized knowledge artifacts land. Default is the consuming repo root; a relative value is resolved against the project directory. Portable non-project roots: an absolute path, a leading ~ (home-relative), or an environment-variable reference ${NAME} / %NAME% (e.g. ${KNOWLEDGE_CORPUS_DIR}) so a machine-varying root never needs a literal machine path in this stored value. A working-notes or artifacts convention declared in your own project's CLAUDE.md or rules takes precedence.", + "description": "Directory where synthesized knowledge artifacts land. Default is the root of the working tree the session is in (its worktree, when it has entered one); a relative value is resolved against that working tree. Portable non-project roots: an absolute path, a leading ~ (home-relative), or an environment-variable reference ${NAME} / %NAME% (e.g. ${KNOWLEDGE_CORPUS_DIR}) so a machine-varying root never needs a literal machine path in this stored value. A working-notes or artifacts convention declared in your own project's CLAUDE.md or rules takes precedence.", "default": "." }, "yt_dlp_js_runtimes": { diff --git a/plugins/knowledge/CHANGELOG.md b/plugins/knowledge/CHANGELOG.md index 5fad44babd..1aaa68bc51 100644 --- a/plugins/knowledge/CHANGELOG.md +++ b/plugins/knowledge/CHANGELOG.md @@ -4,6 +4,51 @@ All notable changes to the `knowledge` plugin are recorded here. The `version` i `.claude-plugin/plugin.json` is the delivery vehicle. A consumer receives a change only after that version increases. +## [0.15.0] - 2026-10-02 + +### Changed + +- **The `docpage-digest` Anthropic docs profile is stated as our decisions plus pointers.** The + docs hosts `platform.claude.com` and `code.claude.com` are pointer hosts; `claude.com/blog`, + `claude.dev/blog` and `anthropic.com/engineering` are correlate-only hosts that the upstream-drift + convention never accepts as a pointer, and the blog hosts share every blog rule. Rules that + quoted or paraphrased a page now say what we do and where the page is read. The pinned-versus-alias + model ID rule stores no generation rule and carries a pointer record. +- **The Anthropic docs queue lists blog posts as digest targets, never pointers.** Each post entry + names the docs page that serves as its pointer and keeps the post as a correlate, and the + queue's custody notes state our finding rather than the page's wording. +- **`docpage-digest` gains a claude.dev blog extractor and a pin-manifest script**, each with a + fixture-driven test suite. The extractor emits table separator rows, drops code-widget and + video-control labels, and never opens a code fence with a blank line; a paragraph inside a + list item keeps its bullet, one inside a table cell stays in the cell, and a table inside a list +item stays indented under it. The pin manifest's + `--check` reports a deleted file as `BLOCKED: missing`, even when its whole category is gone. +- **`docpage-digest` pipeline fixes.** The work root resolves against the session's worktree (the + `library_dir` description says so); the platform docs corpus is searched before a claim is + certified vendor-claimed or blog-only; one standard absence corpus; a split-applicability quote + rule; a tag-exempt blog-apparatus category; per-unit corrections files; an exact-byte write route + that the guardrails hook allows; no edits to the slice, the orchestrating session included, while + a verifier arm runs; the model-alias spawning gap and the Codex verifier's no-network limit are + recorded; the effort gotcha names where an agent's effort comes from; the hedge rule is + links-only. +- **The `docpage-digest` Anthropic profile tags blog rows by what they assert, and correction + rounds search twice.** A blog row-class table gives benchmark method, results and the author's + own test runs a vocabulary tag plus `unverified-inference` or `vendor-claimed` as their content + warrants, and leaves pure widget text tag-exempt. Figure data decoded from a framework payload is + a derived file read for stated values only, never numbers from drawn geometry. Each correction + round pairs its occurrence sweep per fixed class with a differently worded second search. +- docpage-digest's queue drops the two what's-new entries whose URLs now serve model overview + pages, names the model-config section "Work with Fable" with its link, and gives its no-docs- + page note a recheck trigger. The project-root rule and the Codex sandbox note state the + mechanism without a past-run anecdote. +- docpage-digest's effort gotcha points at the marketplace's record on where per-task effort is + set: Workflow's per-call option, or an agent's pin for an Agent tool dispatch. +- docpage-digest's verifier A runs as a Workflow `agent()` call at effort `high` by default, + overridable per run; a named agent keeps its own pin, and the verdict header records the + effective effort and its source. Without the Workflow tool it falls back to an Agent tool + dispatch, where effort is the agent's pin or the session level; below `medium`, Phase 4 stops + and reports the level instead of verifying. + ## [0.14.17] - 2026-10-02 ### Changed diff --git a/plugins/knowledge/README.md b/plugins/knowledge/README.md index c35f5e2a86..456b600731 100644 --- a/plugins/knowledge/README.md +++ b/plugins/knowledge/README.md @@ -85,7 +85,7 @@ defaults keep every pipeline working): | Option | Type | Default | Purpose | |---|---|---|---| -| `library_dir` | directory | `.` (repo root) | Directory where the plugin's ingestion pipelines land synthesized artifacts; a relative value resolves against the project directory. Portable non-project roots: an absolute path, a leading `~` (home-relative), or an env-var reference `${NAME}` / `%NAME%` (e.g. `${KNOWLEDGE_CORPUS_DIR}`) so a machine-varying root never needs a literal machine path in the stored value. Expanded when a pipeline resolves the root (the `video-digest` launcher and the `docpage-digest` work root today), failing loud on an unset variable. `book-distill` is unaffected. It writes to the target skill you name at invocation. A working-notes or artifacts convention declared in your own project's `CLAUDE.md` or rules takes precedence. | +| `library_dir` | directory | `.` (working-tree root) | Directory where the plugin's ingestion pipelines land synthesized artifacts; a relative value resolves against the root of the working tree the session is in, its worktree when it has entered one. Portable non-project roots: an absolute path, a leading `~` (home-relative), or an env-var reference `${NAME}` / `%NAME%` (e.g. `${KNOWLEDGE_CORPUS_DIR}`) so a machine-varying root never needs a literal machine path in the stored value. Expanded when a pipeline resolves the root (the `video-digest` launcher and the `docpage-digest` work root today), failing loud on an unset variable. `book-distill` is unaffected. It writes to the target skill you name at invocation. A working-notes or artifacts convention declared in your own project's `CLAUDE.md` or rules takes precedence. | | `yt_dlp_js_runtimes` | string | `node` | `video-digest`: JavaScript runtime yt-dlp uses for YouTube signature deciphering. Set to `off` to omit the flag entirely. | | `yt_dlp_cookies_file` | string | (empty) | `video-digest`: path to a Netscape cookies.txt for authenticated acquisition. Never commit cookie files. | | `yt_dlp_cookies_from_browser` | string | (empty) | `video-digest`: browser to pull YouTube cookies from (`chrome`, `firefox`, `edge`, …), forcing one instead of the automatic fallback. A cookies file wins over this. | @@ -108,7 +108,7 @@ reads it from. | Option | Type | Default | Environment variable | Description | | --- | --- | --- | --- | --- | -| `library_dir` | directory | `"."` | `CLAUDE_PLUGIN_OPTION_LIBRARY_DIR` | Directory where synthesized knowledge artifacts land. Default is the consuming repo root; a relative value is resolved against the project directory. Portable non-project roots: an absolute path, a leading ~ (home-relative), or an environment-variable reference ${NAME} / %NAME% (e.g. ${KNOWLEDGE_CORPUS_DIR}) so a machine-varying root never needs a literal machine path in this stored value. A working-notes or artifacts convention declared in your own project's CLAUDE.md or rules takes precedence. | +| `library_dir` | directory | `"."` | `CLAUDE_PLUGIN_OPTION_LIBRARY_DIR` | Directory where synthesized knowledge artifacts land. Default is the root of the working tree the session is in (its worktree, when it has entered one); a relative value is resolved against that working tree. Portable non-project roots: an absolute path, a leading ~ (home-relative), or an environment-variable reference ${NAME} / %NAME% (e.g. ${KNOWLEDGE_CORPUS_DIR}) so a machine-varying root never needs a literal machine path in this stored value. A working-notes or artifacts convention declared in your own project's CLAUDE.md or rules takes precedence. | | `yt_dlp_js_runtimes` | string | `"node"` | `CLAUDE_PLUGIN_OPTION_YT_DLP_JS_RUNTIMES` | JavaScript runtime yt-dlp uses for YouTube signature deciphering. Default 'node'. Set to 'off' to omit the --js-runtimes flag entirely. | | `yt_dlp_cookies_file` | string | *(none)* | `CLAUDE_PLUGIN_OPTION_YT_DLP_COOKIES_FILE` | Path to a Netscape-format cookies.txt for authenticated video acquisition (YouTube bot checks; the three login-required X cases). Empty by default (unauthenticated; YouTube adds an automatic browser-cookie fallback on a bot check, and X never iterates browser profiles). Never commit cookie files. | | `yt_dlp_cookies_from_browser` | string | *(none)* | `CLAUDE_PLUGIN_OPTION_YT_DLP_COOKIES_FROM_BROWSER` | Browser to pull cookies from (e.g. chrome, firefox, edge), forcing one instead of the automatic platform-ordered fallback. YouTube only, since the X adapter is cookies-file-only. Empty by default. A cookies file, when set, wins over this. | diff --git a/plugins/knowledge/skills/docpage-digest/SKILL.md b/plugins/knowledge/skills/docpage-digest/SKILL.md index 44eda9c31b..6bc8d33899 100644 --- a/plugins/knowledge/skills/docpage-digest/SKILL.md +++ b/plugins/knowledge/skills/docpage-digest/SKILL.md @@ -29,8 +29,17 @@ session reads it back instead of re-deriving it. Resolution rules for the render - **Unset**, an empty value, or a surviving literal `${user_config.library_dir}` token, means the option was never configured. Use the default `.`; never create a directory named after the token. -- **Relative** (including the default `.`). Resolve against the project directory, - `${CLAUDE_PROJECT_DIR}/`. +- **Relative** (including the default `.`). Resolve against the session's own working tree: + `git rev-parse --show-toplevel` run in the session's working directory, or that directory itself + outside git. Never resolve against the `CLAUDE_PROJECT_DIR` project root: inside a worktree + session it still names the main checkout, where isolation refuses the slice's writes. The slice stays + untracked either way, because the root self-ignores (below). + - **Pointer**: for where `CLAUDE_PROJECT_DIR` points after a session enters a worktree, see + ; for the write + refusal, see . + - **As of**: 2026-10-01 + - **Recheck trigger**: either section changes where `CLAUDE_PROJECT_DIR` points inside a + worktree session or which writes the isolation checks refuse. - **Absolute**. Use verbatim, with no project-directory prefix. - **Leading `~`**, the home directory, with no project-directory prefix. - **`${NAME}` / `%NAME%` env-var reference**. Read the variable yourself (`printenv NAME` in @@ -163,6 +172,17 @@ path stays inside `/digests/` before writing. conditionally. "this brief assumes model X; if you are not X, note the mismatch in your output and continue", never "you are X" as fact. - Each brief carries the untrusted-source rule and ONLY the source section plus SOURCES.md context, not this conversation. + When the matched profile defines applicability evidence rules, the brief names them, including + the corpora an absence or blog-only claim must search, so no agent picks its own search set. +- **Exact bytes:** the Edit and Write tools take their text through a JSON parameter, and a past + run found every agent edit writing a source's literal `\uXXXX` escape as the decoded + character, so do not use them for those bytes. Write those bytes with a script file run as + `python3