Skip to content

chore(harness-ops): validate the native-surface registry against Claude Code 2.1.288 - #5957

Merged
kyle-sexton merged 5 commits into
mainfrom
chore/native-drift-2-1-288
Oct 3, 2026
Merged

kyle-sexton merged 5 commits into
mainfrom
chore/native-drift-2-1-288

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

No related issue: every drift item the 2.1.288 native-drift pass found is resolved in this PR (keys listed under Verification), so none was filed.

Summary

Per-release native-surface pass for Claude Code 2.1.288. Every inventory lane extracts ok and the regex and parser readers agree, so the inventory is revalidated against 2.1.288. The one native surface that moved is the hidden, gated built-in /update, renamed /restart with update kept as an alias. Detect surfaced 14 overlap candidates; each is ruled here by operator decision (2026-10-02, all recommended). The Explore and Plan verification records also named too few disallowed tools.

Fix

  • VALIDATED_AGAINST in plugins/harness-ops/skills/inventory/scripts/inventory.py is now 2.1.288.
  • docs/native-surfaces/records.json: the two update dismissals (vs playbooks:update, firecrawl:update) are orphaned by the rename, so they are removed and re-recorded against restart. Twelve more dismissals:
    • Three resurfaced pairs re-dismissed with their original reason: claude-test vs prototype:pressure-test and testing:test-value (native description now "Run tests on this machine (local)"), and code-review vs review:security-review (native gained --max-findings).
    • Shared-word pairs: claude-test vs mutation-testing:audit, rename vs docs-hygiene:rename-references, agents vs multi-agent:assess and multi-agent:route, workflow-subagent vs multi-agent:assess, and worker vs discovery:sweep-worker.
    • Agent vs multi-agent:assess and multi-agent:route: ours decide how the Agent tool is used and launch nothing.
    • cc-plugin-claude-test vs testing:run-e2e: covered by the existing claude-test -> testing:run-e2e defer row.
  • docs/native-surfaces.md regenerated with overlap.py generate; node scripts/generate-catalog.mjs reported the catalog already in sync.
  • plugins/discovery/skills/explore/reference/native-explore.md and plugins/planning/skills/plan/reference/native-plan-agent.md: on 2.1.288 both agents disallow Agent, Artifact, ArtifactComments, ArtifactData, ArtifactCheck, ExitPlanMode, Edit, Write and NotebookEdit (source literal). The records named four (Explore) and five (Plan) tools from the 2.1.285 extraction; they now state what that means for each skill and point at builtin_agents.<agent>.disallowed_tools instead of copying the list (Codex review).
  • Versions: harness-ops 2.5.2 -> 2.5.3, discovery 0.26.1 -> 0.26.2, planning 0.62.1 -> 0.62.2 (main took the earlier numbers while this PR was open), each with a CHANGELOG entry.

Verification

  • Rename confirmed by string literal in the binaries: name:"update" occurs once in 2.1.287 and 0 times in 2.1.288; name:"restart" 0 then 1. The 2.1.288 extraction reads restart with aliases: ["update"], hidden: true, gated: true.
  • inventory.py --self-check: DEGRADED: cli 2.1.288, validated against 2.1.287 (exit 3, all seven lanes ok) before the bump; OK: cli 2.1.288, validated against 2.1.288 after.
  • inventory.py --reader compare --self-check: reader compare: ok, 2167 of 2167 modules parse: no value->value divergence. A plain --reader compare run shows 4 wrong->unresolved entries: the parser reads the Explore and Plan disallowed_tools as partial, missing the four Artifact tools, as on 2.1.284-2.1.287 (feat(harness-ops): read Explore/Plan disallowed_tools literal under the parser reader, then flip the inventory default #5901). This PR records the literal regex values, which the live Plan agent roster confirms.
  • native_drift.py diff against the 2.1.287 summary: surfaces added/removed/reclassified none, renamed update -> restart; invocability and markers none; fired triggers none; unresolved descriptions unchanged (design). The 15 items: the 14 native-drift:candidate:* keys above plus native-drift:inventory-degraded:2.1.288:inventory.
  • overlap.py detect on the new store: discovered 0, resurfaced 0, orphaned dismissals 0, 120 suppressed.
  • overlap.py self-check: exit 3, only the 2 standing advisories (older recorded extraction versions on rows; upstream SHA not locally decidable). 71 rows checked.
  • TestInstalledBuilds now covers 2.1.288 (passed locally with acorn, 88.8 s); extraction.md names 2.1.288 in its record.
  • Tests: inventory 362 OK (161 skipped without acorn), audit-native-overlap 195 OK, changelog 40 OK. Pinned ruff on inventory.py: all checks passed.

Related

🤖 Generated with Claude Code

kyle-sexton and others added 2 commits October 2, 2026 19:27
…de Code 2.1.288

Revalidate the inventory against 2.1.288, rule the 14 overlap candidates
the pass surfaced, re-key the two /update dismissals to /restart, and
correct the Explore and Plan disallowed-tool lists in their verification
records.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1-288

# Conflicts:
#	plugins/harness-ops/CHANGELOG.md
@kyle-sexton
kyle-sexton marked this pull request as ready for review October 2, 2026 23:35
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T00:02:56.436974Z 0fc5959 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 19s —— View job


/review:code-review — PR #5957

  • Gather context (PR diff, description, changed files)
  • Invoke /review:code-review per custom instructions
  • Post findings as inline PR comments
  • Summarize in this comment

No blocking findings.

This PR is a per-release native-surface data pass: records.json dismissal/defer rows, the
generated docs/native-surfaces.md table, two reference-doc corrections, three plugin version
bumps + CHANGELOGs, and a one-line constant bump in inventory.py
(VALIDATED_AGAINST).
I checked it for the kind of correctness bugs a data/registry PR like this is prone to:

  • records.json fingerprint consistency: every repeated native surface (restart,
    claude-test, agents, Agent, cc-plugin-claude-test) carries one matching fingerprint.native
    hash across its rows, and each row's fingerprint.component hash matches its pre-existing
    component rows where the component itself is unchanged (e.g. the restart rows at
    records.json#L292-325
    reuse the exact component fingerprints the removed update rows had).
  • Orphaned-dismissal handling for the /update → /restart rename: the two old update rows
    are removed and re-recorded against restart, not left as dead entries — matches the PR's claim.
  • The new cc-plugin-claude-test → testing:run-e2e dismissal (records.json L3766-3783) is a
    dismissal (no verdict key), not a verdict row, so it doesn't collide with the existing defer
    verdict row for claude-test → testing:run-e2e it says it's "carried by" — confirmed that row
    exists at L2820-2850 with "verdict": "defer" and a distinct pair key (different native.name),
    so overlap.py's dismissal/verdict mutual-exclusion check (scripts/overlap.py L896-915) isn't
    violated.
  • CHANGELOG tool-count arithmetic: discovery's CHANGELOG claims 5 newly-named Explore
    disallowed tools (Artifact, ArtifactComments, ArtifactData, ArtifactCheck, ExitPlanMode) and
    planning's claims 4 for Plan (same four, since Plan already listed ExitPlanMode) — both
    check out against the actual before/after tool lists in
    native-explore.md and native-plan-agent.md.
  • Version/CHANGELOG bumps (discovery 0.26.0→0.26.1, harness-ops 2.5.1→2.5.2, planning
    0.62.0→0.62.1) each have a matching dated CHANGELOG entry; marketplace.json doesn't duplicate
    plugin versions, so no second place needed updating.
  • No stale prose elsewhere in the repo still refers to the native /update command by its old
    name (the many other update.* hits are unrelated plugin-internal update skills —
    firecrawl:update, playbooks:update, context7, dometrain — not the renamed CLI command).

Note: I could not execute inventory.py --self-check, overlap.py detect/self-check, or the test
suites in this sandbox (Bash tool access is restricted here), so the PR body's verification output
(self-check results, reader-compare counts, test pass counts) is author-claimed and unverified
by me — I did not re-run it.

This is a well-scoped, internally consistent registry update with no code-logic surface beyond the
one-line constant bump. Nothing here rises to a reviewer-blocking finding.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 15s —— View job


Security review (/review:security-review)

  • Gather context (PR diff, changed files)
  • Run /review:security-review criteria against the diff
  • Check instruction-surface deletions against the exception register
  • Post findings (inline + summary) or report clean

No security issues found.

This PR is a per-release data/documentation update for the native-surface overlap registry:

  • docs/native-surfaces/records.json and the generated docs/native-surfaces.md: adds/edits
    human-ruled overlap-dismissal rows (including re-recording the update → restart rename) and
    CHANGELOG entries. No executable logic.
  • plugins/harness-ops/skills/inventory/scripts/inventory.py: a single constant bump
    (VALIDATED_AGAINST = "2.1.288"). No parsing, execution, or I/O behavior changes.
  • plugins/discovery/skills/explore/reference/native-explore.md and
    plugins/planning/skills/plan/reference/native-plan-agent.md: documentation updates that
    expand the recorded disallowed_tools list for the Explore/Plan built-in agents (adding
    Artifact, ArtifactComments, ArtifactData, ArtifactCheck, ExitPlanMode) to match the live
    2.1.288 extraction. These are factual reference records, not enforced guardrails, and the change
    only tightens/corrects the documented restriction set — it doesn't narrow or remove one, so the
    instruction-surface-deletion lens in the security-review skill doesn't flag anything here.
  • plugins/*/.claude-plugin/plugin.json + CHANGELOG.md: routine version bumps.

No authN/authZ logic, injection surfaces, credential/token handling, or Actions workflow changes
are touched by this diff. Nothing met the high-signal bar for a finding.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 63eb7bc293

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/discovery/skills/explore/reference/native-explore.md Outdated
Comment thread plugins/harness-ops/skills/inventory/scripts/inventory.py
Comment thread docs/native-surfaces/records.json
kyle-sexton and others added 2 commits October 2, 2026 19:44
…1-288

# Conflicts:
#	plugins/planning/CHANGELOG.md
… list and pin 2.1.288

The verification records state what the disallowed tools mean for each
skill and point at builtin_agents.<agent>.disallowed_tools instead of
copying a list that changes per release. The installed-build regression
now covers 2.1.288, and the extraction record names it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 22s —— View job


Security review (/review:security-review) — PR #5957

  • Gather context (PR diff, description, changed files)
  • Invoke /review:security-review per custom instructions
  • Check instruction-surface deletions against the exception register
  • Post findings (inline + summary) or report clean

No security issues found.

This is a per-release data/documentation pass over the native-surface overlap registry, reviewed
against the full PR diff (origin/main...HEAD, head 3f2d4e0) across all 13 changed files:

No authN/authZ logic, injection surfaces, credential/token handling, or Actions workflow changes are
touched by this diff. This confirms and extends the earlier security-review pass on 63eb7bc (ready-for-review commit): the one commit added since then (3f2d4e0) is itself doc/test-only and introduces
no new security-relevant surface. Nothing met the high-signal bar for a finding.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 21s —— View job


/review:code-review — PR #5957

  • Gather context (PR diff, description, changed files)
  • Invoke /review:code-review per custom instructions
  • Post findings as inline PR comments
  • Summarize in this comment

No blocking findings.

This is a per-release native-surface data pass (records.json rows, the generated catalog, two
reference-doc corrections, three version bumps). The latest commit (3f2d4e0) already resolves
the two actionable Codex findings from the prior review round:

  • The P1 "replace volatile tool rosters with live pointers" finding: native-explore.md
    and native-plan-agent.md now say what the tool denial means for the skill ("it cannot edit
    files or spawn an agent" / "...or exit plan mode") and point at
    builtin_agents.Explore.disallowed_tools / builtin_agents.Plan.disallowed_tools "in a fresh
    run" instead of enumerating the roster — see
    native-explore.md#L9
    and
    native-plan-agent.md#L9.
  • The P2 "add 2.1.288 to the installed-build regression" finding:
    test_reader_findings.py#L706
    now iterates 2.1.284 through 2.1.288.

I independently re-derived the data-consistency checks the earlier Claude pass reported rather
than re-trusting them: walked records.json for the Agent/agents, claude-test,
cc-plugin-claude-test, restart, and worker rows and confirmed each repeated native surface
shares one fingerprint.native hash across its rows, the new cc-plugin-claude-test →
testing:run-e2e dismissal and the pre-existing claude-test → testing:run-e2e verdict
row (defer) key on different native names so there's no dismissal/verdict collision, and every
table section in docs/native-surfaces.md stays alphabetically ordered by component after the
inserts. No fingerprint or ordering bug found.

Two things that are not blocking:

  • The remaining open Codex finding (restart dismissals don't fingerprint the /update alias, only
    scored_text's description/argument_hint/search_hint —
    discover.py#L133-L138,
    consumed by fingerprints() in
    overlap.py#L1613-L1622)
    is accurate, but it's pre-existing fingerprinting behavior this PR doesn't touch, not a defect
    introduced here — out of scope for this lane per the "pre-existing issues on untouched lines"
    exclusion.
  • The PR body's "Versions" line says planning 0.62.0 -> 0.62.1, but the actual diff and
    CHANGELOG bump planning from 0.62.1 to 0.62.2
    (plugin.json,
    CHANGELOG.md#L6) —
    almost certainly because main had already taken 0.62.1 by the time this branch's merge
    commit landed. Worth a one-line body edit for accuracy, but it's PR prose, not a code defect.

No inline comments posted: nothing in the diff met the high-signal bar for a new, blocking,
code-level finding.

@kyle-sexton

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Nice work!

Reviewed commit: 3f2d4e0835

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

…1-288

# Conflicts:
#	plugins/discovery/CHANGELOG.md
#	plugins/harness-ops/CHANGELOG.md
@kyle-sexton

Copy link
Copy Markdown
Contributor Author

@codex review

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 55s —— View job


Security review (/review:security-review) — PR #5957

  • Gather context (PR diff, description, changed files)
  • Invoke /review:security-review per custom instructions
  • Check instruction-surface deletions against the exception register
  • Post findings (inline + summary) or report clean

No security issues found.

Reviewed the full diff origin/main...HEAD (head 0fc5959) across all 13 changed files. This is a per-release data/documentation pass over the native-surface overlap registry:

  • docs/native-surfaces/records.json and the generated docs/native-surfaces.md: add/edit human-ruled overlap-dismissal rows (new Agent/agents/claude-test/rename/restart/worker/workflow-subagent dismissals, the /update → /restart rename re-record, and the cc-plugin-claude-test carried-dismissal row) with reasons, fingerprints, and dates. Pure data, no executable logic, no secrets or credentials.
  • plugins/harness-ops/skills/inventory/scripts/inventory.py#L65: single constant bump, VALIDATED_AGAINST = "2.1.288". No parsing, execution, or I/O behavior change.
  • plugins/harness-ops/skills/inventory/scripts/test_reader_findings.py#L706: extends the installed-build regression's version tuple to include "2.1.288". Test-only.
  • plugins/discovery/skills/explore/reference/native-explore.md and plugins/planning/skills/plan/reference/native-plan-agent.md: checked under the instruction-surface-deletion lens (docs/conventions/instruction-exception-register/README.md, present in this checkout and consulted) because these verification-record rows change how a tool-denial list is stated. They don't delete, narrow, or soften a guardrail — the native binary enforces the real Explore/Plan tool denials regardless of this doc's wording, and the row now states the consequence ("it cannot edit files or spawn an agent" / "...or exit plan mode") plus points at the live builtin_agents.<agent>.disallowed_tools instead of restating the roster. No protected-class match.
  • plugins/*/.claude-plugin/plugin.json + CHANGELOG.md: routine version bumps and changelog entries.

No authN/authZ logic, injection surfaces, credential/token handling, or Actions workflow changes are touched by this diff. Nothing met the high-signal bar for a finding.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 58s —— View job


/review:code-review — PR #5957

  • Gather context (PR diff, description, changed files)
  • Invoke /review:code-review per custom instructions
  • Post findings as inline PR comments
  • Summarize in this comment

No blocking findings.

This head (0fc5959) is a merge of origin/main into the branch already reviewed twice by this
lane (at 63eb7bc and 3f2d4e0, both "no blocking findings", plus Codex clean at 3f2d4e0). I
diffed 3f2d4e0..0fc5959 to find what the merge actually changed and it's purely mechanical:
origin/main had independently bumped discovery and harness-ops to the same patch versions
this branch was using, so the merge re-bumped this branch's plugin.json/CHANGELOG.md one patch
further (discovery 0.26.1→0.26.2, harness-ops 2.5.2→2.5.3; planning was untouched, no
collision there). I confirmed the version string in every bumped plugin.json matches its
CHANGELOG.md header exactly, and that no other file changed between the two commits — so the
substantive content (records.json, the generated catalog, and the native-explore.md /
native-plan-agent.md reference fixes) is identical to what was already reviewed at 3f2d4e0.

I re-walked the full origin/main...HEAD diff myself rather than trusting that conclusion blindly:

  • docs/native-surfaces/records.json: every inserted dismissal row's fingerprint.native is
    shared across the pair-rows that share a native surface (restart, claude-test, Agent/agents,
    cc-plugin-claude-test, workflow-subagent), and fingerprint.component is shared across
    rows for the same component (multi-agent:assess's fingerprint is identical on its Agent,
    agents, and workflow-subagent rows). The /update→/restart rename correctly removes the
    two old update rows rather than orphaning them, and re-records against the new name with a new
    native fingerprint (description text changed, so the hash correctly changed too).
  • docs/native-surfaces.md: every new/moved row lands in the correct alphabetical slot within its
    table (spot-checked restart landing between rename and resume, workflow-subagent between
    workflow-authoring and workflows, claude-test's mutation-testing:audit row sorting before
    prototype:pressure-test).
  • inventory.py / test_reader_findings.py / extraction.md: the VALIDATED_AGAINST bump,
    the installed-build regression loop extension to 2.1.288, and the as-of annotation update are
    all consistent with each other and with the CHANGELOG's own description of them.
  • native-explore.md / native-plan-agent.md: both now state the consequence of the disallowed
    tools (cannot edit files, spawn an agent, or exit plan mode) and point at
    builtin_agents.<agent>.disallowed_tools instead of re-copying the volatile roster — this is the
    fix for the earlier Codex P1 finding on this PR, and it reads correctly against
    .claude/rules/skill-bodies-state-current-rules.md's pointer-plus-recheck-trigger convention.

On the one disputed Codex finding (restart dismissals not fingerprinting the /update alias,
records.json#L4717):
the PR author's rebuttal is sound — scored_text() in
discover.py#L123-L138
never reads aliases, so if a future release drops the /update alias, restart loses its only
shared token with firecrawl:update/playbooks:update and simply stops scoring as a candidate —
the dismissal can't suppress a real future overlap under that scenario. This is also pre-existing
fingerprinting behavior this PR doesn't change, so it's out of scope for this lane regardless.

No inline comments posted: nothing in the diff meets the high-signal bar for a new, blocking,
code-level finding that the two prior passes on this content (plus my own re-derivation above)
didn't already cover or that wasn't already fixed in 3f2d4e0.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Breezy!

Reviewed commit: 0fc595998f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@kyle-sexton
kyle-sexton merged commit 0100a02 into main Oct 3, 2026
24 checks passed
@kyle-sexton
kyle-sexton deleted the chore/native-drift-2-1-288 branch October 3, 2026 00:03
kyle-sexton added a commit that referenced this pull request Oct 3, 2026
… parser reader's flow check (#5970)

Refs: #5901

This PR leaves the issue open: the default reader is not flipped.

## Summary

#5901 asked for the parser reader to read the Explore and Plan
`disallowed_tools` literal on 2.1.284-2.1.288 and then become the
inventory default. Two causes kept both lists partial: whole-module
(namespace) loads of the re-exporting chunk, and the sink rule. This PR
removes the first and narrows the second. Under the brief's stop rule it
does not flip the default: the sink rule still fires on every build, and
some of what it fires on cannot be cleared without assuming something
the analysis cannot prove (details under Fix).

- **Namespace loads are followed (cause 1 removed).** The helper's new
`namespace` op follows each `import()`, `require()`,
`import.meta.require()` and `import*as` load of the exporting file to
its reads. It accepts only reads of other exports by name: a member read
that is not a call, an object pattern without rest, a `{names:ns}`
record read only by name or tested, and `await Promise.all([...])`
destructured by an array pattern. The built-ins those shapes rely on
join the trusted names the sink rule watches. A global write to a
trusted name is now a sink. A namespace settled through a promise needs
its module to export no `then`. On every installed build, all 13-14
loads of the re-exporting chunk pass.
- **node:vm counts as code built from a string.** Any `vm`/`node:vm`
load except an import naming only `isContext` is a sink like `eval` and
`Function`, and so is a vm runner's name (`runInThisContext`,
`runInNewContext`, `runInContext`, `compileFunction`,
`SourceTextModule`, `SyntheticModule`) read from any object. Code in a
new context still reaches this realm's prototypes through
`this.constructor.constructor`.
- **A load the parser cannot name fails closed (verifier gap 1, also on
main).** Some loads could pull in the exporting file whole without the
parser seeing which file: an aliased `require` or `import.meta.require`,
`.call`, a comma callee, or `import(x)` with no literal specifier. The
`loads` op now reports any module that holds one, and that module fails
every export hop.
- **The sink rule reads computed keys (cause 2 narrowed).** A
computed-key write or define counts only for the trusted names its key
can spell when every value of the key is known: literals, numbers,
boolean or `typeof` results, or variables written only with those.
- harness-ops 3.0.1 -> 3.1.0 (minor). `--reader` stays `regex`; README,
SKILL.md and native-drift.md are unchanged because the default did not
change.

## Fix

What still blocks a literal read, measured on 2.1.288 with the flow's
own trusted names (`some`, `includes`, `has`). Module counts are lower
bounds, because each module reports at most 20 hits:

| Sink kind | Modules | Why it cannot be cleared here |
|---|---|---|
| computed-key write on a target not shown fresh | 178 (201 before key
provenance) | Receivers are parameters, `this` in methods, call results;
clearing them needs interprocedural points-to over the 40 MB bundle |
| computed-key define (`Object.defineProperty(o,k,...)`) | 92 | Same |
| prototype swap (`__proto__=`, `setPrototypeOf`) | 22 | Same |
| definer read other than as a callee
(`Hn=Object.defineProperty.bind(Object)`) | 11 | Aliased definers can be
called with any key |
| node:vm load or runner | 5 | The workflow runtime, plugin loader and a
test kit run code from strings through vm |
| (namespace stage) a load the parser cannot name | 3 | Each could load
the re-exporting chunk whole unseen |
| code built from a string | 8 | ajv runs validator code it generates at
runtime (`Function(self,scope,code)(...)`), protobufjs's `inquire` makes
a direct `eval` call, and 3 modules call a member `.eval(...)` whose
receiver cannot be shown to be something other than the global object.
Clearing ajv's case would mean assuming generated code never patches a
built-in prototype. That assumption is unsound |

The namespace stage needed no unsound assumption. The sink stage does,
for ajv at least.

## Verification

- `INVENTORY_REQUIRE_ACORN=1` full harness-ops suite: all 13 `test_*.py`
modules (`python3 -m unittest` per directory) and 39 of the 40
`*.test.sh` pass. `audit-install-state/scripts/install_state.test.sh`
fails here and on origin/main the same way: under Python 3.14, `python
-m unittest <absolute path>.py` fails to import the module. This PR does
not touch that skill, and `test_install_state.py` passes when run as a
module.
- New tests in `test_reader_findings.py`:
  - `test_a_namespace_read_only_by_name_keeps_the_literal` (7 shapes).
- `test_a_namespace_read_other_than_by_name_stays_partial` (adversarial,
one or more per acceptance rule).
- `test_what_a_namespace_load_trusts_stays_checked` (a replaced
`Promise`, `Promise.resolve`, `then`, `Promise.prototype.constructor`,
an `Object.prototype` getter, a `then` export, an `export*`).
- `test_a_computed_key_known_to_name_no_trusted_name_clears` and
`test_a_computed_key_that_may_name_a_trusted_name_stays_a_sink`.
- `test_a_bundle_calling_vm_run_in_this_context_stays_partial` covers 11
spellings: `runInThisContext`, `runInNewContext` (the verifier's probe),
`runInContext`, `Script#runInNewContext`, `SourceTextModule`, a
destructured `compileFunction`, a computed read and an escaping alias. A
plain named import of `isContext` stays literal.
- `test_a_load_the_parser_cannot_name_stays_partial` covers the
verifier's gap-1 probes, which read a wrong literal before this change:
an aliased `import.meta.require`, `require.call`, `(0,require)`,
`import.meta.require.call`, `import(s)` and `require(s)`. As controls,
`typeof require`, `require.resolve`, a `require` parameter and a
`{require:1}` key stay literal.
- At f2b395c, after both gap fixes:
- `INVENTORY_REQUIRE_ACORN=1` unittest passes for all 13 harness-ops
Python modules.
  - `scripts/run-ruff.sh check` and `format --check` are clean.
- `--reader compare --self-check` on 2.1.288 prints `OK`, with `reader
compare: ok, 2167 of 2167 modules parse`.
- `--reader compare --self-check` on 2.1.284 prints `reader compare: ok,
2151 of 2151 modules parse`. The overall verdict is `DEGRADED`, the same
as on main: there is a version advisory, and the builtin_plugins
canaries are absent on that build.
- The `{names:...}` case in
`test_the_module_table_names_the_exporters_own_file` now reads literal
when the record is never read, and partial with a computed read.
- `python3 inventory.py --binary <build> --reader compare --binary-only`
on this branch (rerun at f2b395c) and on origin/main 57ca27a:

| Build | Explore literal | Plan literal | value->value |
wrong->unresolved | Parser output vs main |
|---|---|---|---|---|---|
| 2.1.284 | no (partial) | no (partial) | 0 | 4 | identical |
| 2.1.285 | no (partial) | no (partial) | 0 | 4 | identical |
| 2.1.286 | no (partial) | no (partial) | 0 | 4 | identical |
| 2.1.287 | no (partial) | no (partial) | 0 | 4 | identical |
| 2.1.288 | no (partial) | no (partial) | 0 | 4 | identical |

The 4 wrong->unresolved entries are the Explore and Plan
`disallowed_tools` and `disallowed_tools_source` on every build, the
same as on main.
- `scripts/run-ruff.sh check` and `format --check` are clean;
`markdownlint-cli2` is clean on the changed docs;
`scripts/check-changelog-parity.sh --check-bump origin/main` passes.

## Related

- #5640 (parser reader plan and the decisions behind it)
- #5891 (P4: cross-module flow and the sink rule)
- #5957 (2.1.288 validation that recorded the partial Explore/Plan
lists)

Known gaps, unchanged by this PR:
- A spread's `Symbol.iterator` and the array iterator's `next` are not
trusted names, so the sink rule does not watch them for a spread.
- `require` and `import()` are recognized by name only. A shadowing
local binding, or a loader reached through another global such as
`module.require` or a `createRequire` result, is caught only when it is
aliased or called with a non-literal.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant