chore(harness-ops): validate the native-surface registry against Claude Code 2.1.288 - #5957
Conversation
…de Code 2.1.288 Revalidate the inventory against 2.1.288, rule the 14 overlap candidates the pass surfaced, re-key the two /update dismissals to /restart, and correct the Explore and Plan disallowed-tool lists in their verification records. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1-288 # Conflicts: # plugins/harness-ops/CHANGELOG.md
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Claude finished @kyle-sexton's task in 2m 19s —— View job
|
|
Claude finished @kyle-sexton's task in 1m 15s —— View job Security review (
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 63eb7bc293
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…1-288 # Conflicts: # plugins/planning/CHANGELOG.md
… list and pin 2.1.288 The verification records state what the disallowed tools mean for each skill and point at builtin_agents.<agent>.disallowed_tools instead of copying a list that changes per release. The installed-build regression now covers 2.1.288, and the extraction record names it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 1m 22s —— View job Security review (
|
|
Claude finished @kyle-sexton's task in 2m 21s —— View job
|
|
@codex review |
|
Codex Review: Didn't find any major issues. Nice work! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
…1-288 # Conflicts: # plugins/discovery/CHANGELOG.md # plugins/harness-ops/CHANGELOG.md
|
@codex review |
|
Claude finished @kyle-sexton's task in 55s —— View job Security review (
|
|
Claude finished @kyle-sexton's task in 1m 58s —— View job
|
|
Codex Review: Didn't find any major issues. Breezy! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
… parser reader's flow check (#5970) Refs: #5901 This PR leaves the issue open: the default reader is not flipped. ## Summary #5901 asked for the parser reader to read the Explore and Plan `disallowed_tools` literal on 2.1.284-2.1.288 and then become the inventory default. Two causes kept both lists partial: whole-module (namespace) loads of the re-exporting chunk, and the sink rule. This PR removes the first and narrows the second. Under the brief's stop rule it does not flip the default: the sink rule still fires on every build, and some of what it fires on cannot be cleared without assuming something the analysis cannot prove (details under Fix). - **Namespace loads are followed (cause 1 removed).** The helper's new `namespace` op follows each `import()`, `require()`, `import.meta.require()` and `import*as` load of the exporting file to its reads. It accepts only reads of other exports by name: a member read that is not a call, an object pattern without rest, a `{names:ns}` record read only by name or tested, and `await Promise.all([...])` destructured by an array pattern. The built-ins those shapes rely on join the trusted names the sink rule watches. A global write to a trusted name is now a sink. A namespace settled through a promise needs its module to export no `then`. On every installed build, all 13-14 loads of the re-exporting chunk pass. - **node:vm counts as code built from a string.** Any `vm`/`node:vm` load except an import naming only `isContext` is a sink like `eval` and `Function`, and so is a vm runner's name (`runInThisContext`, `runInNewContext`, `runInContext`, `compileFunction`, `SourceTextModule`, `SyntheticModule`) read from any object. Code in a new context still reaches this realm's prototypes through `this.constructor.constructor`. - **A load the parser cannot name fails closed (verifier gap 1, also on main).** Some loads could pull in the exporting file whole without the parser seeing which file: an aliased `require` or `import.meta.require`, `.call`, a comma callee, or `import(x)` with no literal specifier. The `loads` op now reports any module that holds one, and that module fails every export hop. - **The sink rule reads computed keys (cause 2 narrowed).** A computed-key write or define counts only for the trusted names its key can spell when every value of the key is known: literals, numbers, boolean or `typeof` results, or variables written only with those. - harness-ops 3.0.1 -> 3.1.0 (minor). `--reader` stays `regex`; README, SKILL.md and native-drift.md are unchanged because the default did not change. ## Fix What still blocks a literal read, measured on 2.1.288 with the flow's own trusted names (`some`, `includes`, `has`). Module counts are lower bounds, because each module reports at most 20 hits: | Sink kind | Modules | Why it cannot be cleared here | |---|---|---| | computed-key write on a target not shown fresh | 178 (201 before key provenance) | Receivers are parameters, `this` in methods, call results; clearing them needs interprocedural points-to over the 40 MB bundle | | computed-key define (`Object.defineProperty(o,k,...)`) | 92 | Same | | prototype swap (`__proto__=`, `setPrototypeOf`) | 22 | Same | | definer read other than as a callee (`Hn=Object.defineProperty.bind(Object)`) | 11 | Aliased definers can be called with any key | | node:vm load or runner | 5 | The workflow runtime, plugin loader and a test kit run code from strings through vm | | (namespace stage) a load the parser cannot name | 3 | Each could load the re-exporting chunk whole unseen | | code built from a string | 8 | ajv runs validator code it generates at runtime (`Function(self,scope,code)(...)`), protobufjs's `inquire` makes a direct `eval` call, and 3 modules call a member `.eval(...)` whose receiver cannot be shown to be something other than the global object. Clearing ajv's case would mean assuming generated code never patches a built-in prototype. That assumption is unsound | The namespace stage needed no unsound assumption. The sink stage does, for ajv at least. ## Verification - `INVENTORY_REQUIRE_ACORN=1` full harness-ops suite: all 13 `test_*.py` modules (`python3 -m unittest` per directory) and 39 of the 40 `*.test.sh` pass. `audit-install-state/scripts/install_state.test.sh` fails here and on origin/main the same way: under Python 3.14, `python -m unittest <absolute path>.py` fails to import the module. This PR does not touch that skill, and `test_install_state.py` passes when run as a module. - New tests in `test_reader_findings.py`: - `test_a_namespace_read_only_by_name_keeps_the_literal` (7 shapes). - `test_a_namespace_read_other_than_by_name_stays_partial` (adversarial, one or more per acceptance rule). - `test_what_a_namespace_load_trusts_stays_checked` (a replaced `Promise`, `Promise.resolve`, `then`, `Promise.prototype.constructor`, an `Object.prototype` getter, a `then` export, an `export*`). - `test_a_computed_key_known_to_name_no_trusted_name_clears` and `test_a_computed_key_that_may_name_a_trusted_name_stays_a_sink`. - `test_a_bundle_calling_vm_run_in_this_context_stays_partial` covers 11 spellings: `runInThisContext`, `runInNewContext` (the verifier's probe), `runInContext`, `Script#runInNewContext`, `SourceTextModule`, a destructured `compileFunction`, a computed read and an escaping alias. A plain named import of `isContext` stays literal. - `test_a_load_the_parser_cannot_name_stays_partial` covers the verifier's gap-1 probes, which read a wrong literal before this change: an aliased `import.meta.require`, `require.call`, `(0,require)`, `import.meta.require.call`, `import(s)` and `require(s)`. As controls, `typeof require`, `require.resolve`, a `require` parameter and a `{require:1}` key stay literal. - At f2b395c, after both gap fixes: - `INVENTORY_REQUIRE_ACORN=1` unittest passes for all 13 harness-ops Python modules. - `scripts/run-ruff.sh check` and `format --check` are clean. - `--reader compare --self-check` on 2.1.288 prints `OK`, with `reader compare: ok, 2167 of 2167 modules parse`. - `--reader compare --self-check` on 2.1.284 prints `reader compare: ok, 2151 of 2151 modules parse`. The overall verdict is `DEGRADED`, the same as on main: there is a version advisory, and the builtin_plugins canaries are absent on that build. - The `{names:...}` case in `test_the_module_table_names_the_exporters_own_file` now reads literal when the record is never read, and partial with a computed read. - `python3 inventory.py --binary <build> --reader compare --binary-only` on this branch (rerun at f2b395c) and on origin/main 57ca27a: | Build | Explore literal | Plan literal | value->value | wrong->unresolved | Parser output vs main | |---|---|---|---|---|---| | 2.1.284 | no (partial) | no (partial) | 0 | 4 | identical | | 2.1.285 | no (partial) | no (partial) | 0 | 4 | identical | | 2.1.286 | no (partial) | no (partial) | 0 | 4 | identical | | 2.1.287 | no (partial) | no (partial) | 0 | 4 | identical | | 2.1.288 | no (partial) | no (partial) | 0 | 4 | identical | The 4 wrong->unresolved entries are the Explore and Plan `disallowed_tools` and `disallowed_tools_source` on every build, the same as on main. - `scripts/run-ruff.sh check` and `format --check` are clean; `markdownlint-cli2` is clean on the changed docs; `scripts/check-changelog-parity.sh --check-bump origin/main` passes. ## Related - #5640 (parser reader plan and the decisions behind it) - #5891 (P4: cross-module flow and the sink rule) - #5957 (2.1.288 validation that recorded the partial Explore/Plan lists) Known gaps, unchanged by this PR: - A spread's `Symbol.iterator` and the array iterator's `next` are not trusted names, so the sink rule does not watch them for a spread. - `require` and `import()` are recognized by name only. A shadowing local binding, or a loader reached through another global such as `module.require` or a `createRequire` result, is caught only when it is aliased or called with a non-literal. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
No related issue: every drift item the 2.1.288 native-drift pass found is resolved in this PR (keys listed under Verification), so none was filed.
Summary
Per-release native-surface pass for Claude Code 2.1.288. Every inventory lane extracts ok and the regex and parser readers agree, so the inventory is revalidated against 2.1.288. The one native surface that moved is the hidden, gated built-in
/update, renamed/restartwithupdatekept as an alias. Detect surfaced 14 overlap candidates; each is ruled here by operator decision (2026-10-02, all recommended). The Explore and Plan verification records also named too few disallowed tools.Fix
VALIDATED_AGAINSTinplugins/harness-ops/skills/inventory/scripts/inventory.pyis now2.1.288.docs/native-surfaces/records.json: the twoupdatedismissals (vsplaybooks:update,firecrawl:update) are orphaned by the rename, so they are removed and re-recorded againstrestart. Twelve more dismissals:claude-testvsprototype:pressure-testandtesting:test-value(native description now "Run tests on this machine (local)"), andcode-reviewvsreview:security-review(native gained--max-findings).claude-testvsmutation-testing:audit,renamevsdocs-hygiene:rename-references,agentsvsmulti-agent:assessandmulti-agent:route,workflow-subagentvsmulti-agent:assess, andworkervsdiscovery:sweep-worker.Agentvsmulti-agent:assessandmulti-agent:route: ours decide how the Agent tool is used and launch nothing.cc-plugin-claude-testvstesting:run-e2e: covered by the existingclaude-test->testing:run-e2edefer row.docs/native-surfaces.mdregenerated withoverlap.py generate;node scripts/generate-catalog.mjsreported the catalog already in sync.plugins/discovery/skills/explore/reference/native-explore.mdandplugins/planning/skills/plan/reference/native-plan-agent.md: on 2.1.288 both agents disallow Agent, Artifact, ArtifactComments, ArtifactData, ArtifactCheck, ExitPlanMode, Edit, Write and NotebookEdit (sourceliteral). The records named four (Explore) and five (Plan) tools from the 2.1.285 extraction; they now state what that means for each skill and point atbuiltin_agents.<agent>.disallowed_toolsinstead of copying the list (Codex review).Verification
name:"update"occurs once in 2.1.287 and 0 times in 2.1.288;name:"restart"0 then 1. The 2.1.288 extraction readsrestartwithaliases: ["update"],hidden: true,gated: true.inventory.py --self-check:DEGRADED: cli 2.1.288, validated against 2.1.287(exit 3, all seven lanes ok) before the bump;OK: cli 2.1.288, validated against 2.1.288after.inventory.py --reader compare --self-check:reader compare: ok, 2167 of 2167 modules parse: no value->value divergence. A plain--reader comparerun shows 4 wrong->unresolved entries: the parser reads the Explore and Plandisallowed_toolsas partial, missing the four Artifact tools, as on 2.1.284-2.1.287 (feat(harness-ops): read Explore/Plan disallowed_tools literal under the parser reader, then flip the inventory default #5901). This PR records the literal regex values, which the live Plan agent roster confirms.native_drift.py diffagainst the 2.1.287 summary: surfaces added/removed/reclassified none, renamedupdate -> restart; invocability and markers none; fired triggers none; unresolved descriptions unchanged (design). The 15 items: the 14native-drift:candidate:*keys above plusnative-drift:inventory-degraded:2.1.288:inventory.overlap.py detecton the new store: discovered 0, resurfaced 0, orphaned dismissals 0, 120 suppressed.overlap.py self-check: exit 3, only the 2 standing advisories (older recorded extraction versions on rows; upstream SHA not locally decidable). 71 rows checked.TestInstalledBuildsnow covers 2.1.288 (passed locally with acorn, 88.8 s);extraction.mdnames 2.1.288 in its record.inventory.py: all checks passed.Related
multi-agentanddiscovery:sweep-worker, the source of several new candidates.🤖 Generated with Claude Code