feat(disk-hygiene): attended deep inventory with justified KEEP reasons - #5585
Conversation
…idator Read-only stdlib module with the shared row schema (name, ext, size, mtime, owner, producer, category, disposition, reason, evidence) and named categories: superseded versioned dirs, unreferenced plugin cache versions, /tmp by producer prefix, transcript dirs whose source path is gone, and dangling symlinks. validate_report fails a KEEP whose reason is empty or only a category phrase unless its evidence shows the named tool still references the entry. Refs #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Declare an `inventory` subcommand in engine_grammar (--target, --data-root, optional --deep), so the engine parses it and the guard admits it from one declaration; apply and preview grammar are unchanged. The guard allows it alongside scan, preview and handoff-verify through one named read-only set. Inventory streams one row per entry to a JSONL report under the data root, with no entry cap. Deep mode lists every level with bottom-up directory sizes and is the default when the target is the home directory; otherwise only immediate children are listed. Category rows (plugin cache versions, transcript dirs, /tmp producers, dotted superseded versions, dangling links) replace the unclassified row at their path; an unreadable or mounted subtree is one UNKNOWN row. Each row runs through validate_report; a failure marks the summary inventory-failed and exits 5. The summary is neither a snapshot nor a plan, and preview refuses it. Refs #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…wner lookups Patch Path.home and the OS-managed check in the inventory tests so they hold on macOS temp paths and Windows homes, and cache uid-to-name lookups so a home-directory walk does one password-database lookup per owner. Refs #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iene pointer Document `--deep` as an attended, report-only mode: default for a whole-home target, forced elsewhere. List the shared row columns and named categories, state that every KEEP needs a specific reason checked by the validator, and that --execute, the low-signal rule and the confirmation gates are unchanged. Add an eval case, and a one-line pointer in repo-hygiene:clean to machine-level listing. Refs #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hangelog entries Refs #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
running_paths returned an empty set when /proc was missing, so on macOS and Windows every non-newest version and old /tmp entry became a CANDIDATE with the reason "no running process uses it", which nothing had checked. It now returns None for an unreadable process table, and the rows that rest on it are UNKNOWN with a reason that says the table was not read. Also drops dangling_symlinks, which only tests called (the walk uses dangling_row), moving its coverage to the walk, and trims SKILL.md back under the 500-line cap. Refs #5221 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Resolve the disk-hygiene and repo-hygiene conflicts against main: renumber the bumps above main (disk-hygiene 0.35.0, repo-hygiene 0.18.1), move the new eval to id 16 and add id 17 for a bare home target, and carry both `catalog` and `inventory` in the guard's read-only tuple. Fold the deep-inventory section of SKILL.md back under the 500-line cap by leaving the detail in scan-flags.md and safety-model.md, state that a home-directory target runs the deep inventory before any scan, and tighten the categories: a release outranks its prerelease, a symlink to a version keeps it, plugin cache candidates carry their .orphaned_at marker age, and the /tmp category reads /tmp instead of $TMPDIR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep the disk-hygiene bump at 0.35.0, above main's 0.34.2, and carry main's 0.34.2 changelog entry beneath it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…category The README lists the scan flags but not --deep. scan-flags.md now says the tmp-producer category yields rows only when the target is /tmp or contains it, so a home inventory has none and /tmp is inventoried as its own target. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep the 0.35.0 entry above main's 0.34.4. Carry main's gated --sizes-only statement (#5587) into the moved SKILL.md and scan-flags.md text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ence, and pass lint
The deep-inventory validator let a KEEP row with an empty reason pass whenever
its evidence named any tool and a reference, and let "tool-managed" pass with
evidence for any tool. An empty reason now always fails, and a category-only
reason passes only when it names the tool ("managed by <tool>") that the
evidence shows still references the entry.
Also: mark deep_inventory.py and test_deep_inventory.py executable (they carry
shebangs), move the test fixture paths outside /home, drop the duplicate
sys.path insert in the engine, and give ORPHAN_SWEEP_DAYS and PROJECT_NAME_CAP a
verification stamp and recheck trigger.
Refs #5221
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Main took disk-hygiene 0.35.0 (handoff-apply, #5541), so the inventory release moves to 0.36.0 above it. The destructive guard's read-only set now lives in main's frozenset beside the mutating set, and inventory joins the read-only side; the tests that named the old tuple or pinned apply as the last subcommand follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Main took disk-hygiene 0.35.1 (#5613); the 0.36.0 entry stays above it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Main took disk-hygiene 0.35.2 (#5526); the 0.36.0 entry stays above it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Renumber the inventory release to disk-hygiene 0.37.0 above main's 0.36.0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Claude finished @kyle-sexton's task in 7m 8s —— View job CI Code Review
This repo runs a separate security lane ( Scope of this review: FindingsOne inline finding posted on 🔴 Important — the two related report files ( What I verified statically (not executed)I traced the Note: I was not able to execute the Python test suites or |
|
Claude finished @kyle-sexton's task in 1m 32s —— View job Security review complete
No security issues found. Scope: the new What I checked:
Out of scope per this lane's charter (and not re-reported here): GitHub Actions/workflow hardening (zizmor's lane — no workflow files changed in this PR), and general code quality/style ( |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3b8e1ddbb8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…s, write the report atomically Merges origin/main and renumbers the release to 0.38.0 above main's 0.37.0. - The deep walk and the size sum skip any mount point read from /proc/self/mountinfo, not only a device change, so a same-device bind mount is one not-walked row. - running_paths also reads /proc/<pid>/fd, so a /tmp entry a process holds open is kept, not offered as a candidate. - The inventory rows stream to a temporary file renamed into place when the walk finishes, and the summary is written atomically, so an interrupted walk leaves no partial report. Refs: #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 3m 14s —— View job Code Review (
|
|
Claude finished @kyle-sexton's task in 2m 49s —— View job Security review complete
Scope: commit range FindingsOne inline finding posted on 🟠 IMPORTANT — the new bind-mount stop silently fails open when What I verified
Everything else in the diff (docs, the merge-brought-in files from |
Renumber the inventory release to disk-hygiene 0.39.0 above main's 0.38.0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 1m 11s —— View job Code Review (
|
|
Claude finished @kyle-sexton's task in 2m 4s —— View job Security review complete
Scope: FindingsOne finding, still open from the prior review round — the diff at this location is byte-for-byte unchanged since it was first flagged on commit 🟠 IMPORTANT — mounts = frozenset(str(p) for p in linux_mount_points()[0])This discards the This is the same finding reported on the prior round for commit What else I checkedSince the diff is unchanged from the last full pass, I re-confirmed rather than re-deriving: the report-only boundary still holds (no delete/move/write outside the guarded data root), the atomic-write fix for Out of scope per this lane's charter: GitHub Actions/workflow hardening (zizmor's lane — no workflow files changed) and general code quality/style ( |
…mmary The inventory summary carries mount_state_error when /proc/self/mountinfo cannot be read, so a report that walked without bind-mount detection does not read as a complete one. Refs: #5221 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude finished @kyle-sexton's task in 52s —— View job Security review complete
No security issues found. Scope:
This resolves the silent-failure concern: a reader of the report can no longer mistake a walk that ran without bind-mount detection for one that ran with it. The walk still proceeds and still reports Everything else in the diff is unchanged since the last full pass at Out of scope per this lane's charter: GitHub Actions/workflow hardening (zizmor's lane — no workflow files changed) and general code quality/style ( |
|
Claude finished @kyle-sexton's task in 1m 39s —— View job Code Review (
|
…ogic with claude-ops (#5632) Closes #5221 ## Summary Items 1, 2, 4, 5 and 7 of #5221 landed in #5585. This PR completes item 3 (reuse the unreferenced-cache-version logic from the claude-ops install-state audit) and item 6 (coordination with the related deep-inventory work). `disk-hygiene`'s deep inventory carried its own plugin-cache-version logic, a divergent copy of the claude-ops `audit-install-state` rule: it read `installed_plugins.json` without the guarded read the audit uses. ## Fix - One module, `lib/plugin_cache_versions.py`, is carried byte-identical in `plugins/claude-ops/lib/` and `plugins/disk-hygiene/lib/`. `install_state.py` and `deep_inventory.py` both call it, so both apply one rule for which cache versions are unreferenced. - The module is registered in `scripts/cross-plugin-source-registry.txt`, so `scripts/check-cross-plugin-source-drift.sh` fails if the copies diverge. - `claude-ops` 0.77.3 and `disk-hygiene` 0.40.1, each with a CHANGELOG entry. - Item 6: no second listing to align. The successor of closed #5214, `/disk-hygiene:audit` (#5590), reads the scan `children_rollup`. The shared listing schema for the managed-state lane (#4006) lives in `plugins/disk-hygiene/skills/clean/scripts/deep_inventory.py` (`ROW_COLUMNS`). #4006 is not folded in and its scope is unchanged. ## Verification - `bash scripts/check-changelog-parity.sh --check`: pass - `bash scripts/check-changelog-parity.sh --check-order`: pass (102 changelogs) - `bash scripts/check-changelog-parity.sh --check-bump origin/main`: pass - `bash scripts/validate-plugins.sh`: all manifests and the catalog validated - `bash scripts/check-changed-skills.sh origin/main`: 2 skills checked, 0 failed - `bash scripts/check-cross-plugin-source-drift.sh`: exit 0; `lib/plugin_cache_versions.py` IDENTICAL [registered] - `python3 -m unittest` on `test_deep_inventory.py` (52 tests) and `test_install_state.py` (91 tests): OK ## Related - #5585: landed items 1, 2, 4, 5, 7 of #5221 - #5420 - #5590: successor of closed #5214; it reads scan `children_rollup`, so there is no second listing to align - #4006: open; the shared listing schema is `ROW_COLUMNS` in `plugins/disk-hygiene/skills/clean/scripts/deep_inventory.py`, for the managed-state lane to emit. Not folded in here. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Refs: #5221
Summary
/disk-hygiene:cleangains an attended, report-only deep inventory. A home-directory target (so a bare/disk-hygiene:clean ~), or any target with--deep, runs the read-onlyinventorysubcommand before anyscan. It lists every entry with name, extension, size, mtime, owner, producer, category, disposition and reason. Every KEEP needs a specific reason: an empty reason always fails the report, and a category-only reason ("tool-managed", "OS-owned", "managed by ") fails unless it names a tool ("managed by ") that the row's evidence shows still references the entry.Fix
inventorysubcommand, row schema, named categories (superseded versions, unreferenced plugin cache versions, /tmp by producer, orphaned transcript dirs, dangling symlinks) and the KEEP-reason validator.toolis that named tool and itsreferencesis set. Tests cover the empty reason with evidence, a reason that names no tool with evidence, and a different tool's evidence.inventoryjoins_READONLY_ENGINE_SUBCOMMANDS(scan, inventory, preview, handoff-verify, catalog) beside main's_MUTATING_ENGINE_SUBCOMMANDS(apply, handoff-apply), and the kill-switch denial text names it. A test pins the denial text.SKILL.mdnames--deepinargument-hintand Arguments, and its Deep inventory section says a bare home target runsinventoryfirst. The large-root paragraph links to it, so the bounded--max-depth 1scan is the step that follows for removal candidates, not the first step.reference/scan-flags.mdholds the columns, categories and the KEEP rule;reference/safety-model.mdholds the report-only boundary. The--sizes-onlytext in both follows main (fix(disk-hygiene): gate --sizes-only behind the large-scan confirmation and keep no per-path entries #5587): it goes through the large-scan question.SKILL.mdis 498 lines (cap 500). The pluginREADME.mddocuments--deep.~/.local/binand~/bin) is KEEP; a plugin cache candidate carries its.orphaned_atmarker age and whether it is past the 14-day sweep window, and says when the registry records no install so the sweep does not run;tmp-producercovers/tmpitself, not$TMPDIR.ORPHAN_SWEEP_DAYSandPROJECT_NAME_CAPcarry a verification stamp and a recheck trigger./procis unreadable (macOS, Windows),running_pathsreturnsNoneand rows that would be CANDIDATE on "no running process uses it" are UNKNOWN, with a reason that says the table was not read.--execute, the low-signal keep rule and the confirmation gates are unchanged; any removal still goes through scan, preview and the removal approval.repo-hygiene:cleangets a one-line pointer to deep mode and no scanner.--deep ~lists with justified KEEP, deletes nothing) and 17 (bare~runs the deep inventory first, no bounded scan first).handoff-apply, two fixes and/disk-hygiene:auditwhile this was open) and repo-hygiene 0.18.0 to 0.18.1, each with a CHANGELOG entry.Known limits, all report-only (deletion stays gated): a version selected by a file rather than a symlink (an nvm alias,
.tool-versions) is not seen, so its row can be a CANDIDATE for a version in use; thetmp-producercategory yields rows only when the target is/tmpor contains it (so a home inventory has none;--deep /tmpcovers it) and none on macOS, where/tmpis a link and/private/tmpan OS-managed root.Verification
Run on the head
b22f973b0, which merges current origin/main (disk-hygiene0.36.0); the only conflict was the disk-hygiene CHANGELOG, resolved by renumbering this release to 0.37.0:*.test.shwrappers: all pass (hygiene.test.sh638 tests, 1 skipped;deep_inventory.test.sh50 tests). All repo-hygiene*.test.shwrappers pass too.scripts/run-ruff.sh check plugins/disk-hygiene: all checks passed;format --checkondeep_inventory.py,test_deep_inventory.py,destructive_guard.pyandengine_grammar.py: already formatted.scripts/check-changed-skills.sh "$(git merge-base origin/main HEAD)": 3 skills checked, 0 failed.markdownlint-cli2onSKILL.md,scan-flags.md,safety-model.md,README.mdand the disk-hygiene CHANGELOG: 0 issues.scripts/check-changelog-parity.sh --check,--check-orderand--check-bump origin/main: pass.scripts/validate-plugins.sh: all manifests and the catalog validated.CI:
lintandci-statuswere red ond0f85a3fbfor the executable bit ondeep_inventory.pyandtest_deep_inventory.pyand for Linux user paths in the test fixtures. Both are fixed, and on6e852df64lint,lint-2,test-linux (0-3),hook-utils,managed-files-guard,ci-lanesandci-statuspass.Related
Closesis not used because two items of the owner's decision on #5221 are still the owner's call:plugin_cache_versionsindeep_inventory.pyapplies theaudit-install-stateregistry rule (_install_paths,_resolve, the registry-doubt rules) as a plugin-local copy, because a plugin cannot import a sibling. It is not inscripts/cross-plugin-source-registry.txtand nosync-*.shgate covers it, so nothing catches drift between the two copies; the tests here pin this copy only. The owner decides between vendoring it through async-*.shscript plus a registry entry and accepting the separate copy.deep_inventory.pyand documented inreference/scan-flags.md#--deep; their scope is not folded in. No comment was posted on either issue from this PR, so a note there is the owner's call.Flipping this PR to ready also stays with the owner.
🤖 Generated with Claude Code