Problem
Inside a claude plugin eval run, a skill under test can read only its hub SKILL.md. Every Read or Grep of the skill's reference/ or scripts/ files is refused with "File is in a directory that is denied by your permission settings", and the refusal lands in the trace's permission_denials. A path-scoped --allow-tools "Read(//<plugin>/skills/**)" grant is accepted by the CLI but does not lift the denial, and claude plugin eval --help lists no flag that makes a directory readable.
So any fact a case depends on must be in the hub. A skill that tells the model to "load the reference file" scores worse inside an eval than it does in real use, and every such read makes the run fail this repo's validity gate (plugins/evals/skills/plugin-eval/scripts/run-validity.py).
This is a limitation of the eval sandbox, tracked here. It can change in any Claude Code release, and it may depend on the account or the context.
Evidence
- Claude Code 2.1.287, 2026-10-02: the evals plugin's 31-case suite, job 5 baseline. 9 with-arm runs were denied reads of
skills/*/reference/*.md, and 5 were denied Greps of skills/*/scripts/.
- The V1 probe: the same denial occurred with the path-scoped grant on, and the run's settings held no deny rule.
- The record, with a recheck trigger, is in
plugins/evals/skills/plugin-eval/SKILL.md (the reference-read row in its facts table).
What this repo does about it
- Hubs state what a case needs. Reference files are labelled as background for a human reader.
run-validity.py fails a run with any with-arm denial at, under or above the plugin's directory.
Recheck
Recheck when a Claude Code release note touches plugin eval or sandbox permissions, when the plugin-evals docs gain a way to make a directory readable, or when a kept trace shows a reference Read succeeding. Then run one case with --keep-temp, read the with-arm trace, and update the facts-table row in plugins/evals/skills/plugin-eval/SKILL.md.
Problem
Inside a
claude plugin evalrun, a skill under test can read only its hubSKILL.md. EveryReadorGrepof the skill'sreference/orscripts/files is refused with "File is in a directory that is denied by your permission settings", and the refusal lands in the trace'spermission_denials. A path-scoped--allow-tools "Read(//<plugin>/skills/**)"grant is accepted by the CLI but does not lift the denial, andclaude plugin eval --helplists no flag that makes a directory readable.So any fact a case depends on must be in the hub. A skill that tells the model to "load the reference file" scores worse inside an eval than it does in real use, and every such read makes the run fail this repo's validity gate (
plugins/evals/skills/plugin-eval/scripts/run-validity.py).This is a limitation of the eval sandbox, tracked here. It can change in any Claude Code release, and it may depend on the account or the context.
Evidence
skills/*/reference/*.md, and 5 were denied Greps ofskills/*/scripts/.plugins/evals/skills/plugin-eval/SKILL.md(the reference-read row in its facts table).What this repo does about it
run-validity.pyfails a run with any with-arm denial at, under or above the plugin's directory.Recheck
Recheck when a Claude Code release note touches
plugin evalor sandbox permissions, when the plugin-evals docs gain a way to make a directory readable, or when a kept trace shows a referenceReadsucceeding. Then run one case with--keep-temp, read the with-arm trace, and update the facts-table row inplugins/evals/skills/plugin-eval/SKILL.md.