Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

skill-quality

A Claude Code plugin for skill-authoring QA: it runs a static, deterministic contract gate over a skill directory, reports the shared listing-budget estimate across a set of skills, and validates a skill's evals.json against a bundled schema plus a deterministic eval-quality lint, and scores description-driven auto-invocation probes (measure-invocation). Default paths invoke no model. The same checks run identically in a session, a pre-commit hook, or CI.

The drift static analysis catches best is a rewrite silently dropping a description trigger phrase, which can degrade a skill's auto-invocation. Check 3 compares the trigger phrases against HEAD and warns on each dropped phrase. It is advisory and never fails the run: a drop is often a deliberate consolidation of near-synonym triggers into a named intent category, so the warning asks the reviewer to confirm the description still names that intent, or to restore the phrase.

Skill What it does
/skill-quality:check Runs the contract gate (check), reports the shared listing budget (listing-budget), schema-validates and quality-lints evals (validate-evals), or scores description auto-invocation probes (measure-invocation).
/skill-quality:setup Check-only: resolves and verifies the skills directory and prints the guidance for routing a non-default skills_root change through Claude Code.

Checks

check runs check-skill.sh. Twenty-six checks, reported as FAIL: (blocking) or WARN: (advisory):

  • Frontmatter parses; description present; a declared name is kebab-case and matches the skill directory (in a plugin skill it also WARNs as redundant, because the field defaults to the directory). A plain (unquoted, non-block) description that contains ": " or a line ending in : on any line FAILs: that is a YAML mapping indicator, and the skills reference says unparsed frontmatter loads the skill with no fields set (https://code.claude.com/docs/en/skills#frontmatter-reference). A quoted or block scalar may contain the indicator. compatibility, when present, is 1-500 characters and FAILs outside that (Agent Skills spec; Claude Code accepts the field and does not act on it). Absence is success: the spec says most skills do not need the field. The effective name (the declared field, else the directory leaf) is at most 64 codepoints (FAIL; the Agent Skills spec's name cap, https://agentskills.io/specification, enforced by its skills-ref validator) and carries neither anthropic nor claude (WARN; a Skills API upload requirement, https://platform.claude.com/docs/en/build-with-claude/skills-guide#creating-a-skill, not a spec rule). Claude Code enforces neither and ships bundled skills named claude-api and claude-in-chrome, so both are portability findings (verified 2026-09-10; recheck when the spec's validator, the upload requirements, or a Claude Code release changes either rule).
  • description + when_to_use within the 1536-char per-skill listing-entry cap (overflow truncates that entry). A different, narrower limit from the shared budget below. The cap and the 1% budget default are upstream's (https://code.claude.com/docs/en/skills#frontmatter-reference, https://code.claude.com/docs/en/settings; verified 2026-08-31; recheck trigger: either default moving re-derives this line and the scripts' encoded constants).
  • description alone within the Agent Skills spec's 1024-codepoint field maximum (https://agentskills.io, "Maximum 1024 characters"; verified 2026-08-31). A separate limit at a separate layer from the listing-entry cap above: a description can sit under 1536 combined and still breach 1024 on its own, which the Skills API rejects at upload. Over the maximum FAILs; within 32 codepoints of it WARNs, naming the skill and how much room is left, so an author sees the ceiling before the next added trigger phrase hits it. Counted in codepoints, not bytes, so a non-ASCII description is measured the way the spec states the limit. CHECK_SKILL_DESC_FIELD_BASELINE=<file> records skills whose breach predates the FAIL, one repo-relative skill path per line; those downgrade to a WARN, and a row whose skill no longer breaches FAILs as stale, so the list can only shrink.
  • Trigger-keyword preservation vs HEAD (advisory: a dropped phrase warns naming it and never fails the run; a phrase moved to a sibling skill warns naming the host; skipped for a new, uncommitted skill).
  • SKILL.md under 500 lines, the cap the official skill-authoring guidance sets.
  • Backtick- and link-cited skill-internal supporting files resolve. When a path that misses instead resolves under a sibling skill, the finding names that sibling and the ${CLAUDE_PLUGIN_ROOT}/skills/<sibling>/... cross-skill form, while keeping the hand-verify caveat (the sibling hit is evidence, not proof: paths can collide). A cited path with a backslash separator (scripts\helper.py), in SKILL.md or in any markdown spoke under reference/, references/, or context/, FAILs outright, naming the citing file and the forward-slash form: Claude Code rejects such a component path at plugin load on macOS and Linux.
  • markdownlint-cli2 clean (advisory-skips when npx is absent).
  • scripts/*.test.sh pass where present.
  • Vendored vendor/ byte-identical vs HEAD; stale-tracking metadata keys preserved; sync age.
  • Gotchas surface present; description carries Use when phrasing; no committed cache artifacts; action-router skills without evals WARN (FAIL with --require-evals for any shape unless a recorded skip exists in scripts/evals-warrant-exemptions.txt); companion spoke dirs are referenced.
  • Precompute opportunity (advisory). A fenced shell block gathers read-only context the skill could inline at load time via ! injection instead of a per-invocation tool call.
  • !-injection portability. Bash-only syntax without a shell: declaration fails; portable-looking but undeclared warns; an injected command with no || <fallback> continuation warns.
  • Fresh-eyes declaration conformance. Same-context judgment language (a curated, advisory heuristic) expects fresh-context delegation wording or a fresh-eyes-exempt directive nearby; malformed or reason-less directives fail. Contract: skills/check/reference/fresh-eyes-declarations.md.
  • metadata.summary within 100 Unicode codepoints. The key is the generated skill cheat sheet's row source; the cap keeps rows scannable. An absent key is no finding.
  • Completion-criteria signal (advisory). A numbered procedure of three or more steps with no observable done-condition token.
  • Explicit invocation mode. Marketplace plugin skills must state disable-model-invocation; elsewhere a missing key warns.
  • Description/verb-contract polarity (FAIL:, blocking). The description lead contradicts the Naming verb contract or the body (read-only vs mutate). --fix in the listing is the compliant override shape.
  • Long spoke files carry a table of contents (advisory). A markdown file under reference/, references/, or context/, at any depth, over 300 lines whose first 40 lines hold fewer than three ](# in-page anchor links warns, naming the file. The threshold is the bundled skill-creator's; the 100-to-300 band stays with docs-hygiene:audit-progressive-disclosure, whose TOC heuristic this check mirrors.
  • ## Next successor section (advisory). A section placed after ## Gotchas, last in the file, in neither the one-invocation nor the two-to-four-outcome-bullet shape, or carrying operative-chain phrasing warns. Absence is an INFO note, because most skills are terminal, except on a stage-bearing skill: a metadata.workflow-stage of explore, research, plan, implement, test, review, verify, pr, or retro with no ## Next and no Handoff, Routing, Integration, or Skill chaining heading warns. contract is not in that list, because those skills route through the slice they write.

listing-budget runs check-listing-budget.sh. An always-advisory report on the shared budget every loaded skill draws from together (skillListingBudgetFraction, default 1% of the model's context window). This is the aggregate limit check's per-skill cap above does not cover: nothing else in the gate checks it, so a marketplace's skill count can silently overflow the live listing with no local signal. It never asserts a live value it cannot observe (the model's context window and a consumer's settings are both unknowable statically). It reports against a documented, overridable default and always exits 0.

Only listing-eligible skills are counted: disable-model-invocation: true keeps a skill's description out of the model-visible listing entirely, so it spends none of the shared budget. A consumer's skillOverrides can free further descriptions with "name-only", which repository content cannot reveal, so the reported figure is an upper bound for anyone who sets it.

check also takes one or more skills roots as positionals, so several trees are gated in one run. Each root is walked on its own under its own header and the run ends with a single N passed, M failed rollup. Nothing is pooled across roots: the cross-skill scans (trigger-move, sibling-ref, plugin-root detection) stay per root, and pooling is listing-budget's job precisely because the budget it reports is the shared one.

/skill-quality:check my-skill                  # gate one skill
/skill-quality:check                           # gate every skill under the resolved root
/skill-quality:check plugins/*/skills          # gate every skill under each plugin's root
/skill-quality:check validate-evals my-skill   # schema-check + quality-lint evals.json
/skill-quality:check listing-budget            # report the shared budget over the resolved root
/skill-quality:check listing-budget plugins/*/skills  # pool every plugin's root into one aggregate
/skill-quality:check measure-invocation        # score description auto-invocation probes

Skills directory is never baked in

The checker resolves the skills root through the convention-resolution ladder, first hit wins:

  1. ${user_config.skills_root}. Set only when your skills live outside .claude/skills.
  2. ${CLAUDE_PROJECT_DIR}/.claude/skills. The conventional default.

CHECK_SKILL_SKILLS_ROOT is honored as a one-run environment override the checker reads directly; the setup skill neither writes nor persists it. When your skills live at the default location, no configuration is needed:

/skill-quality:setup         # check-only: resolve + verify the skills directory (re-runnable);
                             # prints how to route a skills_root change through Claude Code

Evals schema + quality lint

validate-evals checks a skill's evals/evals.json against the bundled reference/evals.schema.json. Every case requires id, prompt, and at least one non-empty grading criterion: expected_output, expectations, or assertions (a case that cannot be graded is not an eval); the rich form adds name (kebab-case) and files. Evals are warranted, not mandatory. A skill shipping none is not a failure.

After the schema, check-evals-quality.sh (bash + jq) lints eval CONTENT deterministically. FAIL: tier: duplicate case ids/names, empty criterion items, files fixture entries that resolve to no path under the skill or evals directory. WARN: tier (advisory, exit 0): a case carrying both expectations and assertions, identical prompt+files pairs, vague whole-item phrasing ("the output is good"), a thin sole-criterion expected_output, a set with no refusal/guardrail or anti-pattern case, and (Q4 prose) an empty files list with path-shaped tokens in prompt/expected_output that resolve nowhere (silence with narration: true or declare fixtures). It deliberately does not flag low case count. Run --help on the script for the full Q1-Q9 list; without jq it exits 2 and the schema verdict stands alone.

Requirements

  • A skills root via CHECK_SKILL_SKILLS_ROOT, CLAUDE_PROJECT_DIR, or a git repository (last-resort default .claude/skills). Git-backed checks (trigger preservation, vendor identity, stale metadata, committed artifacts) skip with a note outside a repo so marketplace plugin-cache installs (plain trees) still run the rest of the gate.
  • npx (Node) is optional; without it the markdownlint check downgrades to a warning and the other twenty-five still gate.

Description-invocation probes

measure-invocation scores whether a skill's listing text would win the requests it should (and stay quiet on the ones it should not). Default method is a deterministic lexical listing-overlap floor; emit-plugin-eval writes claude plugin eval cases for a live run. Contract: reference/invocation-probes.md.

bash plugins/skill-quality/scripts/measure-invocation.sh validate plugins/skill-quality/probes
bash plugins/skill-quality/scripts/measure-invocation.sh score plugins/skill-quality/probes
bash plugins/skill-quality/scripts/measure-invocation.sh compare \
  plugins/skill-quality/probes/baselines/listing-overlap.json /tmp/score.json

Configuration

Options reference

Generated from this plugin's .claude-plugin/plugin.json. Every option Claude Code will prompt for when the plugin is enabled, with the environment variable each hook reads it from.

Option Type Default Environment variable Description
skills_root directory (none) CLAUDE_PLUGIN_OPTION_SKILLS_ROOT Directory holding your skills (each a subdirectory with a SKILL.md). When unset, resolves to .claude/skills under the project root. Set this only when your skills live elsewhere.

How to set these

Three supported routes, in the order most people want them:

  1. Interactively. Claude Code prompts for declared options when you enable the plugin. To change them later: /plugin configure skill-quality@<marketplace>.

  2. Headless. Repeat --config for each option. Replace <marketplace> with the marketplace you installed this plugin from:

    claude plugin install skill-quality@<marketplace> -s <scope> --config skills_root=<value>

    The same command reconfigures a plugin that is already installed: it prints already installed and still writes the value. The short-circuit message is about the install, not the config write. Do not claude plugin uninstall to reconfigure: uninstalling drops this plugin's whole stored pluginConfigs entry, resetting every option in the table above to its default. -s defaults to user, so pass the scope claude plugin list reports for this plugin. The verified-version record lives in the plugin-reconfiguration convention.

    The value is stored immediately; the session you are in does not change. Hooks are handed their CLAUDE_PLUGIN_OPTION_* when the session starts, so start a fresh Claude Code session before expecting new behavior. A check run in the old session still reports the old value, and that is not a failed write.

  3. By hand, in settings. Add the value under pluginConfigs in your user settings (~/.claude/settings.json):

    {
      "pluginConfigs": {
        "skill-quality@<marketplace>": {
          "options": {
            "skills_root": <value>
          }
        }
      }
    }

    Plugin option values are read from user, --settings, and managed settings only, not from a project's .claude/settings.json. To vary behavior per repository, enable or disable the plugin in that project's enabledPlugins instead of setting an option there.

Do not set the CLAUDE_PLUGIN_OPTION_* variables yourself. They are how Claude Code hands a configured value to a hook process; the value comes from the routes above.

Upstream documentation