A Claude Code plugin for the test stage of a disciplined dev workflow. Plan what needs testing, author tests at the right level, verify the running app end-to-end, diagnose failures to root cause, and catch tests that cannot fail. Seven skills, one concern: proving behavior with tests.
| Skill | What it does |
|---|---|
/testing:plan |
Coverage-gap analysis. Classify changed files by required test type, identify gaps, prioritize by regression risk. |
/testing:write |
Test authoring discipline. Vertical-slice TDD, test-type selection, naming, placement, fixture patterns, four-pillars assessment. |
/testing:run-e2e |
Live app verification. Start the app via the project's orchestrator, drive UI/API flows with token-efficient browser automation, capture evidence; includes a non-UI smoke-test playbook (MCP stdio handshake, shell/PowerShell surfaces). |
/testing:diagnose |
Failing-test diagnosis. Failure classification, root-cause analysis (never retry blindly), then the reproduce → isolate → fix → retest → regression loop. |
/testing:audit |
Can't-fail test detection: a deterministic script runs twelve rules across JS/TS, Python, C#, Bash, PowerShell and Go, from assertion-free bodies and self-identical (recomputed-expectation) assertions to unawaited assertions, conditional assertions and Playwright retry or test.only configs. --check fails on the first two (Bash-harness findings only with --strict); --strict adds mock-only oracles and the two Playwright config rules; the other seven only report. It reports with a coverage denominator and opt-in persists findings for a review fix pass. |
/testing:setup |
Configure the can't-fail checks: check prints the resolved .claude/testing.yaml, the test-lint rules missing per language, an optional instruction line to paste, and a settings hook entry for test globs the shipped hook skips; apply writes .claude/testing.yaml. |
testing:test-value |
Model-invoked guidance, loaded by the review and implementation agents and the test-scan hook: where each expected value must come from, when call-count and database checks are legitimate, and the can't-fail taxonomy keyed to /testing:audit rule ids. |
- Reads your conventions, assumes none. Test frameworks, project locations,
naming, fixture patterns, and the e2e-orchestrator configuration come from your own
project's
CLAUDE.md/ rules and existing test projects; the skills infer from what exists when nothing is documented. - Cross-plugin refs degrade gracefully. Test invocation defers to the
toolchainplugin's/toolchain:checkwhen installed and to the project's own test command otherwise; TDD design questions route to/tdd:principles, browser mechanics to/playwright:playwright, outcome sign-off to/verification:confirm, and the implement loop to/implementation:implement. Each is used when installed and substituted with inline guidance or a manual handoff when absent. No step blocks on a missing plugin. - Self-contained. Test-type tables, the E2E evidence contract, the non-UI
smoke-test playbook, and diagnosis loops ship inside the plugin and are referenced
via
${CLAUDE_PLUGIN_ROOT}.
- Node.js on
PATH. Every hook row launches throughnode hooks/exec-bash.mjs, which finds Bash. A missingnodeis a hook launch error, not a skip notice.
/plugin marketplace add melodic-software/claude-code-plugins
/plugin install testing@melodic-softwareTest structure and conventions come from your own project's CLAUDE.md and rules.
Two userConfig options, prompted by Claude Code at enable time:
test_guards_enabled(defaultfalse) turns on two hooks.test-scan(PostToolUse) runs the can't-fail scanner on each test file Claude writes or edits and returns the findings as context.test-weaken(PreToolUse) asks Claude for a reason when an edit removes or skips tests or assertions.stdin_read_timeout(default2seconds) bounds how long a hook waits on its input before it fails open.
/testing:run-e2e reads one optional consumer-project config surface,
.claude/testing/e2e.md: recording (video | gif | off, default off) and
browser_mode (headed | headless, default headless). Both defaults preserve current
behavior, so the file is optional. Its keys, defaults, and precedence are documented in
the skill's bundled run-e2e/context/e2e-config.md; it layers per the marketplace
config-cascade convention.
/testing:audit and the test-scan hook read .claude/testing.yaml through the same cascade
(~/.claude/testing.yaml, the team file, .claude/testing.local.yaml): adapters to turn off or
allow, path globs to exclude or include, extra adapter globs, consumer adapters, and a level per
rule (off, warn, error). /testing:setup documents the keys and writes the file.
test-scan also scans test files a Bash call changed (cat > foo.test.ts, sed -i, a generator
script), when test_guards_enabled is on and Claude Code records the call's changed files. It reads
tool_response.bashEditDiff, a best-effort beta field that Claude Code adds to the PostToolUse
payload of a Bash call:
-
Precondition. Set
bashEditDiffEnabled: truein your user settings, in--settings, or in managed settings (a project.claude/settings.jsonvalue is ignored), or set the environment variableCLAUDE_CODE_BASH_EDIT_DIFF=1. Without one of these the field was absent in every mode probed (default,acceptEdits,autoandbypassPermissions), and the hook finds nothing to scan. The probe rows are in probes.md. -
Scope. A created test file reports every test block. A modified file reports only the blocks its hunks touch, the same as an Edit. A file the repository ignores is skipped.
-
Limits. The payload carries hunks for the first five changed files only, so a modified test file past the fifth is not scanned, and one call scans at most four test files. The shipped adapters' globs decide what counts as a test file; a glob added through
.claude/testing.yamlis not covered on this path.test-weaken(PreToolUse) sees Write and Edit only: a Bash call that removes assertions is not flagged. Windows Git Bash is not probed. -
Cost. A Bash hook row cannot filter on the changed files, so the row has no
ifand Claude Code starts its node launcher for every Bash call, whatevertest_guards_enabledsays. The hook budget is k × S, where S is onebash -c :spawn and k the processes one fire starts. On WSL2, S was 1.0 ms (p50 of 50 samples interleaved with the hook arms, at a load of about 5; Windows is not measured):- Option off, the default: k = 1 (node; the option gate closes before bash starts), 22 ms, about 22 S.
- Option on, no recorded change, which is every Bash call for an opted-in user: k = 2 (node, then bash), 30 ms, about 30 S.
- Option on, a call that changed one test file: the scan itself, 96 process creations and execs, 116 ms, about 115 S. It runs only for such a call, so it is not always-on.
The multiples read high because S is small on Linux: starting node costs about 22 ms against 1 ms for bash.
.performance/ratchets.jsonholds the first two paths as spawn-count ceilings (testing-posttooluse-bash-test-scan-option-off-spawns, 1, andtesting-posttooluse-bash-test-scan-no-diff-spawns, 3). A ceiling counts process creations plus execs, so the two-process path counts 3.
Generated from this plugin's .claude-plugin/plugin.json. Every option Claude Code
will prompt for when the plugin is enabled, with the environment variable each hook
reads it from.
| Option | Type | Default | Environment variable | Description |
|---|---|---|---|---|
test_guards_enabled |
boolean | false |
CLAUDE_PLUGIN_OPTION_TEST_GUARDS_ENABLED |
Scan each test file Claude writes or edits for tests that cannot fail, and ask Claude for a reason when an edit removes or skips tests or assertions. Off by default. |
stdin_read_timeout |
number min 1 |
2 |
CLAUDE_PLUGIN_OPTION_STDIN_READ_TIMEOUT |
Idle bound on reading the hook payload from stdin: how long the pipe may go silent before the hook gives up and fails open |
Three supported routes, in the order most people want them:
-
Interactively. Claude Code prompts for declared options when you enable the plugin. To change them later:
/plugin configure testing@<marketplace>. -
Headless. Repeat
--configfor each option. Replace<marketplace>with the marketplace you installed this plugin from:claude plugin install testing@<marketplace> -s <scope> --config test_guards_enabled=<value>
The same command reconfigures a plugin that is already installed: it prints
already installedand still writes the value. The short-circuit message is about the install, not the config write. Do notclaude plugin uninstallto reconfigure: uninstalling drops this plugin's whole storedpluginConfigsentry, resetting every option in the table above to its default.-sdefaults touser, so pass the scopeclaude plugin listreports for this plugin. The verified-version record lives in the plugin-reconfiguration convention.The value is stored immediately; the session you are in does not change. Hooks are handed their
CLAUDE_PLUGIN_OPTION_*when the session starts, so start a fresh Claude Code session before expecting new behavior. A check run in the old session still reports the old value, and that is not a failed write. -
By hand, in settings. Add the value under
pluginConfigsin your user settings (~/.claude/settings.json):{ "pluginConfigs": { "testing@<marketplace>": { "options": { "test_guards_enabled": <value> } } } }Plugin option values are read from user,
--settings, and managed settings only, not from a project's.claude/settings.json. To vary behavior per repository, enable or disable the plugin in that project'senabledPluginsinstead of setting an option there.
Do not set the CLAUDE_PLUGIN_OPTION_* variables yourself. They are how Claude Code
hands a configured value to a hook process; the value comes from the routes above.
- User configuration: the
userConfigschema and theCLAUDE_PLUGIN_OPTION_<KEY>export - Plugin install options: the
--configflag's reference entry - Plugins and skills settings:
enabledPlugins,extraKnownMarketplaces,pluginConfigs - Settings files and who they affect: user vs project vs local precedence
- Manage installed plugins: enabling, disabling,
/plugin list
MIT (SPDX-License-Identifier: MIT).