Problem
The evals plugin's eval suite reached a VALID run with the plugin scoring 0.99 and 0.31 without (delta +0.68, 95% interval +0.52 to +0.83; 29 of 31 cases at 1.00). The work left gaps in the suite's judge calibration, and a wider conversion task, none of them blocking.
Follow-ups
- Two judges fail correct answers that their calibration samples don't cover.
noise-report-before-posting/gate-before-numbers failed correct answers in four runs. Every one used the shape "run run-validity.py first; on INVALID, post the INVALID line and its reasons". Candidate causes: an absolute script path, a second gate on the noise report, or a trailing partial check.
noise-before-gain/not-established failed an answer that opens with "the numbers point the right way", which its rubric names as allowed.
- Add must-pass samples of these shapes, then recalibrate.
- Three must-fail samples are never tested, because the reply-only agent will not reproduce them word for word:
measurable-criterion four-properties sample 3: the model drops its "launch a medical bot with 'answers are safe'" paragraph as unsafe advice.
worktree-guard-stop no-way-around sample 4, and zeros-after-usage-limit check-and-rerun sample 3: both get paraphrased.
- Rewrite each so it still fails for one reason, with wording the model will repeat, for example a retraction that isn't safety-sensitive.
measurable-criterion must-fail sample 2 now misses two properties. It was built to miss only the trial set. Since the Achievable rule changed, its invented "2% measured last month" baseline fails too, so it no longer tests one gap on its own.
- Convert the other plugins' evals (V7). Only
source-control has a claude plugin eval suite (17 cases). 82 plugins carry skill-creator evals.json files only (326 files).
Context
The loop that produced these, and the reasoning behind each, is in the pull request that adds the evals suite and validity gate (link to follow).
Problem
The evals plugin's eval suite reached a VALID run with the plugin scoring 0.99 and 0.31 without (delta +0.68, 95% interval +0.52 to +0.83; 29 of 31 cases at 1.00). The work left gaps in the suite's judge calibration, and a wider conversion task, none of them blocking.
Follow-ups
noise-report-before-posting/gate-before-numbersfailed correct answers in four runs. Every one used the shape "run run-validity.py first; on INVALID, post the INVALID line and its reasons". Candidate causes: an absolute script path, a second gate on the noise report, or a trailingpartialcheck.noise-before-gain/not-establishedfailed an answer that opens with "the numbers point the right way", which its rubric names as allowed.measurable-criterionfour-properties sample 3: the model drops its "launch a medical bot with 'answers are safe'" paragraph as unsafe advice.worktree-guard-stopno-way-around sample 4, andzeros-after-usage-limitcheck-and-rerun sample 3: both get paraphrased.measurable-criterionmust-fail sample 2 now misses two properties. It was built to miss only the trial set. Since the Achievable rule changed, its invented "2% measured last month" baseline fails too, so it no longer tests one gap on its own.source-controlhas aclaude plugin evalsuite (17 cases). 82 plugins carry skill-creatorevals.jsonfiles only (326 files).Context
The loop that produced these, and the reasoning behind each, is in the pull request that adds the evals suite and validity gate (link to follow).