Skip to content

feat(evals): close the eval suite's calibration gaps and convert other plugins' evals #5879

Description

@kyle-sexton

Problem

The evals plugin's eval suite reached a VALID run with the plugin scoring 0.99 and 0.31 without (delta +0.68, 95% interval +0.52 to +0.83; 29 of 31 cases at 1.00). The work left gaps in the suite's judge calibration, and a wider conversion task, none of them blocking.

Follow-ups

  1. Two judges fail correct answers that their calibration samples don't cover.
    • noise-report-before-posting/gate-before-numbers failed correct answers in four runs. Every one used the shape "run run-validity.py first; on INVALID, post the INVALID line and its reasons". Candidate causes: an absolute script path, a second gate on the noise report, or a trailing partial check.
    • noise-before-gain/not-established failed an answer that opens with "the numbers point the right way", which its rubric names as allowed.
    • Add must-pass samples of these shapes, then recalibrate.
  2. Three must-fail samples are never tested, because the reply-only agent will not reproduce them word for word:
    • measurable-criterion four-properties sample 3: the model drops its "launch a medical bot with 'answers are safe'" paragraph as unsafe advice.
    • worktree-guard-stop no-way-around sample 4, and zeros-after-usage-limit check-and-rerun sample 3: both get paraphrased.
    • Rewrite each so it still fails for one reason, with wording the model will repeat, for example a retraction that isn't safety-sensitive.
  3. measurable-criterion must-fail sample 2 now misses two properties. It was built to miss only the trial set. Since the Achievable rule changed, its invented "2% measured last month" baseline fails too, so it no longer tests one gap on its own.
  4. Convert the other plugins' evals (V7). Only source-control has a claude plugin eval suite (17 cases). 82 plugins carry skill-creator evals.json files only (326 files).

Context

The loop that produced these, and the reasoning behind each, is in the pull request that adds the evals suite and validity gate (link to follow).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-humanHuman-in-the-loop required; autonomous sessions must not resolve items carrying this.priority: lowNice-to-have, cosmetic, or speculative; opportunistic.status: needs-decisionAwaiting a human or maintainer judgment call.work-class: structuralRefactors, migrations, contract changes; cross-cutting and hard to reverse.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions