Problem
In R6 (PR #5895, head 50c2769, run-validity VALID), the measurable-criterion case scored 0.33 with the evals plugin loaded. R5 had scored it 1.00 on the earlier tip. Two of three with-plugin runs failed the calibrated four-properties grader:
- Run 2 set a red-flag escalation target of "at least 99%" with no stated basis and called it a placeholder needing sign-off (judges: FAIL FAIL FAIL).
- Run 3 relied on a target relative to the current baseline, then called that target "a placeholder for you to adjust once you have data" (judges: PASS FAIL FAIL).
The rubric's Achievable property follows Anthropic's "Define your success criteria" page: a target rests on benchmarks, prior experiments, AI research or expert knowledge, and a target stated relative to the current baseline counts. The skill's answers drift from that, either by inventing a number or by withdrawing the basis they just gave.
The PR's commits only respelled this case's calibration samples, which no run reads, so this is model variance the skill does not yet constrain, not a regression.
What to do
- Read the three R6 with-arm replies (kept traces from the R6 run) beside the methodology hub's success-criteria guidance.
- Fix the methodology skill's hub so the answer states each target's basis and does not call a grounded target a placeholder. Do not ease the case or its grader.
- Recalibrate
four-properties if its samples change, then confirm with a run that run-validity calls VALID.
Related
Problem
In R6 (PR #5895, head 50c2769, run-validity VALID), the
measurable-criterioncase scored 0.33 with the evals plugin loaded. R5 had scored it 1.00 on the earlier tip. Two of three with-plugin runs failed the calibratedfour-propertiesgrader:The rubric's Achievable property follows Anthropic's "Define your success criteria" page: a target rests on benchmarks, prior experiments, AI research or expert knowledge, and a target stated relative to the current baseline counts. The skill's answers drift from that, either by inventing a number or by withdrawing the basis they just gave.
The PR's commits only respelled this case's calibration samples, which no run reads, so this is model variance the skill does not yet constrain, not a regression.
What to do
four-propertiesif its samples change, then confirm with a run that run-validity calls VALID.Related