Need
The performance comparator reports eligibility and descriptive changes, not a regression decision. Integrations must not interpret exit 0, overlapping ranges or missing percentages as PASS.
Proposed bounded change
Add an explicit policy captured with each run before dispatch and an offline decide command over verified baseline/candidate/reference directories. Policy pins the binary/workload, all required cells and metrics, minimum trials, practical loss tolerance and reference-spread allowance. No implicit recipe tolerances.
Decision scope is observed policy compliance, not statistical significance, causality or universal no-regression. Use checked rational comparisons of raw verified measurements: latency higher is worse, rates lower are worse. Incomplete required evidence, unusable/reused repeat or excessive reference variability -> INCONCLUSIVE; invalid inputs -> ERROR. PASS requires all declared gates; a qualified REGRESSION cannot be offset by another metric improving. Return a versioned envelope with reason codes, evidence/policy/evaluator identities and eligibility separately.
Preserve current run/compare semantics and legacy receipt bytes when the option is absent. No new collector, deployment operations, automatic retries, model-quality scoring or dashboard.
Acceptance
Real CLI fixtures cover all outcomes and exact boundaries, policy capture/tamper/pins, ordered output and metric coverage, reference qualification, deterministic output, legacy load and privacy-safe machine reason codes. Independent correctness/resource/integrity review and exact-head CI. Actual recipe live qualification is separate and requires a coordinated serving window.
Need
The performance comparator reports eligibility and descriptive changes, not a regression decision. Integrations must not interpret exit 0, overlapping ranges or missing percentages as PASS.
Proposed bounded change
Add an explicit policy captured with each run before dispatch and an offline
decidecommand over verified baseline/candidate/reference directories. Policy pins the binary/workload, all required cells and metrics, minimum trials, practical loss tolerance and reference-spread allowance. No implicit recipe tolerances.Decision scope is observed policy compliance, not statistical significance, causality or universal no-regression. Use checked rational comparisons of raw verified measurements: latency higher is worse, rates lower are worse. Incomplete required evidence, unusable/reused repeat or excessive reference variability -> INCONCLUSIVE; invalid inputs -> ERROR. PASS requires all declared gates; a qualified REGRESSION cannot be offset by another metric improving. Return a versioned envelope with reason codes, evidence/policy/evaluator identities and eligibility separately.
Preserve current run/compare semantics and legacy receipt bytes when the option is absent. No new collector, deployment operations, automatic retries, model-quality scoring or dashboard.
Acceptance
Real CLI fixtures cover all outcomes and exact boundaries, policy capture/tamper/pins, ordered output and metric coverage, reference qualification, deterministic output, legacy load and privacy-safe machine reason codes. Independent correctness/resource/integrity review and exact-head CI. Actual recipe live qualification is separate and requires a coordinated serving window.