You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Establish what TheGrill's simple comparison can reliably conclude within a finite budget, and make unsupported conclusions difficult to mistake for a regression certificate. Keep the user experience simple; users should not have to design statistical studies or tune significance parameters.
Existing foundation — this is not another policy engine
#24 / #26 delivered the captured observed-envelope policy. #28 delivered the separate baseline/check capture-period comparison and measured-period presentation labels. #14 records earlier concerns about unchanged-run variability; its original suggestions are not a mandate to reintroduce obsolete heuristics.
The current baseline/check method uses acquisition summaries and a nominal model interval assuming stable, independent acquisition-level variation. The observed-envelope policy answers a different question. Neither matching declarations nor fresh clients establish independence, and more extrema can widen an envelope rather than improve precision.
The gap is calibration and claim discipline, not missing verdict labels or a reason to build a statistics framework.
Bounded work
Evaluate the existing simple structured-C1 method first. State separately what would support:
detecting a directional change versus zero;
ruling out a regression worse than a prospective tolerance;
establishing an improvement exceeding a prospective useful-effect threshold.
Do not let one label stand for all three. Retain measured-period language unless stronger claims are actually supported.
Create a small, reproducible offline calibration corpus/driver with known timing relationships: unchanged IID observations, correlated/drifting observations, controlled slowdowns, heterogeneous variance, missing/interrupted evidence and zero estimated variance. Use existing arithmetic and lightweight independent calculations; avoid a new production dependency or general simulation platform.
Quantify the relevant false-direction/coverage behaviour under stated assumptions. A directional unchanged control can occur by chance; do not demand zero false positives or treat one control result as proof of a software defect. Distinguish formula correctness, simulated coverage and real-serving qualification.
Treat output variation explicitly. Equal prompts and declared settings need not produce byte-identical output. Define semantic requirements, actual reported token exposure, and how output differences limit interpretation before collection. A harmless blank line must not silently become a serving-regression claim; neither should failures or genuine content changes be normalized away after seeing a result.
For optional speculative diagnostics, define the numerator/denominator and minimum usable exposure. Short responses may have coarse ratios and accepted-but-not-emitted terminal work. Retain raw counts, missing/reset states and scope; do not clamp inconsistent counters into plausible rates or claim a bookkeeping bracket is a confidence interval. Longer output alone does not establish independence or remove every bias.
If evidence demonstrates an actual product defect or misleading claim, implement the smallest correction and version any changed measurement/decision contract. A replacement inference method needs a concise reviewed specification before implementation. Do not silently change historical policies, verdicts, thresholds or exit semantics.
Acceptance criteria
Publish a concise calibration report tied to exact code/workload identities, assumptions, sample units, budgets, observed coverage/error behaviour and limitations. Results are reproducible offline without model access.
The existing simple CLI/report makes scope, observed effect, uncertainty and unavailable evidence clear without requiring users to manipulate statistical settings.
Known deterministic arithmetic/boundary failures have behaviour-based regression tests; expensive/stochastic calibration work is separate from the fast default test suite.
Requests/lanes are not counted as independent repetitions when they share a wave; acquisitions/windows are grouped at the justified level. Wider multi-cell claims explicitly address coverage/multiplicity rather than cherry-picking a passing cell.
No decision claims guaranteed 5% sensitivity, equivalence or noninferiority merely because collection completed, an interval overlaps zero, or an unchanged control was inconclusive.
Sample count, exclusions and stopping rules are prospective and budgeted. If precision is unaffordable, report that limitation; no automatic extensions, replacements, outcome-conditioned exclusions or retry-until-PASS.
Historical receipts and their original interpretation remain immutable. Independent arithmetic/method review distinguishes verified facts from model assumptions.
Lean implementation bar / non-goals
Choose one clear supported default path; reuse current collectors, schemas and assessment boundaries. Research must terminate in a bounded conclusion and, only where justified, a small product correction. Do not turn this into an endless effort to eliminate natural serving variation.
No parallel decision engine, statistics-option matrix, dashboard, benchmark suite expansion, automatic server tuning, learned classifier, universal regression threshold, changed defaults solely to obtain a pass, or automatic PR merge gate. Do not copy C1's interval blindly into concurrent/history workloads.
Any necessary live calibration requires its own finite, coordinated budget and cannot be substituted with simulated proof. This issue does not preauthorize GPU campaigns. Decision calibration should accompany each explicitly added workload before maintainers rely on a stronger mandatory gate.
Keep each issue within its stated scope. Agree once on shared workload-selection and versioning interfaces before overlapping implementation; reuse the same collection/evidence path. Recipe integration can land first, with concurrency and conversation coverage added deliberately. Calibration accompanies each scope before stronger decision claims become a merge requirement; it is not permission for an open-ended benchmark campaign.
Goal
Establish what TheGrill's simple comparison can reliably conclude within a finite budget, and make unsupported conclusions difficult to mistake for a regression certificate. Keep the user experience simple; users should not have to design statistical studies or tune significance parameters.
Existing foundation — this is not another policy engine
#24 / #26 delivered the captured observed-envelope policy. #28 delivered the separate baseline/check capture-period comparison and measured-period presentation labels. #14 records earlier concerns about unchanged-run variability; its original suggestions are not a mandate to reintroduce obsolete heuristics.
The current baseline/check method uses acquisition summaries and a nominal model interval assuming stable, independent acquisition-level variation. The observed-envelope policy answers a different question. Neither matching declarations nor fresh clients establish independence, and more extrema can widen an envelope rather than improve precision.
The gap is calibration and claim discipline, not missing verdict labels or a reason to build a statistics framework.
Bounded work
Do not let one label stand for all three. Retain measured-period language unless stronger claims are actually supported.
Acceptance criteria
Lean implementation bar / non-goals
Choose one clear supported default path; reuse current collectors, schemas and assessment boundaries. Research must terminate in a bounded conclusion and, only where justified, a small product correction. Do not turn this into an endless effort to eliminate natural serving variation.
No parallel decision engine, statistics-option matrix, dashboard, benchmark suite expansion, automatic server tuning, learned classifier, universal regression threshold, changed defaults solely to obtain a pass, or automatic PR merge gate. Do not copy C1's interval blindly into concurrent/history workloads.
Any necessary live calibration requires its own finite, coordinated budget and cannot be substituted with simulated proof. This issue does not preauthorize GPU campaigns. Decision calibration should accompany each explicitly added workload before maintainers rely on a stronger mandatory gate.
Related follow-ups and shared boundaries
Keep each issue within its stated scope. Agree once on shared workload-selection and versioning interfaces before overlapping implementation; reuse the same collection/evidence path. Recipe integration can land first, with concurrency and conversation coverage added deliberately. Calibration accompanies each scope before stronger decision claims become a merge requirement; it is not permission for an open-ended benchmark campaign.