feat(perf): decide captured observed-performance policies offline - #26
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #24.
Delivered
run --policy FILEvalidates and captures exact policy bytes before dispatch, with a skip-absent plan binding. Every baseline/candidate/repeat must have captured the same policy; no retrospective binding through decide.decide A B --reference A2 --jsonreturns a versioned PASS / REGRESSION / INCONCLUSIVE / ERROR envelope, separate from measurement eligibility and conventional compare exits.Claim boundary
PASS means complete observed evidence meets the declared practical policy. It does not establish statistical significance, causal effect, future performance, universal no-regression or model quality. Reference declarations and captured policy are not authenticated restoration or independent preregistration. Only a printed decision envelope constitutes a decision; process success from run/compare is not PASS.
Verification
Independent SIX, HELP and CYBERGLM review; one malformed-unbound-collector identity emission defect was reproduced, fixed and re-reviewed. HELP independently ran exact boundary/determinism/privacy regressions. Seventeen new real-CLI policy tests; one older comparison fixture now uses explicitly synthetic fixed observations instead of a scheduler-sensitive intersection premise. Exact-head fmt, Clippy -D warnings, full workspace tests and release build passed. Combined shared workflow was exercised on a clean-source native Linux ARM64 build using synthetic loopback responses.
No inference/server/SSH/deployment calls. Live recipe qualification remains separately coordinated, DeepSeek first and GLM independently later. No claim of actual source authenticity or absence of all historical tool defects.