feat: route on a calibrated classifier, live load, and measured quality - #31
Merged
Merged
Conversation
Section 7.2 asks for a lightweight classifier or calibrated model for task and complexity, and lists structured-output requirement, current queue delay, GPU capacity, and historical quality by task and model as routing features. Task was keyword matching, complexity did not exist, queue delay was a static catalog estimate, and quality was one figure per model. Task and complexity now come from a multinomial naive Bayes model trained at start-up from a committed dataset, with posteriors temperature-scaled against a held-out split. A prediction under 0.5 abstains to the general task, and a declared task is never overridden. Keyword rules remain the fallback when no dataset is deployed. Capability stays deterministic: a structured request never reaches a model that cannot produce structured output. Scoring uses benchmarked quality for the task, an exponentially weighted average of observed queue delay in place of the estimate, and engine saturation from the last scrape. Predicted complexity pulls a request toward higher measured quality. Saturation can tip an eligible request to the external model but never overrides privacy. The response and the trace report how the task was established. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fourth PR closing gaps between the design spec and the implementation.
Gap this closes
§7.2: "Use a lightweight classifier or calibrated model for task and complexity prediction", plus the routing features it lists that were missing or static: structured-output requirement, current queue delay, GPU capacity, and historical quality by task and model.
What changed
classifier.py: multinomial naive Bayes over unigrams and bigrams, trained at start-up fromconfig/routing/task-classifier-v1.jsonl. No new dependency. Posteriors are temperature-scaled on a held-out split; a prediction under 0.5 abstains togeneral. A declaredrouting.taskis never overridden. Keyword rules remain the fallback when the dataset is absent.low/medium/high) is predicted by a second head and pulls a request toward higher measured quality.routing.structuredis now a hard filter againstsupports_structured_outputon the model card.BenchmarkRun.task), falling back to the card's headline figure.load.py: observed queue delay (EWMA per model) replaces the catalog estimate; engine saturation from the last metrics scrape is a penalty on local models, discarded after 30 seconds.task_source,task_confidence, andcomplexity.Limits worth knowing
Test plan
ruff format --check .,ruff check .,mypycleanpytest tests/unit tests/integration: 241 passed, coverage 98%🤖 Generated with Claude Code