Skip to content

feat: route on a calibrated classifier, live load, and measured quality - #31

Merged
github-actions[bot] merged 1 commit into
mainfrom
feat/18-calibrated-routing
Oct 3, 2026
Merged

github-actions[bot] merged 1 commit into
mainfrom
feat/18-calibrated-routing

Conversation

@Yash-Chindam

Copy link
Copy Markdown
Owner

Fourth PR closing gaps between the design spec and the implementation.

Gap this closes

§7.2: "Use a lightweight classifier or calibrated model for task and complexity prediction", plus the routing features it lists that were missing or static: structured-output requirement, current queue delay, GPU capacity, and historical quality by task and model.

What changed

  • classifier.py: multinomial naive Bayes over unigrams and bigrams, trained at start-up from config/routing/task-classifier-v1.jsonl. No new dependency. Posteriors are temperature-scaled on a held-out split; a prediction under 0.5 abstains to general. A declared routing.task is never overridden. Keyword rules remain the fallback when the dataset is absent.
  • Complexity (low/medium/high) is predicted by a second head and pulls a request toward higher measured quality.
  • routing.structured is now a hard filter against supports_structured_output on the model card.
  • Quality is looked up per task from benchmark history (BenchmarkRun.task), falling back to the card's headline figure.
  • load.py: observed queue delay (EWMA per model) replaces the catalog estimate; engine saturation from the last metrics scrape is a penalty on local models, discarded after 30 seconds.
  • The response and the trace report task_source, task_confidence, and complexity.

Limits worth knowing

  • The training and evaluation prompts were written by hand, in one voice, for this PR. The held-out result (42/42 tasks, 88% complexity, ECE 0.035) checks that the mechanism works; it is not evidence of accuracy on real traffic.
  • One gateway faces one engine, so saturation costs all local models equally. It only affects local versus external.
  • Required modality is not implemented; message content is text only.
  • The only benchmark run in the catalog covers extraction, so quality history changes nothing in the committed configuration yet.

Test plan

  • 35 new tests (classifier accuracy and calibration on unseen prompts, abstention, structured hard filter, per-task quality, queue feedback, saturation and its privacy limit)
  • ruff format --check ., ruff check ., mypy clean
  • pytest tests/unit tests/integration: 241 passed, coverage 98%
  • Playwright end-to-end suite: 9 passed locally against the running gateway

🤖 Generated with Claude Code

Section 7.2 asks for a lightweight classifier or calibrated model for task
and complexity, and lists structured-output requirement, current queue
delay, GPU capacity, and historical quality by task and model as routing
features. Task was keyword matching, complexity did not exist, queue delay
was a static catalog estimate, and quality was one figure per model.

Task and complexity now come from a multinomial naive Bayes model trained
at start-up from a committed dataset, with posteriors temperature-scaled
against a held-out split. A prediction under 0.5 abstains to the general
task, and a declared task is never overridden. Keyword rules remain the
fallback when no dataset is deployed.

Capability stays deterministic: a structured request never reaches a
model that cannot produce structured output. Scoring uses benchmarked
quality for the task, an exponentially weighted average of observed queue
delay in place of the estimate, and engine saturation from the last
scrape. Predicted complexity pulls a request toward higher measured
quality. Saturation can tip an eligible request to the external model but
never overrides privacy.

The response and the trace report how the task was established.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation area/api area/tests labels Oct 3, 2026
@github-actions
github-actions Bot merged commit 1055e9f into main Oct 3, 2026
6 checks passed
@github-actions
github-actions Bot deleted the feat/18-calibrated-routing branch October 3, 2026 14:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/api area/tests documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant