Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 38 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,44 @@ workflow tags `main` (`vMAJOR.MINOR.PATCH`) and publishes a GitHub Release with
auto-generated notes. Tags are never created by hand, and nothing is ever tagged off a
branch other than `main`.

## Routing

Privacy, tenant entitlement, context size, and capability are hard filters: deterministic, and
applied before any score is computed. Among the models that remain, the router scores measured
quality for the task against observed queue delay, cost, engine saturation, and how complex the
request is predicted to be.

**Task and complexity** come from a calibrated classifier — multinomial naive Bayes over word
unigrams and bigrams, trained at start-up from
[`config/routing/task-classifier-v1.jsonl`](config/routing/task-classifier-v1.jsonl). It needs no
accelerator and no extra dependency, so a routing decision never waits on the models it is
choosing between. Its posteriors are temperature-scaled against a held-out split, so a confidence
reads as a probability, and a prediction under 0.5 abstains to the `general` task rather than
being trusted. A task the caller declares in `routing.task` is never overridden. Without the
dataset the router falls back to keyword rules.

On the held-out prompts in
[`benchmarks/datasets/routing-tasks-v1.jsonl`](benchmarks/datasets/routing-tasks-v1.jsonl) the
classifier gets every task right, 88% of complexity labels, and an expected calibration error of
0.035. Both datasets are small and were written by hand in one voice, so treat that as a floor
check on the mechanism, not as evidence of accuracy on real traffic: replace them with labelled
production prompts before relying on the numbers.

| Routing feature | Source |
|---|---|
| Task and complexity | The classifier; `low` leaves work on the cheapest capable model, `high` outweighs the specialization and cost terms. |
| Structured-output requirement | `routing.structured` excludes any model whose card sets `supports_structured_output: false`. |
| Quality by task and model | Mean benchmarked quality for the task from the catalog; the card's headline `quality` only for a task never measured. |
| Current queue delay | An exponentially weighted average of what requests for that model actually waited, replacing the catalog estimate after the first observation. |
| GPU capacity | Engine saturation — the worse of KV-cache occupancy and the share of admitted work not yet started — from the last metrics scrape, discarded after 30 seconds. |

One gateway faces one engine, so saturation costs every local model equally: it can tip an
eligible request to the approved external model, and it never reorders local models or overrides
privacy. Required modality is not a routing feature yet; message content is text only.

Every response reports `task_source` (`declared`, `classifier`, `abstained`, `keyword`, or
`cached`), `task_confidence`, and `complexity` beside the route reason.

## Serving backends

`ROUTER_BACKEND=mock` (the default) keeps CI deterministic and GPU-free.
Expand Down
42 changes: 42 additions & 0 deletions benchmarks/datasets/routing-tasks-v1.jsonl
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
{"prompt": "Extract the order number and delivery date from this confirmation.", "task": "extraction", "complexity": "low"}
{"prompt": "Return the patient name and appointment time as JSON.", "task": "extraction", "complexity": "low"}
{"prompt": "Pull the VAT number out of the supplier letter.", "task": "extraction", "complexity": "low"}
{"prompt": "Extract the amount due from this bill.", "task": "extraction", "complexity": "low"}
{"prompt": "Parse the following address into its fields.", "task": "extraction", "complexity": "low"}
{"prompt": "Extract the serial number and model from the warranty card.", "task": "extraction", "complexity": "low"}
{"prompt": "Classify this message as complaint, question or praise.", "task": "classification", "complexity": "low"}
{"prompt": "Choose one label for this support email.", "task": "classification", "complexity": "low"}
{"prompt": "Is this ticket a bug or a feature request?", "task": "classification", "complexity": "low"}
{"prompt": "Classify the sentiment of this review.", "task": "classification", "complexity": "low"}
{"prompt": "Which category does this expense belong to?", "task": "classification", "complexity": "low"}
{"prompt": "Label this transaction as personal or business.", "task": "classification", "complexity": "low"}
{"prompt": "According to the documents, what is the cancellation fee?", "task": "rag", "complexity": "medium"}
{"prompt": "Using the provided context, answer how refunds are issued.", "task": "rag", "complexity": "medium"}
{"prompt": "Answer only from the passages below: what is the SLA?", "task": "rag", "complexity": "low"}
{"prompt": "Based on the retrieved context, who signs off the release?", "task": "rag", "complexity": "medium"}
{"prompt": "Use the provided context to answer the question about holidays.", "task": "rag", "complexity": "medium"}
{"prompt": "From the documents provided, what is the retention period?", "task": "rag", "complexity": "low"}
{"prompt": "Summarize the annual review.", "task": "summarization", "complexity": "medium"}
{"prompt": "Give me a summary of this long email.", "task": "summarization", "complexity": "medium"}
{"prompt": "Summarise the transcript in five bullet points.", "task": "summarization", "complexity": "medium"}
{"prompt": "Write a brief summary of the audit findings.", "task": "summarization", "complexity": "medium"}
{"prompt": "Condense the article into its key points.", "task": "summarization", "complexity": "medium"}
{"prompt": "Provide a short overview of this policy document.", "task": "summarization", "complexity": "medium"}
{"prompt": "Reason step by step about why the test fails.", "task": "reasoning", "complexity": "high"}
{"prompt": "Prove that the algorithm always halts.", "task": "reasoning", "complexity": "high"}
{"prompt": "Analyze deeply the trade-offs between these two designs.", "task": "reasoning", "complexity": "high"}
{"prompt": "Work through the puzzle and explain each step.", "task": "reasoning", "complexity": "high"}
{"prompt": "Deduce the correct ordering from the constraints.", "task": "reasoning", "complexity": "high"}
{"prompt": "Reason about whether the claim follows from the premises.", "task": "reasoning", "complexity": "high"}
{"prompt": "Critique this system design.", "task": "critique", "complexity": "high"}
{"prompt": "Find flaws in this proposal.", "task": "critique", "complexity": "high"}
{"prompt": "Review this essay and point out its weaknesses.", "task": "critique", "complexity": "medium"}
{"prompt": "Give a critical review of this architecture.", "task": "critique", "complexity": "high"}
{"prompt": "What is wrong with this argument?", "task": "critique", "complexity": "high"}
{"prompt": "Point out the weaknesses in this security plan.", "task": "critique", "complexity": "high"}
{"prompt": "Hello", "task": "general", "complexity": "low"}
{"prompt": "Tell me a joke about computers.", "task": "general", "complexity": "low"}
{"prompt": "What is the capital of Japan?", "task": "general", "complexity": "low"}
{"prompt": "Write a short poem about the sea.", "task": "general", "complexity": "medium"}
{"prompt": "Thanks for your help.", "task": "general", "complexity": "low"}
{"prompt": "Draft an email inviting the team to lunch.", "task": "general", "complexity": "medium"}
1 change: 1 addition & 0 deletions config/registry.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,7 @@ benchmarks:
container_digest: sha256:mock-container
engine_revision: mock-engine-0.1.0
model_revision: mock-small@sha256:dev
task: extraction
concurrency: 32
prompt_tokens_p50: 420
prompt_tokens_p95: 1100
Expand Down
Loading
Loading