Problem
The live eval layer runs one model. There is no way to compare tool-use
reliability, latency, or cost across models, so model-selection and routing
changes (including sim-auto) are unmeasured.
Proposal
Run the same scenarios across a list of models and report a scenario × model
matrix.
EVAL_MODELS=provider:model,... selects the models; a bare model id defaults
to DeepSeek.
- A small provider registry (DeepSeek, OpenAI, Groq, OpenRouter) reads each key
from <PROVIDER>_API_KEY, so adding a provider is one entry.
- The report carries per-model pass rate, iterations, latency, and tokens, plus
a per-scenario pass-rate matrix.
- Opt-in live suite, never in CI.
EVAL_MIN_PASS_RATE fails a model below a
floor.
Acceptance criteria
Problem
The live eval layer runs one model. There is no way to compare tool-use
reliability, latency, or cost across models, so model-selection and routing
changes (including
sim-auto) are unmeasured.Proposal
Run the same scenarios across a list of models and report a scenario × model
matrix.
EVAL_MODELS=provider:model,...selects the models; a bare model id defaultsto DeepSeek.
from
<PROVIDER>_API_KEY, so adding a provider is one entry.a per-scenario pass-rate matrix.
EVAL_MIN_PASS_RATEfails a model below afloor.
Acceptance criteria
EVAL_MODELSruns two or more models