Skip to content

feat(evals): compare agent tool-use across models #8471

Description

@sudoKrishna

Problem

The live eval layer runs one model. There is no way to compare tool-use
reliability, latency, or cost across models, so model-selection and routing
changes (including sim-auto) are unmeasured.

Proposal

Run the same scenarios across a list of models and report a scenario × model
matrix.

  • EVAL_MODELS=provider:model,... selects the models; a bare model id defaults
    to DeepSeek.
  • A small provider registry (DeepSeek, OpenAI, Groq, OpenRouter) reads each key
    from <PROVIDER>_API_KEY, so adding a provider is one entry.
  • The report carries per-model pass rate, iterations, latency, and tokens, plus
    a per-scenario pass-rate matrix.
  • Opt-in live suite, never in CI. EVAL_MIN_PASS_RATE fails a model below a
    floor.

Acceptance criteria

  • EVAL_MODELS runs two or more models
  • Comparison JSON/Markdown with per-model metrics and a scenario matrix
  • A model is selectable by env var with no code change
  • README documents the spec format and supported providers

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions