Skip to content

feat(evals): add agent evaluation suite - #8409

Open
sudoKrishna wants to merge 2 commits into
simstudioai:stagingfrom
sudoKrishna:feat/agent-tool-use-evals
Open

sudoKrishna wants to merge 2 commits into
simstudioai:stagingfrom
sudoKrishna:feat/agent-tool-use-evals

Conversation

@sudoKrishna

@sudoKrishna sudoKrishna commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Adds a deterministic, CI-runnable evaluation layer for the agent harness, plus
opt-in live tooling. Scripted scenarios drive the real code — the
OpenAI-compatible streaming tool loop and the DAGExecutor — and score tool
selection, planning, retrieval, recovery, and context assembly. Live runs
compare models and support record/replay and rubric judging. No provider key is
required for CI.

Targets staging per CONTRIBUTING. One squash-ready commit.

What's included

Deterministic (CI, no key)

  • agent-tool-use/ — 13 tool-loop scenarios (reliability + adversarial:
    no-tool-needed, near-duplicate tools, empty results, four-tool chain) and 4
    executor scenarios (Start→Agent, variable resolution, block retry, model
    fallback)
  • agent-context/ — provider-request assembly with conversation memory
    (history content, role, and order; system prompt; conversation isolation)
  • replay.ts — record live transcripts once, replay task and judge turns
    through the real loop
  • judge.ts — LLM judge with a weighted rubric and a judge-identity envelope
    (model, rubric digest, parser version, decoding) so scores are only compared
    when the evaluator matches
  • report.ts — JSON + Markdown reports, including the model-comparison matrix

Opt-in live (EVAL_LIVE=1, never in CI)

  • live.ts / models.ts — real transport for any OpenAI-compatible provider;
    model comparison across EVAL_MODELS
  • the live suite runs the judge on scenarios with a rubric
    (EVAL_JUDGE_MODEL)

How to run

cd apps/sim
bun run test:evals            # tool-use + executor + judge + report + replay
bun run test:evals:context    # context assembly

Live (needs a key):

DEEPSEEK_API_KEY=... bun run test:evals:live
EVAL_MODELS=deepseek:deepseek-chat,deepseek:deepseek-reasoner DEEPSEEK_API_KEY=... bun run test:evals:compare
DEEPSEEK_API_KEY=... bun run test:evals:judge

The scripted suite is collected by the normal bun run test, so a regression
fails CI without the dedicated commands.

Results from live runs

  • Tool use on DeepSeek: 100% (33/33) across 11 scenarios, including the
    adversarial ones.
  • Model comparison: deepseek-chat 100% / deepseek-reasoner 89%, with the
    reasoner spending ~27% more tokens — a counterintuitive result the comparison
    surfaced.
  • Every live failure found was a test-design bug (case-sensitive phrasing,
    an over-specific id, a required input the prompt never supplied, an
    over-strict retry count), not a model failure. The adversarial cases exposed
    them.

Test plan

  • bun run test:evals → 35/35 (1 replay suite skips with no fixtures)
  • bun run test:evals:context → 2/2
  • Negative checks: executor retry, model fallback, context isolation, and
    tool-feedback forwarding each fail when broken, then reverted
  • Live runs and a key-free model-comparison report validated
  • bun run check:test-patterns, check:import-specifiers, check:boundaries,
    check:utils, check:source-text pass
  • Eval modules type-check against the real loop/executor signatures
  • Full bun run type-check — run in CI

Follow-up

  • Record real fixtures so the replay suite runs in CI (needs a key).
  • Add the frozen human-adjudicated judge calibration slice.
  • Subagent/orchestration evals (parent → child workflow).

@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 29, 2026 09:51
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 29, 2026 9:51am UTC

Request Review

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation suite for the agent tool-use loop.

The evaluation suite has no identified runtime blocker, but its import paths must satisfy the repository requirement before merging.

Findings

  1. P2 Tool feedback goes unchecked ▶
  2. P2 Tool arguments are not checked ▶
  3. P2 Relative imports violate app requirement ▶

Summary

The PR adds eight deterministic scenarios that drive the production OpenAI-compatible streaming tool loop, score outcomes, and write JSON and Markdown reports.

  • The suite exercises dispatch and error accounting without a provider key.
  • Its scripted model does not inspect tool feedback, and scoring does not check dispatched arguments, limiting the regressions it can detect.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Scenario script] --> B[Scripted model turns]
  B --> C[Production streaming tool loop]
  C --> D[Stub tool results]
  D --> C
  C --> E[Scoring]
  E --> F[JSON and Markdown reports]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add agent tool-use evaluati..."

Comment thread apps/sim/evals/agent-tool-use/harness.ts Outdated
Comment thread apps/sim/evals/agent-tool-use/harness.ts
Comment thread apps/sim/evals/agent-tool-use/harness.ts
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@sudoKrishna

Copy link
Copy Markdown
Author

Addressed all three findings in f619479:

  • the scripted model now reads and asserts tool feedback, so a dropped result fails the case;
  • a new tool-arguments check compares each executed call's arguments to the expected ones;
  • eval imports are absolute now.

Deterministic, CI-runnable evaluation for the agent harness plus opt-in live
tooling. Scripted scenarios drive the real OpenAI-compatible streaming tool
loop and the DAGExecutor; scenarios cover tool selection, planning, retrieval,
recovery, adversarial cases, and conversation-context assembly. Live runs
compare models and support record/replay and rubric judging.

- agent-tool-use/: 13 loop scenarios + 4 executor scenarios + reports
- agent-context/: provider-request assembly with conversation memory
- replay.ts / judge.ts / models.ts: record-replay, LLM judge, model registry
- test:evals, test:evals:{live,record,context,compare,judge}
- rebased onto staging per CONTRIBUTING (PRs target staging, not main)
@sudoKrishna
sudoKrishna force-pushed the feat/agent-tool-use-evals branch from f619479 to 39f1a5a Compare October 4, 2026 15:40
@sudoKrishna
sudoKrishna changed the base branch from main to staging October 4, 2026 15:42
@sudoKrishna

Copy link
Copy Markdown
Author

Rewrote this branch onto the current staging as a single commit, per CONTRIBUTING (PRs target staging; the previous history was based on an old main). The suite still passes on the new base — 35/35 + 2/2, plus the static audits.
The earlier Greptile findings (tool feedback, tool-argument checks, relative imports) are all addressed in this commit; those comments are from the previous history and are outdated.
Ready for review when the workflows can be approved.

…ioai#8570)

Adds an agent-orchestration suite: a parent workflow invokes a child through
the Workflow block, driven by the real DAGExecutor. The child definition loader
and child Executor are mocked, so spawn -> wait -> aggregate -> recover runs
without a database; the parent executor, workflow handler, serializer, and
output mapping are real.

- two cases: child output reaches the parent; a child failure fails the parent
- report via the shared writeEvalReport; test:evals:orchestration script
- README documents the suite

Held off simstudioai#8409: the suite needs its eval infrastructure and upstream takes no
stacked PRs, so it opens after simstudioai#8409 merges.

Refs simstudioai#8570

This branch was previously deployed

1 inactive (outdated) deployment
Preview — 86c79d78 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant