Repository navigation
feat(evals): add agent evaluation suite - #8409
sudoKrishna wants to merge 2 commits into
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
|
Addressed all three findings in f619479:
|
Deterministic, CI-runnable evaluation for the agent harness plus opt-in live
tooling. Scripted scenarios drive the real OpenAI-compatible streaming tool
loop and the DAGExecutor; scenarios cover tool selection, planning, retrieval,
recovery, adversarial cases, and conversation-context assembly. Live runs
compare models and support record/replay and rubric judging.
- agent-tool-use/: 13 loop scenarios + 4 executor scenarios + reports
- agent-context/: provider-request assembly with conversation memory
- replay.ts / judge.ts / models.ts: record-replay, LLM judge, model registry
- test:evals, test:evals:{live,record,context,compare,judge}
- rebased onto staging per CONTRIBUTING (PRs target staging, not main)
f619479 to
39f1a5a
Compare
|
Rewrote this branch onto the current staging as a single commit, per CONTRIBUTING (PRs target staging; the previous history was based on an old main). The suite still passes on the new base — 35/35 + 2/2, plus the static audits. |
…ioai#8570) Adds an agent-orchestration suite: a parent workflow invokes a child through the Workflow block, driven by the real DAGExecutor. The child definition loader and child Executor are mocked, so spawn -> wait -> aggregate -> recover runs without a database; the parent executor, workflow handler, serializer, and output mapping are real. - two cases: child output reaches the parent; a child failure fails the parent - report via the shared writeEvalReport; test:evals:orchestration script - README documents the suite Held off simstudioai#8409: the suite needs its eval infrastructure and upstream takes no stacked PRs, so it opens after simstudioai#8409 merges. Refs simstudioai#8570
Summary
Adds a deterministic, CI-runnable evaluation layer for the agent harness, plus
opt-in live tooling. Scripted scenarios drive the real code — the
OpenAI-compatible streaming tool loop and the
DAGExecutor— and score toolselection, planning, retrieval, recovery, and context assembly. Live runs
compare models and support record/replay and rubric judging. No provider key is
required for CI.
Targets
stagingper CONTRIBUTING. One squash-ready commit.What's included
Deterministic (CI, no key)
agent-tool-use/— 13 tool-loop scenarios (reliability + adversarial:no-tool-needed, near-duplicate tools, empty results, four-tool chain) and 4
executor scenarios (Start→Agent, variable resolution, block retry, model
fallback)
agent-context/— provider-request assembly with conversation memory(history content, role, and order; system prompt; conversation isolation)
replay.ts— record live transcripts once, replay task and judge turnsthrough the real loop
judge.ts— LLM judge with a weighted rubric and a judge-identity envelope(model, rubric digest, parser version, decoding) so scores are only compared
when the evaluator matches
report.ts— JSON + Markdown reports, including the model-comparison matrixOpt-in live (
EVAL_LIVE=1, never in CI)live.ts/models.ts— real transport for any OpenAI-compatible provider;model comparison across
EVAL_MODELS(
EVAL_JUDGE_MODEL)How to run
Live (needs a key):
The scripted suite is collected by the normal
bun run test, so a regressionfails CI without the dedicated commands.
Results from live runs
adversarial ones.
deepseek-chat100% /deepseek-reasoner89%, with thereasoner spending ~27% more tokens — a counterintuitive result the comparison
surfaced.
an over-specific id, a required input the prompt never supplied, an
over-strict retry count), not a model failure. The adversarial cases exposed
them.
Test plan
bun run test:evals→ 35/35 (1 replay suite skips with no fixtures)bun run test:evals:context→ 2/2tool-feedback forwarding each fail when broken, then reverted
bun run check:test-patterns,check:import-specifiers,check:boundaries,check:utils,check:source-textpassbun run type-check— run in CIFollow-up