Skip to content

feat(agent-cli): forward Switchyard candidate models to the in-sandbox pi agent (2/2) - #129

Merged
ayush-or merged 4 commits into
mainfrom
codex/sandbox-candidate-models
Oct 1, 2026
Merged

ayush-or merged 4 commits into
mainfrom
codex/sandbox-candidate-models

Conversation

@ayush-or

@ayush-or ayush-or commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

TL;DR

Terminal-Bench (and deep_swe / swe_atlas with an Ori agent) can now send Switchyard's two candidate models from inside the sandbox. Before, any run that set models was rejected because the in-sandbox agent never put a models array on its requests, so nvidia/switchyard had nothing to route between. Stacked on #127, which forwards the switchyard-router plugin the same way.

What changed?

  • New env var TB_OPENROUTER_MODELS. buildAgentCliEnv writes opts.models as JSON when it is set.
  • The existing pi before_provider_request extension sets payload.models from that env var, replacing any models already on the request, next to the plugin merge from fix(terminal-bench): forward auto-router cost tier to the in-sandbox pi agent #126/feat(agent-cli): forward the switchyard-router plugin through sandbox request plugins (1/2) #127. The extension is written and loaded when plugins or models are configured (hasRequestPlugins is renamed to loadsRequestExtension).
  • sandboxAgentPluginError accepts models for pi. It still rejects models for claude, prime-agent and omp, which cannot rewrite outgoing requests. The runner keeps a matching defensive SolverError.
  • terminal_bench, deep_swe and swe_atlas pass benchmarkConfig.models into the agent CLI options.
  • The extension removes max_output_tokens when it sets models. pi sizes that field from its catalog entry for the requested model. nvidia/switchyard isn't in the catalog, so pi used a default of 235,929 tokens. That is above the 128,000 limit of GPT-6-Sol, GPT-6-Astra and Opus 5.5, so routing dropped every endpoint of the capable model. Switchyard still picked it, but the request was silently served by the efficient model. Without the cap, each candidate uses its own provider default.

Resulting in-sandbox request body for a Switchyard cell:

{
  "model": "nvidia/switchyard",
  "models": ["z-ai/glm-5.3-flash", "anthropic/claude-opus-5.5"],
  "plugins": [{ "id": "switchyard-router", "algorithm": "stage" }]
}

How to test

  • bun test src/benchmarks/agent-cli src/benchmarks/terminal-bench src/benchmarks/deep-swe src/benchmarks/swe-atlas: 180 pass. New cases:
    • the extension adds models and the switchyard-router plugin together, and adds models alone when no plugins are configured;
    • the pi solver env carries TB_MODEL=nvidia/switchyard, TB_OPENROUTER_MODELS and TB_OPENROUTER_PLUGINS, and loads the extension;
    • the terminal_bench layer accepts models + switchyardAlgorithm for pi and rejects models for claude.
  • Full bun test: 1599 pass, 2 fail. The same 2 fail on the base branch without this change (makeResponsesLayer > serializes input_video.processing=*).
  • bun run typecheck and bun run check pass.

Live validation (commit 1410e8c, production OpenRouter, default deployed runtime)

1. Real ori pi + pi 0.84.2 on this machine, running the exact run script and env generated by harness.buildRunScript / buildAgentCliEnv. A capture proxy recorded each request body and the switchyard-router pipeline entry from x-openrouter-metadata. The task was a two-step prompt with a failing tool call. Every request carried model: nvidia/switchyard, both candidate models, and plugins: [{id: "switchyard-router", algorithm: "stage"}].

Pairing Strategy Picked → served
DeepSeek v4.1 Flash + GPT-6-Sol stage Sol → Sol, then Flash → Flash ×2
GLM 5.3 Flash + Opus 5.5 stage GLM → GLM ×2
GLM 5.3 Flash + Kimi K3 caller_order (equal_pricing) GLM ×3. No routing decision.
GPT-6-Sol + GPT-6-Astra stage Sol → Sol ×2

Before the max_output_tokens fix, the same run showed DeepSeek + Sol picking Sol but serving Flash. Replaying pi's captured body with algorithm: random and 8 fresh sessions per arm:

Pairing pi's 235,929 cap Cap removed
DeepSeek + Sol picked Sol 5, served Sol 0 picked Sol 5, served Sol 5
GLM + Opus picked Opus 2, served Opus 0 picked Opus 2, served Opus 2

2. Real Terminal-Bench on Modal through the harness CLI (bun src/cli/index.ts --benchmark terminal_bench --model nvidia/switchyard --solver-config '{…, "agent":"pi", "taskSubset":["fix-git"]}'):

Pairing Result Turns Served models (from /generation) Cost
GLM 5.3 Flash + Opus 5.5 ✅ C 8 Opus 5.5 ×8, router=nvidia/switchyard $0.150
DeepSeek v4.1 Flash + GPT-6-Sol ✅ C 14 GPT-6-Sol ×14, router=nvidia/switchyard $0.063

Total live spend for this validation was under $1.

Reviewer focus

@ayush-or
ayush-or requested a review from a team as a code owner October 1, 2026 13:15
@ayush-or
ayush-or added this pull request to stack #130 October 1, 2026 13:18
Base automatically changed from devin/1790733057-sandbox-switchyard-plugin to main October 1, 2026 18:06
@ayush-or
ayush-or merged commit 1a8b00b into main Oct 1, 2026
4 checks passed
@ayush-or
ayush-or deleted the codex/sandbox-candidate-models branch October 1, 2026 18:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant