You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
forcedToolUse: false: Sonnet 5.5 rejects tool_choicetool/any with a 400, so a forced tool is sent as auto (and the Force option is hidden in tool input), same as Opus 5.5 / Fable 5.1
Map the none thinking level to thinking: { type: "between_tools" } via a new capabilities.thinking.noneMode field. Sonnet 5.5 rejects disabled and is adaptive by default, so previously none silently ran full adaptive thinking at high. between_tools carries no effort/display (both 400). Other models keep sending no thinking config for none
Entry sits before claude-sonnet-5 so the prefix-matching catalog lookups (startsWith(id + '-')) resolve it to its own capabilities
Make Sonnet 5.5 the recommended model and the default for the Agent, Router, and Evaluator blocks, the combobox model fallback, and the Anthropic provider; mark claude-sonnet-5 legacy (Anthropic now lists it as legacy). Router/Evaluator use responseFormat (native structured outputs), not forced tools
Update docs defaults and regenerate the agent streaming table
Type of Change
New feature
Testing
Live against the Anthropic API: Models API confirms 1M input / 128K output / 5 effort levels / structured outputs / adaptive only; raw probes confirm disabled, enabled, non-default temperature, forced/any tool_choice, and between_tools + xhigh/display all 400, while every shape Sim sends returns 200; a 683-token prompt caches
Live end-to-end through executeAnthropicProviderRequest with the real SDK: none + tools (streaming and non-streaming tool loops, between_tools on every turn), high + tools streaming with agent events (adaptive summarized, thinking-block replay accepted), forced tool downgraded to auto, router-style responseFormat with temperature set, xhigh/max, streaming without tools, and prompt-cache write then read
providers/anthropic/core.test.ts: forced-tool and none → between_tools tests, both confirmed red with the catalog fields reverted
vitest run providers executor blocks lib/model-router combobox (4,055 tests), bun run type-check, bun run lint, bun run check:audits (51 audits), block registry check, docs-manifest:check
Pricing cross-checked against Anthropic pricing and OpenRouter
Checklist
Code follows project style guidelines
Self-reviewed my changes
Tests added/updated and passing (new tests pass the test-audit authoring gate)
[Medium risk] Changes default AI model across the application.
The PR appears safe to merge based on the reviewed changes and resolved prior threads.
Summary
The PR adds Claude Sonnet 5.5 to the model catalog, makes it the default for Agent, Router, and Evaluator blocks, and maps its none thinking selection to between_tools. The latest revision replaces mock-call assertions with recorded request payloads and retains multi-turn coverage. No new actionable issue was identified.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
A[Model selection] --> B[Sonnet 5.5 capabilities]
B --> C{Thinking level}
C -->|none| D[between_tools]
C -->|other supported level| E[adaptive thinking]
D --> F[Anthropic request]
E --> F
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
claude-sonnet-5-5: $2/$10 per MTok, $0.20 cache read, 1M context, 128K max output, adaptive thinking (low–max, defaulthigh), native structured outputs, 512-token cache minimumforcedToolUse: false: Sonnet 5.5 rejectstool_choicetool/anywith a 400, so a forced tool is sent asauto(and the Force option is hidden in tool input), same as Opus 5.5 / Fable 5.1nonethinking level tothinking: { type: "between_tools" }via a newcapabilities.thinking.noneModefield. Sonnet 5.5 rejectsdisabledand is adaptive by default, so previouslynonesilently ran full adaptive thinking athigh.between_toolscarries no effort/display (both 400). Other models keep sending no thinking config fornoneclaude-sonnet-5so the prefix-matching catalog lookups (startsWith(id + '-')) resolve it to its own capabilitiesclaude-sonnet-5legacy (Anthropic now lists it as legacy). Router/Evaluator useresponseFormat(native structured outputs), not forced toolsType of Change
Testing
disabled,enabled, non-defaulttemperature, forced/anytool_choice, andbetween_tools+xhigh/displayall 400, while every shape Sim sends returns 200; a 683-token prompt cachesexecuteAnthropicProviderRequestwith the real SDK:none+ tools (streaming and non-streaming tool loops,between_toolson every turn),high+ tools streaming with agent events (adaptive summarized, thinking-block replay accepted), forced tool downgraded to auto, router-styleresponseFormatwith temperature set,xhigh/max, streaming without tools, and prompt-cache write then readproviders/anthropic/core.test.ts: forced-tool andnone→between_toolstests, both confirmed red with the catalog fields revertedvitest run providers executor blocks lib/model-router combobox(4,055 tests),bun run type-check,bun run lint,bun run check:audits(51 audits), block registry check,docs-manifest:checkChecklist
test-auditauthoring gate)