Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion apps/docs/content/docs/agents/index.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ The example throughout is an agent that scores inbound sales leads.

## The Agent block

You set up the reasoning step by giving the Agent block a **model** and a **prompt**. The model is the LLM that powers it; you pick one from the available providers, and the default is `claude-sonnet-5`. The prompt is a system message that defines who the agent is and how it should behave, plus a user message that carries the input, usually a reference like `<start.input>`.
You set up the reasoning step by giving the Agent block a **model** and a **prompt**. The model is the LLM that powers it; you pick one from the available providers, and the default is `claude-sonnet-5-5`. The prompt is a system message that defines who the agent is and how it should behave, plus a user message that carries the input, usually a reference like `<start.input>`.

When it runs, the Agent block reasons, calls any tools it needs, and stores its result under its own name. By default that result is free text in `content`, read by a later block as `<agent.content>`, alongside run details like the model used, token counts, tool calls, and cost. Every setting and output field is in the [Agent block reference](/workflows/blocks/agent).

Expand Down
2 changes: 1 addition & 1 deletion apps/docs/content/docs/getting-started/index.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Build a people research agent in 10 minutes. It takes a name through a chat inte
- **System**: "You are a people research agent. When given a person's name, use your search tools to find their location, profession, educational background, and other relevant details."
- **User**: insert `<start.input>` so the agent reads whatever the chat receives.

Leave the **Model** on the default (`claude-sonnet-5`), or pick any other.
Leave the **Model** on the default (`claude-sonnet-5-5`), or pick any other.

<div className="mx-auto w-full overflow-hidden rounded-lg">
<Video src="getting-started/started-2.mp4" width={700} height={450} />
Expand Down
4 changes: 2 additions & 2 deletions apps/docs/content/docs/workflows/blocks/agent.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Answer in two sentences, cite the doc you used, and never guess a price.

### Model

The model that runs the step. Defaults to `claude-sonnet-5`. Type or pick any model from OpenAI, Anthropic, Google, xAI, Groq, Cerebras, DeepSeek, Azure, AWS Bedrock, Google Vertex, or OpenRouter, or a local model through Ollama or VLLM.
The model that runs the step. Defaults to `claude-sonnet-5-5`. Type or pick any model from OpenAI, Anthropic, Google, xAI, Groq, Cerebras, DeepSeek, Azure, AWS Bedrock, Google Vertex, or OpenRouter, or a local model through Ollama or VLLM.

For a custom cloud deployment, enter its provider prefix and model ID: `azure/my-deployment`, `azure-anthropic/my-deployment`, `bedrock/my-inference-profile`, or `vertex/my-gemini-model`. The prefix selects the provider and shows its credential fields even when the ID is absent from the catalog. Bedrock accepts full inference profile ARNs after `bedrock/`; Vertex uses the Gemini API and accepts Google model resource names. The deployment must support the selected provider's API. Custom IDs have no catalog pricing or token limits.

Expand Down Expand Up @@ -159,7 +159,7 @@ Live tool-call chips stream for **OpenAI, Anthropic, Azure Anthropic, Google, Ve
| Provider | Streamed thinking | Models |
|----------|-------------------|--------|
| OpenAI | Summaries only — Requires OpenAI organization verification; falls back to no summaries. | `gpt-6-astra`, `gpt-6-sol`, `gpt-6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5-pro`, `gpt-5.5`, `gpt-5.4-pro`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-5.3-codex`, `gpt-5.2-pro`, `gpt-5.2`, `gpt-5.1`, `gpt-5-pro`, `gpt-5`, `gpt-5-mini`, `gpt-5-nano`, `o4-mini`, `o3`, `o3-mini`, `o1` |
| Anthropic | Summaries only — These generations omit full thinking; Sim requests summarized thinking on streaming runs. | `claude-fable-5-1`, `claude-fable-5`, `claude-sonnet-5`, `claude-opus-5-5`, `claude-opus-5`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-opus-4-6`, `claude-sonnet-4-6`, `claude-opus-4-5`, `claude-sonnet-4-5`, `claude-haiku-4-5` |
| Anthropic | Summaries only — These generations omit full thinking; Sim requests summarized thinking on streaming runs. | `claude-fable-5-1`, `claude-fable-5`, `claude-sonnet-5-5`, `claude-sonnet-5`, `claude-opus-5-5`, `claude-opus-5`, `claude-opus-4-8`, `claude-opus-4-7`, `claude-opus-4-6`, `claude-sonnet-4-6`, `claude-opus-4-5`, `claude-sonnet-4-5`, `claude-haiku-4-5` |
| Azure OpenAI | Summaries only — Requires OpenAI organization verification; falls back to no summaries. | `azure/gpt-6-astra`, `azure/gpt-5.6-sol`, `azure/gpt-5.6-terra`, `azure/gpt-5.6-luna`, `azure/gpt-5.5`, `azure/gpt-5.4-pro`, `azure/gpt-5.4`, `azure/gpt-5.4-mini`, `azure/gpt-5.4-nano`, `azure/gpt-5.2`, `azure/gpt-5.1`, `azure/gpt-5.1-codex`, `azure/gpt-5`, `azure/gpt-5-mini`, `azure/gpt-5-nano`, `azure/o3`, `azure/o4-mini` |
| Azure Anthropic | Summaries only — These generations omit full thinking; Sim requests summarized thinking on streaming runs. | `azure-anthropic/claude-fable-5-1`, `azure-anthropic/claude-opus-5`, `azure-anthropic/claude-opus-4-8`, `azure-anthropic/claude-opus-4-7`, `azure-anthropic/claude-opus-4-6`, `azure-anthropic/claude-opus-4-5`, `azure-anthropic/claude-sonnet-5`, `azure-anthropic/claude-sonnet-4-6`, `azure-anthropic/claude-sonnet-4-5`, `azure-anthropic/claude-opus-4-1`, `azure-anthropic/claude-haiku-4-5` |
| Google | Summaries only | `gemini-3.8-flash`, `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash-lite`, `gemini-3.5-flash`, `gemini-3.1-pro-preview`, `gemini-3.1-flash-lite`, `gemini-3-flash-preview`, `gemini-2.5-pro`, `gemini-2.5-flash`, `gemini-2.5-flash-lite` |
Expand Down
2 changes: 1 addition & 1 deletion apps/docs/content/docs/workflows/blocks/evaluator.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ The content to score. Usually an earlier output like `<agent.content>`. Structur

### Model

The model that does the scoring, defaulting to `claude-sonnet-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you.
The model that does the scoring, defaulting to `claude-sonnet-5-5`. Stronger reasoning models give more consistent scores. Type or pick any supported model. **Temperature** and a **System Prompt** are available under advanced, and on hosted Sim the API key is supplied for you.

### Fallback models

Expand Down
2 changes: 1 addition & 1 deletion apps/docs/content/docs/workflows/blocks/router.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ Each route is a **title** and a **description** of when to choose it ("Route her

### Model

The model that makes the decision, defaulting to `claude-sonnet-5`. Stronger reasoning models route more accurately; a faster, cheaper model is fine when the routes are clearly distinct. Type or pick any supported model, or a local one through Ollama or VLLM. On hosted Sim the API key is supplied for you.
The model that makes the decision, defaulting to `claude-sonnet-5-5`. Stronger reasoning models route more accurately; a faster, cheaper model is fine when the routes are clearly distinct. Type or pick any supported model, or a local one through Ollama or VLLM. On hosted Sim the API key is supplied for you.

### Fallback models

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,7 @@ export const BLOCKS: BlockDef[] = [
bgColor: 'var(--text-primary)',
sentence: {
segments: ['Prompt', { subBlockId: 'model', noun: 'a model' }],
values: { model: 'claude-sonnet-5' },
values: { model: 'claude-sonnet-5-5' },
},
rows: [],
x: 300,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -98,7 +98,7 @@ export const DEMO_BLOCKS: BlockDef[] = [
bgColor: 'var(--text-primary)',
sentence: {
segments: ['Prompt', { subBlockId: 'model', noun: 'a model' }],
values: { model: 'claude-sonnet-5' },
values: { model: 'claude-sonnet-5-5' },
},
rows: [],
x: col(2),
Expand Down Expand Up @@ -127,7 +127,7 @@ export const DEMO_BLOCKS: BlockDef[] = [
bgColor: 'var(--text-primary)',
sentence: {
segments: ['Prompt', { subBlockId: 'model', noun: 'a model' }],
values: { model: 'claude-sonnet-5' },
values: { model: 'claude-sonnet-5-5' },
},
rows: [],
x: col(4),
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ import { useSubBlockStore } from '@/stores/workflows/subblock/store'
/**
* Constants for ComboBox component behavior
*/
const DEFAULT_MODEL = 'claude-sonnet-5'
const DEFAULT_MODEL = 'claude-sonnet-5-5'
const ZOOM_FACTOR_BASE = 0.96
const MIN_ZOOM = 0.1
const MAX_ZOOM = 1
Expand Down Expand Up @@ -296,7 +296,7 @@ export const ComboBox = memo(function ComboBox({

/**
* Determines the default option value to use.
* Priority: explicit defaultValue > claude-sonnet-5 for model field > first option
* Priority: explicit defaultValue > DEFAULT_MODEL for model field > first option
*/
const defaultOptionValue = useMemo(() => {
if (defaultValue !== undefined) {
Expand All @@ -308,7 +308,6 @@ export const ComboBox = memo(function ComboBox({
// Default not available (e.g. provider disabled) — fall through to other fallbacks
}

// For model field, default to claude-sonnet-5 if available
if (subBlockId === 'model') {
const defaultModelOption = evaluatedOptions.find(
(opt) => getOptionValue(opt) === DEFAULT_MODEL
Expand Down
4 changes: 2 additions & 2 deletions apps/sim/blocks/blocks/agent.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ import type { ToolResponse } from '@/tools/types'
const logger = createLogger('AgentBlock')

/** Model the agent block falls back to when `model` is unset or the auto pseudo-model. */
const AGENT_FALLBACK_MODEL = 'claude-sonnet-5'
const AGENT_FALLBACK_MODEL = 'claude-sonnet-5-5'

const MODELS_WITH_REASONING_EFFORT = getModelsWithReasoningEffort()
const MODELS_WITH_VERBOSITY = getModelsWithVerbosity()
Expand Down Expand Up @@ -154,7 +154,7 @@ Return ONLY the JSON array.`,
type: 'combobox',
placeholder: 'Type or select a model...',
required: true,
defaultValue: 'claude-sonnet-5',
defaultValue: 'claude-sonnet-5-5',
options: getAgentModelOptions,
commandSearchable: true,
},
Expand Down
2 changes: 1 addition & 1 deletion apps/sim/blocks/blocks/evaluator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -186,7 +186,7 @@ export const EvaluatorBlock: BlockConfig<EvaluatorResponse> = {
type: 'combobox',
placeholder: 'Type or select a model...',
required: true,
defaultValue: 'claude-sonnet-5',
defaultValue: 'claude-sonnet-5-5',
options: getModelOptions,
},
...getProviderCredentialSubBlocks(),
Expand Down
4 changes: 2 additions & 2 deletions apps/sim/blocks/blocks/router.ts
Original file line number Diff line number Diff line change
Expand Up @@ -183,7 +183,7 @@ export const RouterBlock: BlockConfig<RouterResponse> = {
type: 'combobox',
placeholder: 'Type or select a model...',
required: true,
defaultValue: 'claude-sonnet-5',
defaultValue: 'claude-sonnet-5-5',
options: getModelOptions,
},
...getProviderCredentialSubBlocks(),
Expand Down Expand Up @@ -302,7 +302,7 @@ export const RouterV2Block: BlockConfig<RouterV2Response> = {
type: 'combobox',
placeholder: 'Type or select a model...',
required: true,
defaultValue: 'claude-sonnet-5',
defaultValue: 'claude-sonnet-5-5',
options: getModelOptions,
},
...getProviderCredentialSubBlocks(),
Expand Down
6 changes: 3 additions & 3 deletions apps/sim/executor/constants.ts
Original file line number Diff line number Diff line change
Expand Up @@ -228,7 +228,7 @@ export const HTTP = {
} as const

export const AGENT = {
DEFAULT_MODEL: 'claude-sonnet-5',
DEFAULT_MODEL: 'claude-sonnet-5-5',
get DEFAULT_FUNCTION_TIMEOUT() {
return getMaxExecutionTimeout()
},
Expand All @@ -243,13 +243,13 @@ export const MCP = {
} as const

export const ROUTER = {
DEFAULT_MODEL: 'claude-sonnet-5',
DEFAULT_MODEL: 'claude-sonnet-5-5',
DEFAULT_TEMPERATURE: 0,
INFERENCE_TEMPERATURE: 0.1,
} as const

export const EVALUATOR = {
DEFAULT_MODEL: 'claude-sonnet-5',
DEFAULT_MODEL: 'claude-sonnet-5-5',
DEFAULT_TEMPERATURE: 0.1,
RESPONSE_SCHEMA_NAME: 'evaluation_response',
JSON_INDENT: 2,
Expand Down
84 changes: 83 additions & 1 deletion apps/sim/providers/anthropic/core.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -293,7 +293,7 @@ describe('executeAnthropicProviderRequest forced tool use', () => {
expect(payload.tool_choice).toEqual({ type: 'tool', name: 'publish' })
})

it.each(['claude-fable-5-1', 'claude-opus-5-5'])(
it.each(['claude-fable-5-1', 'claude-opus-5-5', 'claude-sonnet-5-5'])(
'drops forced tool_choice when the catalog model disables Force (%s)',
async (model) => {
const { payload, warn } = await runWithForcedTool(model)
Expand Down Expand Up @@ -745,6 +745,87 @@ describe('streaming', () => {
}
})
})

/**
* Both tool loops rebuild the request for every turn after a tool call, so
* the `none` mapping must hold past the first request. Claude Sonnet 5.5
* rejects `thinking.type: "disabled"` and names `between_tools` (which takes
* no effort or display field) as its lowest setting; a later turn without it
* would silently run adaptive thinking. Every other model keeps `none` as
* "send no thinking config".
*/
describe('executeAnthropicProviderRequest none thinking level across tool turns', () => {
const lookupTool = {
id: 'lookup',
name: 'lookup',
description: 'Lookup',
params: {},
parameters: { type: 'object', properties: {}, required: [] },
}
const turns = [
message([{ type: 'tool_use', id: 'tool-1', name: 'lookup', input: {} }], 'tool_use'),
message([{ type: 'text', text: 'done' }], 'end_turn'),
]

/** Runs a tool exchange and returns every request body sent over the SDK boundary. */
async function runToolExchange(model: string, streaming: boolean) {
mockExecuteTool.mockResolvedValue({ success: true, output: { value: 'tool result' } })
const sent: Anthropic.Messages.MessageCreateParams[] = []
const nextTurn = (payload: Anthropic.Messages.MessageCreateParams) => {
sent.push(payload)
return turns[sent.length - 1]
}
const result = await executeAnthropicProviderRequest(
{
model,
apiKey: 'test-key',
stream: streaming,
maxTokens: 1024,
thinkingLevel: 'none',
agentEvents: true,
messages: [{ role: 'user', content: 'Look this up' }],
tools: [lookupTool],
},
{
providerId: 'anthropic',
providerLabel: 'Anthropic',
createClient: () =>
({
messages: streaming
? {
stream: (payload: never) =>
stream([{ type: 'message_stop' }], nextTurn(payload)),
}
: { create: async (payload: never) => nextTurn(payload) },
}) as never,
logger: { info: vi.fn(), warn: vi.fn(), error: vi.fn(), debug: vi.fn() },
}
)
if (streaming) await collectEvents(result as StreamingExecution)
return sent
}

it.each([false, true])(
'sends bare between_tools on every turn (streaming: %s)',
async (streaming) => {
const sent = await runToolExchange('claude-sonnet-5-5', streaming)
expect(sent).toHaveLength(2)
for (const payload of sent) {
expect(payload.thinking).toEqual({ type: 'between_tools' })
expect(payload.output_config).toBeUndefined()
}
}
)

it.each(['claude-sonnet-5', 'claude-opus-5-5'])(
'sends no thinking config on %s',
async (model) => {
const sent = await runToolExchange(model, false)
expect(sent).toHaveLength(2)
for (const payload of sent) expect(payload.thinking).toBeUndefined()
}
)
})
})

/**
Expand All @@ -758,6 +839,7 @@ describe('buildThinkingConfig', () => {
for (const model of [
'claude-fable-5-1',
'claude-fable-5',
'claude-sonnet-5-5',
'claude-sonnet-5',
'claude-opus-5-5',
'claude-opus-5',
Expand Down
41 changes: 31 additions & 10 deletions apps/sim/providers/anthropic/core.ts
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,7 @@ const ANTHROPIC_THINKING_OUTPUT_HEADROOM = 4096
/**
* Checks if a model supports adaptive thinking (thinking.type: "adaptive").
* Fable 5, Fable 5.1, and Opus 5.5 support ONLY adaptive thinking (always on; type: "disabled" is rejected).
* Sonnet 5.5 is adaptive by default and rejects type: "disabled"; its lowest setting is "between_tools".
* Sonnet 5 supports ONLY adaptive thinking (manual budget_tokens returns a 400 error).
* Opus 5, Opus 4.8, and Opus 4.7 support ONLY adaptive thinking (no extended thinking / budget_tokens).
* Opus 4.6 and Sonnet 4.6 support both extended and adaptive thinking — use adaptive.
Expand All @@ -169,10 +170,15 @@ function supportsAdaptiveThinking(modelId: string): boolean {
/**
* Builds the thinking configuration for the Anthropic API based on model capabilities and level.
*
* - Fable 5.1, Fable 5, Sonnet 5, Opus 5.5, Opus 5, Opus 4.8, Opus 4.7: Uses adaptive thinking only (no extended thinking support)
* - Fable 5.1, Fable 5, Sonnet 5.5, Sonnet 5, Opus 5.5, Opus 5, Opus 4.8, Opus 4.7: Uses adaptive thinking only (no extended thinking support)
* - Opus 4.6, Sonnet 4.6: Uses adaptive thinking with effort parameter
* - Other models: Uses budget_tokens-based extended thinking
*
* The `none` level returns null (send no thinking config) unless the model
* declares `capabilities.thinking.noneMode`: Sonnet 5.5 rejects
* `type: "disabled"`, so `none` becomes `type: "between_tools"`, which turns
* off up-front thinking and takes no effort or display field.
*
* The newest Claude generations default `thinking.display` to `omitted`
* (empty thinking blocks, no thinking deltas). Their registry entries mark
* `capabilities.thinking.streamed: 'summary'`, and for those models Sim opts
Expand All @@ -190,7 +196,19 @@ export function buildThinkingConfig(
outputConfig?: Anthropic.Messages.OutputConfig
} | null {
const capability = getThinkingCapability(modelId)
if (!capability || !capability.levels.includes(thinkingLevel)) {
if (!capability) {
return null
}

if (thinkingLevel === 'none') {
if (capability.noneMode !== 'between_tools') return null
return {
// double-cast-allowed: @anthropic-ai/sdk 0.115 predates the between_tools thinking type (typed from 0.129)
thinking: { type: 'between_tools' } as unknown as Anthropic.Messages.ThinkingConfigParam,
}
}

if (!capability.levels.includes(thinkingLevel)) {
return null
}

Expand Down Expand Up @@ -358,9 +376,8 @@ export async function executeAnthropicProviderRequest(
}
}

// Add extended thinking configuration if supported and requested
// The 'none' sentinel means "disable thinking" — skip configuration entirely.
if (request.thinkingLevel && request.thinkingLevel !== 'none') {
// The 'none' sentinel means "disable thinking": no config, unless the model declares a noneMode.
if (request.thinkingLevel) {
const thinkingConfig = buildThinkingConfig(
request.model,
request.thinkingLevel,
Expand Down Expand Up @@ -405,11 +422,15 @@ export async function executeAnthropicProviderRequest(
// Per Anthropic docs: thinking is not compatible with temperature or top_k modifications.
payload.temperature = undefined

const isAdaptive = thinkingConfig.thinking.type === 'adaptive'
logger.info(
`Using ${isAdaptive ? 'adaptive' : 'extended'} thinking for model: ${modelId} with ${isAdaptive ? `effort: ${request.thinkingLevel}` : `budget: ${(thinkingConfig.thinking as { budget_tokens: number }).budget_tokens}`}`
)
} else {
if (request.thinkingLevel === 'none') {
logger.info(`Using between_tools thinking for model: ${modelId}`)
} else {
const isAdaptive = thinkingConfig.thinking.type === 'adaptive'
logger.info(
`Using ${isAdaptive ? 'adaptive' : 'extended'} thinking for model: ${modelId} with ${isAdaptive ? `effort: ${request.thinkingLevel}` : `budget: ${(thinkingConfig.thinking as { budget_tokens: number }).budget_tokens}`}`
)
}
} else if (request.thinkingLevel !== 'none') {
logger.warn(
`Thinking level "${describeModelLevel(request.thinkingLevel)}" not supported for model: ${modelId}, ignoring`
)
Expand Down
Loading
Loading