fix(providers): retry transient failures on OpenAI Responses and Gemini - #8383
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
…, resolve shadowed payloads
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
1 issue found across 14 files
Confidence score: 4/5
- In
apps/sim/providers/gemini/core.ts, exhausted-quota 429 errors may be retried twice before surfacing, adding unnecessary delay; apply the policy’sinsufficient_quotaclassifier to this retry path.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="apps/sim/providers/gemini/core.ts">
<violation number="1" location="apps/sim/providers/gemini/core.ts:1221">
P2: This SDK retry path does not apply the policy's `insufficient_quota` exemption. A Gemini/Vertex 429 carrying an exhausted-quota error will be replayed twice before surfacing; apply the quota-body classifier to `withProviderRetry` as well.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
…retries, pin SDK retry budgets
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 17 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 17 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 17 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
Summary
providers/retry.ts: one retry policy for calls no vendor SDK retries, matching whatopenai,@anthropic-ai/sdk, and the AI SDK agree onPROVIDER_MAX_RETRIES(2) replays with jittered exponential backoffTypeError, BunECONNRESET)x-should-retry, thenretry-after-ms, thenretry-afterfetch failedcarries it oncause; Bun setsECONNRESET/ConnectionRefused/ConnectionCloseddirectly), never by message text. ATypeErrorfrom building the request is a bug and is not retriedinsufficient_quota) is not retried, and neither is a failure whose requested delay exceeds 30s (a daily quota, a provider-wide pause) — the AI SDK likewise ignores long server delaysclone(): cancelling one branch of a tee waits on the other, so an oversized body could hangx-request-id, and bound SDK error text to 200 charsfetchWithProviderRetry; the payload is prepared once outside the retry, since preparing can compact the conversation with its own model callgenerateContent/generateContentStreaminwithProviderRetry, honoring thegoogle.rpc.RetryInfodelay from the error body (the SDK exposes no headers). The SDK's ownretryOptionsstays off: it replaces every error (4xx included) with a genericstatusTextmessage and ignores the abort signal. Deep researchinteractions.*is left alone — that client already retries, and wrapping would double itmaxRetriestoPROVIDER_MAX_RETRIES(Azure OpenAI viaopenAICompatTransport()), so an SDK bump cannot silently move their retry budget. No behavior change: both equal the SDK defaults. Anthropic pins onlymaxRetries; its SDK derives the timeout frommax_tokenslib/core/errors/provider-quota.ts(now two consumers)constbound to a prepared payload, and follows thepostOnce→postrename (the old name silently dropped the OpenAI body from coverage)Accepted risk
/v1/responsestakes no idempotency key. Bun's ~300s socket wall surfaces asTimeoutError, which is not retried, so a runaway generation is never replayedFollow-ups (separate PRs)
fetchand no retry; wrap them infetchWithProviderRetryType of Change
Testing
providers/retry.test.ts: one test per failure mode (transient status, budget exhaustion, permanent 4xx, exhausted quota,x-should-retryboth ways,retry-after-ms, dropped connection, unknown error, abort during backoff, pre-aborted request, SDK status errors); each guard mutated and its test watched go redproviders/retry.tsmutated and its test watched go red (long delay on both paths, TypeError narrowing, Bun codes, uncappedretry-after, 429 tee hang, abort during body read)providers+lib/embeddingssuites (1,958 tests)bun run type-check,bun run lint,bun run check:audits(51 audits),docs-manifest:checkChecklist
test-auditauthoring gate)