UN-4223 [FIX] Make streamed Gemini calls honour the adapter timeout - #2313
Open
johnyrahul wants to merge 2 commits into
Open
johnyrahul wants to merge 2 commits into
johnyrahul wants to merge 2 commits into
Conversation
LiteLLM's Gemini handler (gemini/* and vertex_ai/*gemini*) drops the per-call timeout on the sync streaming path and streams on litellm.module_level_client, whose deadline is litellm.request_timeout (6000 s by default). The adapter's timeout (600 s) never reached the request, so a stalled Gemini stream could sit for 100 minutes per attempt. Pass a shared HTTPHandler built with the adapter's timeout for Gemini models on both streaming paths (complete and stream_complete). Other providers already forward timeout to the request and are untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
|
Cap the per-timeout client cache at 8 entries so a worker that sees many distinct adapter timeouts does not hold a connection pool for each. Replace the mock-only Vertex test with request-level tests: the HTTP request's read timeout is now asserted for the Vertex adapter (auth stubbed) and for stream_complete(), as well as the Gemini adapter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Contributor
Unstract test resultsPer-group results
Critical paths
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



What
gemini/*via the Gemini adapter,vertex_ai/*gemini*via the Vertex AI adapter) now apply the adapter'stimeoutto the HTTP request. Before this change the request used LiteLLM's 6000 s default.LLM.complete()(which streams by default) andLLM.stream_complete().Why
UN-4223: on 2026-10-01, 10-02 and 10-03, single
vertex_ai/gemini-3.1-flash-litecalls inunstract-worker-pg-executorhung for 2+ hours, and pod liveness eventually took whole executor pods offline (P0 on the US Checkly check). The SDK setstimeout=600on these calls, but it never took effect.The root cause is in LiteLLM 1.96.2 (
vertex_and_google_ai_studio_gemini.py):make_sync_callis bound withouttimeoutand posts onlitellm.module_level_client.litellm.request_timeout, which defaults to 6000 s.timeout. Streaming has been on by default since we added it, so in practice every Gemini call was capped at 100 minutes per read instead of 10.Anthropic's handler forwards
timeoutper request, so it is unaffected. I checked this in the same LiteLLM version.This is not the whole fix for UN-4223. Prod logs show the hung Gemini calls never timed out even at 6000 s. No
Retry, timeout or error line was logged over 2+ hours, while memory climbed to around 9–10 GB. That points to a stream that keeps sending data, which a per-read timeout cannot catch. A total wall-clock deadline on streamed completions follows in a separate PR. This PR closes a real gap that was hiding behind it: a genuinely silent Gemini stream should fail in 10 minutes, not 100.How
_with_gemini_stream_timeout()addsclient=HTTPHandler(timeout=httpx.Timeout(<adapter timeout>))to the streaming call's kwargs when:gemini/prefix, orvertex_ai/withgeminiin the name), andtimeoutis set, andclient.clientas anHTTPHandler(gemini_client=client if isinstance(client, HTTPHandler)), so passing a client is the only way to get the timeout through.module_level_client.Can this PR break any existing features. If yes, please list possible items. If no, please explain why. (PS: Admins do not merge the PR without this section filled)
timeout(default 600 s; the Vertex form exposes no timeout field) now fails and retries instead of waiting up to 6000 s. A healthy stream sends chunks continuously and the read timeout resets on each one, so long generations are unaffected.vertex_ai/claude-*) take LiteLLM's partner route, which already forwardstimeout. They don't match the Gemini check and are pinned by a test.Database Migrations
Env Config
Relevant Docs
Related Issues or PRs
Dependencies Versions
Notes on Testing
tests/test_gemini_stream_timeout.pywith 11 cases.6000.0, reproducing the prod bug, and the Vertex adapter test fails on the missing client.ruff0.3.4 (the pre-commit pin) is clean.Screenshots
Checklist
I have read and understood the Contribution Guidelines.
🤖 Generated with Claude Code