UN-4224 [FIX] Stop sending temperature to GPT-6 models on Bedrock, OpenAI and Azure OpenAI - #2310
Open
praveen-formido wants to merge 6 commits into
Open
praveen-formido wants to merge 6 commits into
praveen-formido wants to merge 6 commits into
Conversation
…enAI and Azure OpenAI GPT-6 models (luna, sol, astra, 6.1-sol) reject the `temperature` parameter. Our LLM adapters always pass one (the pydantic default, or 1 when reasoning is on), and LiteLLM forwards it on the Bedrock Mantle route, so every GPT-6 Luna call on Mantle failed with "temperature not permitted for this model". GPT-5.6 was never affected because LiteLLM has a GPT-5 rule that drops a non-1 temperature; that rule matches `gpt-5` names only. The Converse route (`us.`/`global.` profile ids) already drops it. - Add a `gpt-6` stem to `_SAMPLING_DEPRECATED_MODEL_STEMS`, the strip already used for Claude Opus 4.7+. The Bedrock, Azure AI Foundry, Vertex and Anthropic adapters call it on their return path. The trailing-edge anchor keeps `gpt-5.6-*` (normalised to `gpt-5-6-*`), `gpt-60` and `gpt-oss-*` out. - Call the strip from the native OpenAI adapter too. - Azure OpenAI: the deployment name need not name the model, so detect GPT-6 from the optional Model field as well. `LLM` sets `cost_model` aside and re-validates without it, where Azure's default temperature of 1 would return, so pin `temperature` to None instead of popping it (LiteLLM omits a None temperature) and stop reasoning from forcing 1. LiteLLM is intentionally not upgraded here: the first release whose bundled registry carries the GPT-6 Mantle entries (1.104.0) needs boto3>=1.43.1, which conflicts with the boto3 1.34 / s3fs 2024.x pins. See UN-4224 for the routing gap that leaves on air-gapped deployments. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
|
…e per call Review (Greptile P1): for an Azure OpenAI deployment whose name hides GPT-6, the only record of the model that survived `LLM`'s re-validation was the pinned `temperature: None`. `complete()` merges per-call kwargs over the stored ones, so a caller passing `temperature=` (the cloud agentic_table / agentic_extraction workers do) overwrote that marker and the temperature reached a model that rejects it. Detect from the real model id instead of a caller-writable marker. `LLM._revalidate` now feeds the `cost_model` it set aside back into re-validation, after the per-call kwargs so they cannot displace it, and the four completion paths use it. Azure prefers that `cost_model` as the original model id, so detection holds on every pass and the sampling params are simply popped like every other adapter -- the `None` pin and its sticky-marker check are gone. Other adapters are unaffected: none declares `cost_model`, Pydantic ignores unknown keys, and every call site already pops `cost_model` before LiteLLM sees the kwargs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Deepak-Kesavan
approved these changes
Oct 5, 2026
harini-venkataraman
approved these changes
Oct 5, 2026
…ge stack LiteLLM 1.104.0 is the first release whose bundled model registry carries the GPT-6 Bedrock Mantle entries, so air-gapped deployments (which fall back to the bundled registry) now route `openai.gpt-6-luna` to Mantle and price GPT-6 calls instead of recording $0. LiteLLM >= 1.100 hard-requires boto3 >= 1.43.1, which drags the storage stack with it; these move together: - litellm 1.96.2 -> 1.104.0 - boto3 / botocore 1.34.x -> 1.43.106, pinned exactly in backend, workers and sdk1: it must stay inside the botocore range aiobotocore 3.9.2 accepts (1.43.101-1.43.106) - s3fs / fsspec 2024.10.0 -> 2026.9.0 (s3fs drops its `[boto3]` extra) - gcsfs 2024.10.0 -> 2026.10.0, adlfs 2024.7 -> 2026.8.0 - google-cloud-storage 2.9.0 -> 3.16.0: every gcsfs compatible with fsspec 2026.x needs google-cloud-storage 3.x. Our direct use (Client, bucket, get_blob, md5_hash, upload_from_*) is unchanged in 3.x. Follow-on fixes: - workers: bound `requires-python` to <3.13 like every other project. The open bound made uv resolve Python 3.14, where google-api-core >= 2.27 needs protobuf >= 6.31 but the connectors' secret-manager / bigquery pins cap it below 5. The image runs Python 3.12. - MinioFS: drop the UN-3487 `walk()` override. fsspec 2026.x's `DirFileSystem.walk()` relpaths each entry's `name` itself, so the override relpathed twice and tripped fsspec's assertion. The UN-3487 regression test now guards the upstream behaviour. - cohere embed timeout patch: re-pointed to 1.104.0 after re-diffing upstream; the sync `embedding()` still builds an untimed HTTPHandler. Known gap, deliberately not fixed here: boto3 >= 1.36 stops sending Content-MD5 on DeleteObjects, which MinIO older than RELEASE.2025-01-20 rejects. Verified live against the enterprise chart's bitnami-minio:2024.12.18: single-file deletes (recursive=False, raw fsspec) fail with MissingContentMD5, directory cleanup survives through the UN-3421 fallback. MinIO 2026-09-22 (OSS compose) is unaffected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…than 2025-01-20 botocore 1.36+ sends a CRC32 flexible checksum instead of Content-MD5 on DeleteObjects. MinIO older than RELEASE.2025-01-20 -- including the enterprise chart's bitnami-minio:2024.12.18 -- rejects that with MissingContentMD5, and s3fs routes every `rm` through DeleteObjects, so with the boto3 1.43 upgrade single-file deletes (recursive=False, raw fsspec) and connector deletes failed on those servers. The UN-3421 fallback only covered recursive deletes through sdk1 FileStorage. Register a `before-call.s3.DeleteObjects` handler that adds Content-MD5 over the final (XML-escaped) body. It is appended to botocore's BUILTIN_HANDLERS, so every session created afterwards gets it -- boto3, botocore and the aiobotocore sessions s3fs creates. AWS S3 and new MinIO accept both headers; S3 Express (which rejects MD5) is skipped. `AWS_REQUEST_CHECKSUM_CALCULATION=when_required` does not help (the operation requires a checksum), and botocore's deprecated `conditionally_calculate_md5` skips requests that carry a flexible checksum, so neither could be reused. unstract-connectors does not depend on sdk1, so each carries an identical, idempotent copy: sdk1 loads it from the file-storage helper, connectors from the MinIO connector. Whichever loads first registers. Verified live on the new stack against bitnami-minio 2024.12.18 and MinIO 2026-09-22: FileStorage.rm(recursive=False), raw fsspec single and bulk rm, and MinioFS rm (file and recursive) all succeed with no fallback warnings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SonarCloud failed the quality gate on duplicated new code (39.2% vs 3%): connectors carried a verbatim copy of sdk1's hook and its tests, on the assumption that connectors does not depend on sdk1. It does at runtime: `unstract_file_system.py`, the base class every filesystem connector extends, imports `unstract.filesystem`, which depends on sdk1. So the MinIO connector now imports `unstract.sdk1.patches.s3_delete_objects_md5` directly, and the copy and its duplicate tests are gone. Importing the connector still registers the handler (verified), and connector deletes against MinIO 2024-12-18 still succeed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Contributor
Unstract test resultsPer-group results
Critical paths
|
Deepak-Kesavan
approved these changes
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



What
temperatureto OpenAI GPT-6 models (gpt-6-luna,gpt-6-sol,gpt-6-astra,gpt-6.1-sol) from the AWS Bedrock, native OpenAI and Azure OpenAI LLM adapters. Azure AI Foundry is covered through the existing strip.Why
GPT-6 Luna on AWS Bedrock failed every call with
temperature not permitted for this model. GPT-6 models reject the parameter.Our adapters always pass a temperature: the pydantic default (0.1, or 1 for Azure OpenAI), and 1 when reasoning / extended thinking is on.
What reaches AWS depends on the route. I captured the outgoing request on LiteLLM 1.96.2 with
drop_params=True:openai.gpt-6-luna"temperature": 0.1, the reported failureus./global.openai.gpt-6-lunaopenai.gpt-5.6-terraGPT-5.6 worked only because LiteLLM drops a non-1 temperature for
gpt-5names. That rule doesn't match GPT-6.Not a regression from UN-4020 [FIX] Route AWS Bedrock Mantle models (GPT-5.6 Terra) via bedrock_mantle #2248 (UN-4020). That PR didn't handle temperature for any Mantle model.
How
_SAMPLING_DEPRECATED_MODEL_STEMSgets agpt-6stem. This strip already handles Claude Opus 4.7+ and is called by the Bedrock, Azure AI Foundry, Vertex and Anthropic adapters..→-normalisation matches every Bedrock encoding:bedrock/,bedrock_mantle/,us./global.profiles andgpt-6.1-*.gpt-5.6-*(which normalises togpt-5-6-*),gpt-60andgpt-oss-*.cost_model.LLMsetscost_modelaside at construction. A newLLM._revalidatehelper, used by all four completion paths, passes it back into re-validation. It goes in after the per-call kwargs, so a caller passingtemperature=cannot displace it. The cloud agentic workers do passtemperature=; this was flagged by review.Can this PR break any existing features. If yes, please list possible items. If no, please explain why. (PS: Admins do not merge the PR without this section filled)
gpt-60and older Claude models to keep their temperature.LLM._revalidatepassescost_modelback toadapter.validate(). No parameter model declares it, Pydantic ignores unknown keys, and every call site already popscost_modelbefore LiteLLM, so no other adapter's request changes.Database Migrations
Env Config
Relevant Docs
Related Issues or PRs
Dependencies Versions
LiteLLM 1.104.0 is the first release whose bundled registry includes the GPT-6 Bedrock Mantle entries. With it, air-gapped deployments now route
openai.gpt-6-lunato Mantle and price GPT-6 calls instead of recording $0. LiteLLM ≥ 1.100 hard-requiresboto3>=1.43.1, so the storage stack moves with it:Follow-on changes:
requires-pythonis now bounded to<3.13, like every other project. The open bound made uv resolve Python 3.14, where protobuf ≥ 6.31 conflicts with the connectors' secret-manager/bigqueryprotobuf<5caps. The image runs 3.12.walk()override is removed. fsspec 2026.x'sDirFileSystem.walk()now relpaths entry names itself, so the override relpathed twice and tripped fsspec's assertion.azure-datalake-store(ADLS Gen1) drops out with adlfs 2026; we only useAzureBlobFileSystem.MinIO older than RELEASE.2025-01-20. boto3 ≥ 1.36 sends a CRC32 checksum instead of
Content-MD5onDeleteObjects, and older MinIO rejects that withMissingContentMD5. That includes the enterprise chart'sbitnami-minio:2024.12.18.s3fsroutes everyrmthroughDeleteObjects.before-call.s3.DeleteObjectshandler re-addsContent-MD5over the final (XML-escaped) body. It is appended to botocore'sBUILTIN_HANDLERS, so boto3, botocore and the aiobotocore sessions thats3fscreates all get it. S3 Express, which rejects MD5, is skipped.unstract/sdk1/patches/s3_delete_objects_md5.py. It's loaded by the sdk1 file-storage helper and imported by the MinIO connector. Connectors already depends on sdk1 at runtime, viaunstract.filesysteminunstract_file_system.py.AWS_REQUEST_CHECKSUM_CALCULATION=when_requireddoesn't help, becauseDeleteObjectsrequires a checksum. botocore's deprecatedconditionally_calculate_md5skips requests that carry a flexible checksum.Notes on Testing
bitnami-minio:2024.12.18and MinIO 2026-09-22, with the hook:FileStorage.rm(recursive=False), raw fsspec single and bulkrm, andMinioFSrm (file and recursive) all succeed with no fallback warnings. Without the hook, the single-file and raw-fsspec deletes fail on 2024.12.18.MinioFSls/walk names stay bucket-relative on both.DeleteObjectsrequest: the digest matches the XML-escaped body, the handler runs afterescape_xml_payload, and registration is idempotent. With the registration removed, 3 tests fail.bitnami-minio:2024.12.18:MissingContentMD5.Bulk delete failed with MissingContentMD5in the logs; it only completed because of the UN-3421 fallback. New tests intests/test_sampling_strip.pycover:LLM._revalidate, including a per-calltemperature=0.5, a deployment that names the model, and non-GPT-6 controls (which keep a per-call temperature);openai.gpt-6-lunaon Mantle: 200, completion returned (25/32 tokens). This is the route that failed before.global.openai.gpt-6-lunaandus.openai.gpt-6-lunaon Converse: 200.openai.gpt-6-lunaon Mantle returns "model does not exist" inap-south-1. That's regional availability, not this change.ruff0.3.4 check and format are clean.Screenshots
Checklist
I have read and understood the Contribution Guidelines.
🤖 Generated with Claude Code