You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
@@ -294,6 +295,7 @@ Point any OpenAI-style client (Cursor, Continue, Aider, OpenWebUI, Hermes, your
294
295
295
296
-**OpenAI `/v1/chat/completions`** — streaming SSE + non-streaming; tool calling (parallel tools, streamed `tool_calls` deltas); vision (`image_url` base64/data-URL); `reasoning_effort` snapped down to per-model tiers; `max_completion_tokens`; usage passthrough from upstream `totalUsage`
296
297
-**Anthropic `/v1/messages`** — full streaming block lifecycle (`message_start` → `content_block_start/delta/stop` → `signature_delta` → `message_delta` → `message_stop`); `tool_use` / `tool_result` round-trip; thinking blocks with signature compatibility; system block arrays
298
+
-**Usage detail reaches the client** — the Anthropic route reports `input_tokens` / `cache_read_input_tokens` / `cache_creation_input_tokens` in the closing `message_delta` (non-streaming: `message.usage`), and the OpenAI route reports `prompt_tokens_details.cached_tokens`. Clients can therefore see **cache hits** and the **upstream's real input count** instead of a locally estimated total. `input_tokens` / `prompt_tokens` already include cache reads, with the cache fields as a subset — do not add them together
297
299
-**Faithful wire translation**, verified line-by-line against the original CLI: raw-base64 image parts with `mediaType`, `tool_search→search_tools` aliasing, terminal-error no-retry list
298
300
-**Full tool passthrough** — no truncation to 15; the 30+ tools issued by multi-tool agent hosts are all forwarded
299
301
-**Fuzzy model resolution + per-plan filtering** — unknown models pass through as-is (upstream returns an accurate error instead of a silently substituted default); `GET /v1/models?plan=…&available=1` (fail-open)
0 commit comments