Skip to content

Commit bd2fbe7

Browse files
committed
fix(usage): report cache hits and real input tokens to clients
`message_start` 只能带本地估算的 input_tokens(此刻上游还没回话),而收尾的 `message_delta` 只发了 output_tokens —— 上游 finish 事件里的缓存命中明细与真实 输入量被整个丢掉。Anthropic 兼容客户端(实测 ZCode)因此记录的 input 是估算值、 cache_read 恒为 0,用量界面上完全看不到缓存命中,而代理自己其实拿到了准确值。 - Anthropic 路由:收尾 delta 补 input_tokens / cache_read_input_tokens / cache_creation_input_tokens;非流式 buildAnthropicResponse 同步补齐 - OpenAI 路由:非流式响应与 include_usage 收尾 chunk 补 prompt_tokens_details.cached_tokens - 缓存拆分收敛到 usage.splitInput 导出复用,去掉 adapter 内的重复实现 (非流式路径此前压根没做拆分,正是该 bug 的成因) - 集成 mock 带缓存明细 + 三条路径断言,全量 229 项通过 版本 4.11.0 -> 4.11.1
1 parent 0be26d3 commit bd2fbe7

11 files changed

Lines changed: 88 additions & 18 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,21 @@
22

33
所有主要版本更新都记录在此文件。
44

5+
## [4.11.1] - 2026-09-14
6+
7+
### 修复
8+
- **客户端看不到缓存命中,输入量还是估算值** — Anthropic 路由的流式响应里,`message_start` 在拿到上游 usage 之前就已发出(`input_tokens` 只能是本地估算),而真正的收尾 `message_delta` 只带了 `output_tokens`:上游在 `finish` 事件里给出的**缓存命中明细**(`usageAcc.cacheReadTokens`)与**真实输入量**被整个丢掉。后果是 Anthropic 兼容客户端(实测 ZCode)记录的 `input_tokens` 是估算值、`cache_read_input_tokens` 恒为 0,用量界面上完全看不到缓存命中——而代理自己其实拿到了准确值(缓存读占输入 99%)。现在收尾 delta 一并补报 `input_tokens` / `cache_read_input_tokens` / `cache_creation_input_tokens`;非流式 `buildAnthropicResponse` 同样补齐。
9+
- 同源的 OpenAI 路由缺字段:非流式响应与流式 `include_usage` 收尾 chunk 都没有缓存明细,现补 `prompt_tokens_details.cached_tokens`。
10+
11+
### 兼容性
12+
- **`usage` 新增字段,纯增量,既有字段语义不变** — Anthropic 路由新增 `cache_read_input_tokens` / `cache_creation_input_tokens`(流式在收尾 `message_delta`,非流式在 `message.usage`);OpenAI 路由新增 `prompt_tokens_details.cached_tokens`。`input_tokens` / `prompt_tokens` 仍是**含缓存读的输入总量**,缓存字段是其中的子集,客户端不要重复相加。
13+
14+
### 清理(行为不变)
15+
- 缓存拆分逻辑收敛到 `usage.ts` 的 `splitInput` 并导出复用,删掉 adapter 内的重复实现(此前非流式路径压根没做这一步,正是上面那个 bug 的成因);`StreamEncoderState` 补 `cacheWriteTokens`,与既有 `cacheReadTokens` / `noCacheTokens` 对齐。
16+
17+
### 测试
18+
- 集成 mock 上游的 `__THINK__` 场景带上缓存明细,新增断言锁定:流式收尾 `message_delta` 必须报出真实 `input_tokens`(=10,而非估算)与 `cache_read_input_tokens`(=8);两条非流式路由同样锁定。全量 **229 项通过**。
19+
520
## [4.11.0] - 2026-09-12
621

722
### 修复

‎README.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -90,6 +90,7 @@
9090
- **打包与 CI** — TypeScript + esbuild + `pkg` 单文件 exe;GitHub Actions 全量回归(typecheck → build → vitest)
9191
- **离线可用的仪表盘** — Tailwind / Font Awesome / Chart.js 全部本地化,不依赖公共 CDN
9292
- **Anthropic SDK 兼容** — `POST /v1/messages/count_tokens` 本地估算(CJK 感知,不发起上游请求);OpenAI `stream_options.include_usage` 在收尾 chunk 附带 usage
93+
- **用量明细透传给客户端** — Anthropic 路由在收尾 `message_delta`(非流式为 `message.usage`)报出 `input_tokens` / `cache_read_input_tokens` / `cache_creation_input_tokens`,OpenAI 路由报出 `prompt_tokens_details.cached_tokens`。客户端因此能看到**缓存命中量**与**上游真实输入量**,而不是只有本地估算的输入总量;`input_tokens` / `prompt_tokens` 为含缓存读的总量,缓存字段是其子集(勿重复相加)
9394
- **每日预算告警 + 更新检查** — `DAILY_BUDGET_USD` 当日花费超阈值弹 toast;启动时查询 GitHub Releases,仪表盘头部显示"新版本"徽章
9495
- **明细导出** — 用量页一键导出 CSV(含 BOM,Excel 直开)
9596

@@ -294,6 +295,7 @@ Point any OpenAI-style client (Cursor, Continue, Aider, OpenWebUI, Hermes, your
294295

295296
- **OpenAI `/v1/chat/completions`** — streaming SSE + non-streaming; tool calling (parallel tools, streamed `tool_calls` deltas); vision (`image_url` base64/data-URL); `reasoning_effort` snapped down to per-model tiers; `max_completion_tokens`; usage passthrough from upstream `totalUsage`
296297
- **Anthropic `/v1/messages`** — full streaming block lifecycle (`message_start` → `content_block_start/delta/stop` → `signature_delta` → `message_delta` → `message_stop`); `tool_use` / `tool_result` round-trip; thinking blocks with signature compatibility; system block arrays
298+
- **Usage detail reaches the client** — the Anthropic route reports `input_tokens` / `cache_read_input_tokens` / `cache_creation_input_tokens` in the closing `message_delta` (non-streaming: `message.usage`), and the OpenAI route reports `prompt_tokens_details.cached_tokens`. Clients can therefore see **cache hits** and the **upstream's real input count** instead of a locally estimated total. `input_tokens` / `prompt_tokens` already include cache reads, with the cache fields as a subset — do not add them together
297299
- **Faithful wire translation**, verified line-by-line against the original CLI: raw-base64 image parts with `mediaType`, `tool_search→search_tools` aliasing, terminal-error no-retry list
298300
- **Full tool passthrough** — no truncation to 15; the 30+ tools issued by multi-tool agent hosts are all forwarded
299301
- **Fuzzy model resolution + per-plan filtering** — unknown models pass through as-is (upstream returns an accurate error instead of a silently substituted default); `GET /v1/models?plan=…&available=1` (fail-open)

‎package-lock.json‎

Lines changed: 2 additions & 2 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.

‎package.json‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
{
22
"name": "commandcode-proxy-v4",
3-
"version": "4.11.0",
3+
"version": "4.11.1",
44
"description": "Local OpenAI Chat Completions & Anthropic Messages compatible proxy for Command Code AI. Chinese dashboard, hardened translation, retries, idle timeouts.",
55
"main": "dist/index.js",
66
"type": "module",

‎src/adapters/commandcode/adapter.ts‎

Lines changed: 26 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -26,7 +26,7 @@ import {
2626
StreamEncoderState,
2727
} from '../../types/index.js';
2828
import { resolveModelName } from '../../utils/models.js';
29-
import { parseUsd } from './usage.js';
29+
import { parseUsd, splitInput } from './usage.js';
3030
import { estimateTextTokens } from './upstream.js';
3131

3232
// ─── 推理强度(reasoning effort)映射表(按 CLI wire 契约)───────────────────────
@@ -554,6 +554,7 @@ export class CommandCodeAdapter {
554554
inputTokens: 0,
555555
outputTokens: 0,
556556
cacheReadTokens: 0,
557+
cacheWriteTokens: 0,
557558
noCacheTokens: 0,
558559
};
559560
}
@@ -710,11 +711,10 @@ export class CommandCodeAdapter {
710711
if (usage.inputTokens != null) {
711712
state.inputTokens = usage.inputTokens;
712713
// 缓存命中量必须拆出来:单价仅为输入价的 1/50,混在 inputTokens 里会虚高成本。
713-
const details = usage.inputTokenDetails || {};
714-
const cacheRead = details.cacheReadTokens ?? usage.cachedInputTokens ?? 0;
715-
const cacheWrite = details.cacheWriteTokens ?? 0;
714+
const { cacheRead, cacheWrite, noCache } = splitInput(usage);
716715
state.cacheReadTokens = cacheRead;
717-
state.noCacheTokens = details.noCacheTokens ?? Math.max(0, usage.inputTokens - cacheRead - cacheWrite);
716+
state.cacheWriteTokens = cacheWrite;
717+
state.noCacheTokens = noCache;
718718
}
719719
if (usage.outputTokens != null) state.outputTokens = usage.outputTokens;
720720
}
@@ -735,7 +735,13 @@ export class CommandCodeAdapter {
735735
created: state.created,
736736
model: state.model,
737737
choices: [{ index: 0, delta: {}, finish_reason: finishReason }],
738-
usage: { prompt_tokens: pt, completion_tokens: state.outputTokens, total_tokens: pt + state.outputTokens },
738+
usage: {
739+
prompt_tokens: pt,
740+
completion_tokens: state.outputTokens,
741+
total_tokens: pt + state.outputTokens,
742+
// 缓存命中明细按 OpenAI 语义放在 prompt_tokens_details,客户端据此算缓存折扣。
743+
prompt_tokens_details: { cached_tokens: state.cacheReadTokens || 0 },
744+
},
739745
})}\n\n`);
740746
} else {
741747
chunks.push(this.openAIDelta(state, {}, finishReason));
@@ -770,6 +776,8 @@ export class CommandCodeAdapter {
770776
let reasoningText = '';
771777
const toolCalls: Array<{ type: 'tool_use'; id: string; name: string; input: Record<string, unknown> }> = [];
772778
let outputTokens = 0;
779+
let cacheReadTokens = 0;
780+
let cacheWriteTokens = 0;
773781
let stopReason: string | null = null;
774782

775783
for (const event of events) {
@@ -804,7 +812,12 @@ export class CommandCodeAdapter {
804812
// Original CLI: totalUsage at top level; rawFinishReason ?? finishReason.
805813
const usage = event.totalUsage ?? event.data?.usage;
806814
if (usage) {
807-
if (usage.inputTokens != null) inputTokens = usage.inputTokens;
815+
if (usage.inputTokens != null) {
816+
inputTokens = usage.inputTokens;
817+
const { cacheRead, cacheWrite } = splitInput(usage);
818+
cacheReadTokens = cacheRead;
819+
cacheWriteTokens = cacheWrite;
820+
}
808821
if (usage.outputTokens != null) outputTokens = usage.outputTokens;
809822
}
810823
const rawFR = event.rawFinishReason || event.finishReason || event.data?.finishReason;
@@ -848,7 +861,12 @@ export class CommandCodeAdapter {
848861
model: modelName,
849862
stop_reason: stopReason || (toolCalls.length > 0 ? 'tool_use' : 'end_turn'),
850863
stop_sequence: null,
851-
usage: { input_tokens: inputTokens, output_tokens: outputTokens },
864+
usage: {
865+
input_tokens: inputTokens,
866+
output_tokens: outputTokens,
867+
cache_read_input_tokens: cacheReadTokens,
868+
cache_creation_input_tokens: cacheWriteTokens,
869+
},
852870
};
853871
}
854872
}

‎src/adapters/commandcode/usage.ts‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -50,7 +50,7 @@ export function parseUsd(value: unknown): number | undefined {
5050
}
5151

5252
/** 从 usage 明细中拆出缓存读/写与非缓存输入量。 */
53-
function splitInput(usage: CCEventUsage): { cacheRead: number; cacheWrite: number; noCache: number } {
53+
export function splitInput(usage: CCEventUsage): { cacheRead: number; cacheWrite: number; noCache: number } {
5454
const details = usage.inputTokenDetails || {};
5555
const total = usage.inputTokens ?? 0;
5656
const cacheRead = details.cacheReadTokens ?? usage.cachedInputTokens ?? 0;

‎src/routes/chat.ts‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -286,6 +286,8 @@ export async function chatRoutes(fastify: FastifyInstance) {
286286
prompt_tokens: inputTokens,
287287
completion_tokens: outputTokens,
288288
total_tokens: inputTokens + outputTokens,
289+
// 缓存命中明细按 OpenAI 语义放在 prompt_tokens_details,客户端据此算缓存折扣。
290+
prompt_tokens_details: { cached_tokens: usageAcc.cacheReadTokens || 0 },
289291
},
290292
});
291293
} catch (err: any) {

‎src/routes/messages.ts‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -258,7 +258,16 @@ export async function messagesRoutes(fastify: FastifyInstance) {
258258
sse('message_delta', {
259259
type: 'message_delta',
260260
delta: { stop_reason: stopReason, stop_sequence: null },
261-
usage: { output_tokens: outputTokens },
261+
// 输入侧用量也只有上游收尾的 finish 事件才给得准:message_start 里
262+
// 发出去的 input_tokens 是本地估算(缓存明细此刻尚不存在)。这里在
263+
// 收尾 delta 补报真实值,口径与 message_start 一致 —— 均为含缓存的
264+
// 输入总量,cache_read_* 是其中的子集,客户端不要重复相加。
265+
usage: {
266+
input_tokens: usageAcc.inputTokens,
267+
output_tokens: outputTokens,
268+
cache_read_input_tokens: usageAcc.cacheReadTokens || 0,
269+
cache_creation_input_tokens: usageAcc.cacheWriteTokens || 0,
270+
},
262271
})
263272
);
264273
reply.raw.write(sse('message_stop', { type: 'message_stop' }));

‎src/types/index.ts‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -357,6 +357,8 @@ export interface StreamEncoderState {
357357
outputTokens: number;
358358
/** 输入中命中缓存的 token 数(计费按 cacheRead 单价)。 */
359359
cacheReadTokens: number;
360+
/** 写入缓存的输入 token 数(多数模型为 0)。 */
361+
cacheWriteTokens: number;
360362
/** 输入中未命中缓存的 token 数。 */
361363
noCacheTokens: number;
362364
/** 上游 provider-metadata 给出的权威账单金额(USD);缺省为 undefined。 */

‎tests/adapter.test.ts‎

Lines changed: 8 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -289,11 +289,17 @@ describe('Stream encoding', () => {
289289
{ type: 'reasoning-delta', text: 'thinking...' },
290290
{ type: 'text-delta', text: 'Let me check.' },
291291
{ type: 'tool-call', toolCallId: 'tu_1', toolName: 'search', input: { q: 'x' } },
292-
{ type: 'finish', finishReason: 'tool-calls', data: { usage: { inputTokens: 10, outputTokens: 20 } } },
292+
{ type: 'finish', finishReason: 'tool-calls', data: { usage: { inputTokens: 10, outputTokens: 20, inputTokenDetails: { cacheReadTokens: 7, noCacheTokens: 3 } } } },
293293
];
294294
const msg = adapter.buildAnthropicResponse(events, 'msg_x', 'claude-sonnet-5', 5);
295295
expect(msg.stop_reason).toBe('tool_use');
296-
expect(msg.usage).toEqual({ input_tokens: 10, output_tokens: 20 });
296+
// 缓存命中明细必须随 usage 透出,否则客户端无法区分"输入总量"与"其中命中缓存的部分"。
297+
expect(msg.usage).toEqual({
298+
input_tokens: 10,
299+
output_tokens: 20,
300+
cache_read_input_tokens: 7,
301+
cache_creation_input_tokens: 0,
302+
});
297303
expect(msg.content.some((b: any) => b.type === 'thinking')).toBe(true);
298304
expect(msg.content.some((b: any) => b.type === 'text')).toBe(true);
299305
expect(msg.content.some((b: any) => b.type === 'tool_use' && b.name === 'search')).toBe(true);

0 commit comments

Comments
 (0)