Repository navigation
Conversation
Squash merge after successful product CI and fallback review. The primary Codex action failed due hosted runner communication loss; no code findings were reported.
…1471) Legacy hedge racing did not emit attempt lifecycle events or summary snapshots to the proxy session routing trace, causing the decision-chain view to show no provider attempts for hedged streaming requests. Instrument sendStreamingWithHedge with DiscoveryRequestMetrics and routing trace events for attempt startup, settlement, retries, and winner commitment, then capture final trace summaries on completion.
…ure groups in model redirects (#1475) * feat(providers): support regex capture groups in model redirects Allow provider regex redirection rules to reference capture groups ($1, $&, $<name>) in the target model. Expand matched groups while preserving full model replacement semantics, update the tester UI and proxy redirector to resolve targets dynamically, and add regex guidance across locale messages. Fixes #1454 * fix(proxy): inject session header for OpenCode upstreams OpenCode Zen requires the x-opencode-session header on all requests for routing and prompt cache affinity, returning a 503 error when missing. Derive a stable UUID-formatted hash from the session ID to preserve cache affinity without exposing client identity values to upstream providers. Inject the header in proxy forwarder requests and provider connectivity test probes when targeting opencode.ai hosts, while preserving explicit custom headers. Fixes #1470 * fix(proxy): normalize opencode session headers and preserve shadow seed Ensure x-opencode-session is only injected after merging custom headers and performing a case-insensitive check to avoid duplicate header values. Retain upstreamSessionSeed on ProxySession so hedge shadow sessions maintain cache affinity with the parent request.
Node 24 global fetch rejects npm undici ProxyAgent/socksDispatcher (UND_ERR_INVALID_ARG invalid onRequestStart) in ~12ms without opening the proxy socket. Provider tests and webhook sends now use the same undici fetch that created the dispatcher.
* fix(logs): harden session suggestion indexes * chore: resolve session suggestion PR conflicts * fix(logs): resolve session suggestion PR conflicts
Restore complete JSON generation outputs from Claude, OpenAI Chat, Responses, and Gemini streams instead of storing raw SSE text. Name traces as user:shortModel, record mashed original client headers in client_metadata, and pass LANGFUSE_TRACING_ENVIRONMENT/RELEASE into LangfuseSpanProcessor. Create observations inside propagateAttributes for the JS SDK v5 model, snapshot and cap bodies at 1 MiB before async export, and add coverage for stream finalization and trace metadata.
* fix(proxy): 修复首内容串行排队并按可用内存分配预算 * fix(proxy): 修复额度恢复与裸 JSON 流兼容性 * 修复请求额度确定性回收并隔离容器压力与流式准入范围 * 修复子限额排队预占并完整计量 Discovery 前缀生命周期 * 保留实时观测刷新与清理任务的请求内存所有权 * 等待全部 worker 就绪后建立内存准入基线 * 修复启动和门控清理阻塞并补齐失败请求回收 --------- Co-authored-by: tesgth032 <254944484+tesgth032@users.noreply.github.com> Co-authored-by: tesgth032 <tesgth032@users.noreply.github.com>
* feat(providers): resolve custom header templates at request time
Provider customHeaders values can copy inbound request headers and session IDs via {{header.Name}}, {{session.id}}, and {{session.client_id}}. Missing sources skip that outbound header. Sensitive auth/cookie sources and protected destination names are rejected; empty static values stay empty. Shadow sessions resolve {{session.id}} from upstreamSessionSeed.
* feat(providers): optional OpenCode Go session-header adapter
Saving an official https://opencode.ai provider can prompt to persist x-opencode-session as {{session.id}}. Skip the dialog to keep the hashed auto-inject from #1475.
parseCustomHeaderExpr fell through without a return value, which failed tsgo (TS2366) and broke the build after #1480. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… shrinking limit (#1487) * fix(memory): clamp governor limit by real headroom instead of re-discounting sample() recomputed 0.6 * (remaining RAM - reserve) every second, so the process's own RSS growth (Next base heap, off-heap buffers) lowered the lease limit even with zero leases. On a 2C/4G host the limit fell from 1405 MiB at boot to 1110 MiB at idle. The 0.6 factor already lives in the boot ceiling. At runtime the limit is now min(ceiling, used + headroom) where headroom is remaining RAM minus the reserve, so it only tightens when real headroom is short or under memory pressure. Same change in the multicore coordinator. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(memory): add tagged lease ledger to governor snapshot The governor only kept a scalar used counter, so a lost release() left no trace of who held the bytes. Leases now carry an optional tag and are kept in a ledger; snapshot().leases reports count, bytes and oldest age per tag (body_read, body_decode, body_materialize, gate) in the 30s worker_memory_stats log. Body-store acquire sites are tagged; the stream-gate site is tagged with the gate backstop change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(memory): bound request memory lifetime against stuck background owners RequestMemoryLifetime freed a request's leases only after every background retainer settled. AsyncTaskManager tasks and un-timed Redis/DB promises that never settle therefore pinned the multi-MB body lease and the body buffers forever, leaving used == limit with zero connections. Once the response has ended, background owners now get a grace window (REQUEST_MEMORY_BACKGROUND_GRACE_MS, default 150s, never shorter than the hedge loser drain timeout + 30s). The window is idle-based: task touch() refreshes it, so detached streams that keep draining survive. On expiry the lifetime force-releases its leases, drops the session's body references, and logs the stuck owner labels. The timer never runs while the response is still streaming. Retainers are labelled, attaching to an ended lifetime returns false instead of throwing (stranding the lease), and worker_memory_stats gains requestMemory (draining count, oldest age, labels, forcedTotal). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(proxy): back stream-gate leases with the request lifetime Gate leases were passed hand to hand outside the request lifetime. After commit() the only owner of the prefix lease (and its spill file) was the buildBufferedPrefixStream closure, released only on pull-past-prefix or cancel, so a dropped unread stream leaked it permanently. Attach the raw and committed gate leases to the request lifetime as a backstop; explicit release stays the primary path and release is idempotent. Tag the gate lease in the ledger, and release the incoming lease before DiscoveryPrebuffer.attachLease throws on a double attach. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(proxy): record local capacity 429s in usage logs Capacity rejections happen before the message_request row exists (the body lease is taken in ProxySession.fromContext, ahead of auth), and logErrorToDatabase silently returned without a messageContext, so the dashboard never showed them. Add recordLocalCapacityRejection, following the sensitive-word/warmup blocked-request convention (providerId 0, statusCode 429, blockedBy local_capacity, not billed). Post-auth rejections without a messageContext write a row from the error handler. Pre-auth rejections resolve the key from headers on the rejection path only, and always emit a structured warn log. Rows are throttled to one per key per 10s with the suppressed count carried into the next row so retry loops cannot flood the DB. Add the local_capacity label to the log detail dialog in all 5 locales. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(memory): document headroom clamp, bounded lifetime and diagnostics Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…1481) Bumps the npm_and_yarn group with 1 update in the / directory: [next](https://github.com/vercel/next.js). Updates `next` from 16.3.2 to 16.3.3 - [Release notes](https://github.com/vercel/next.js/releases) - [Commits](vercel/next.js@v16.3.2...v16.3.3) --- updated-dependencies: - dependency-name: next dependency-version: 16.3.3 dependency-type: direct:production dependency-group: npm_and_yarn ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…1483) * fix(responses-ws): fall back to HTTP for oversized upstream requests Recognize request-size errors before accepting the first upstream event and reuse the existing same-provider HTTP fallback. Keep these failures out of the unsupported cache and discard the rejected persistent socket. Cover classification, forwarding, session reuse, and cleanup races while preserving the current handling of errors after streaming starts. * fix(responses-ws): recognize size codes without synthesizing messages Normalize underscore and hyphen separators before matching upstream size signals. Keep missing upstream messages optional so payload fallback does not introduce a hardcoded display message. Cover code-only errors in both the classifier and real WebSocket adapter tests. * fix(responses-ws): inspect top-level upstream size errors Read size signals from top-level Responses error fields as well as nested error objects. Preserve upstream messages from either form and keep the explicit status and first-event gates. Cover top-level and mixed envelopes in classifier tests and the real WebSocket adapter regression suite.
Migration 0122 only ran SELECT 1, so databases migrated outside runMigrations() never got the session-prefix indexes. The migration now creates them with IF NOT EXISTS and stamps the preflight marker, and runMigrations() builds them concurrently before migrating any database older than 0122, so the in-migration CREATE INDEX is skipped there. The concurrent index lock and statement timeouts are configurable through MIGRATION_INDEX_LOCK_TIMEOUT_MS and MIGRATION_INDEX_STATEMENT_TIMEOUT_MS. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The OpenCode Go save-time adapter wrote x-opencode-session: {{session.id}}
into provider custom headers. A configured header suppresses the hashed
header injected for opencode.ai hosts, so upstream received the raw
session ID, which can originate from the client's metadata.user_id.
Remove the adapter and its confirmation dialog. Requests to opencode.ai
keep the per-session stable hashed header from #1475, which serves the
same routing and prompt-cache purpose.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pre-auth capacity rejections recorded raw request keys in the throttle table before validating them, so a flood of forged keys filled the 4096-entry table and cleared it, dropping the throttle state of real keys and letting their retries write rows again. Only validated keys now enter the table, pre-auth key lookups are capped at 20 per second per process, and a full table evicts its oldest entry instead of clearing everything. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Regex redirect targets can now carry captured text from the client's model name. The Gemini path rewrite interpolated that value into a String.replace template, so $&, $1 and $` in the name were expanded again. Use a replacement function instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Streaming traces no longer carry responseText, so a stream with a reconstructed final output and an error message was reported with responseMissing: true. Treat a final reconstructed output as a response. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…treams
A tool without parameters streams a single input_json_delta with an empty
partial_json; JSON.parse("") failed and the whole trace output became
malformed_frame. Keep the input from content_block_start in that case.
SSE frames without an event: line are parsed with the default event name
"message", so relays that only send data: lines never reached the
Anthropic event handlers. Prefer data.type, as the OpenAI finalizer does.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Responses API usage was forwarded as-is: nested *_tokens_details objects that flat Langfuse usage_details do not map, and input/output totals that already include cached and reasoning tokens. Langfuse stores flat keys verbatim and prices each key as a separate bucket, so cached and reasoning tokens were lost or double counted. Map the usage to input, input_cached_tokens, output, output_reasoning_tokens and total, subtracting the details from their inclusive totals. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the auto worker ceiling at 32, the default DB_POOL_MAX=20 allowed up to 20 workers with one connection each, collapsing the data, control and writer lanes onto a single one-connection pool. Auto mode now also caps the worker count at floor(DB_POOL_MAX / 4). Explicit worker counts keep the one-connection minimum. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The governor test called requestCredits a second time while the first request was still pending, so it received the already resolved promise and never exercised the IPC callback error. Advance fake timers to the one-second retry and assert the same request ID is resent. The hedge lifecycle test now fails with an explicit error when the occupied gate lease was never acquired, instead of dereferencing a possibly undefined lease. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Migration 0122 applied outside runMigrations() built its four indexes with a plain CREATE INDEX, which blocks writes to message_request and usage_ledger until every build finishes. The migration now raises an error with a hint when any of the indexes is missing and either table is larger than 64 MiB, pointing to runMigrations(), which builds them with CREATE INDEX CONCURRENTLY before the migration. Small databases still get the indexes from the migration itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tatements The guard stopped the migration whenever any prefix index was missing and either table was large, so a small usage_ledger missing its indexes was blocked by a large, already indexed message_request. It now checks the size of each missing index's own table. The error detail lists the CREATE INDEX CONCURRENTLY statements for the blocked indexes, so deployments that apply migrations outside runMigrations() can build them without blocking writes and rerun the migration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(logs): preserve upstream errors in routing decision chains * fix(logs): keep hedge setup attempts and circuit state accurate
Legacy Hedge rectifier retries share one participant sequence, so legacy-hedge-2-1 and legacy-hedge-2-2 both resolved to the same decision-chain entry and showed the wrong upstream error. The retry entry also used the request attempt count as attemptNumber, which made the terminal failure of an initial-provider retry look like a duplicate and dropped it from the chain. Legacy Hedge and Discovery now write routingAttemptId and routingRound on attempt-level chain entries, the chain keeps entries whose routingAttemptId differs, and the request detail matches routed entries by attempt ID only. Reconstructed Discovery attempts use the recorded round; entries without one are grouped under a new "Round not recorded" heading in all five locales, instead of being counted in the winner's round. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(providers): show upstream balance for each provider key
供应商管理页现在默认展示每个供应商用自己的密钥查询到的上游余额。
余额来源(参考 all-api-hub 的遥测适配器做法,按域名排序后逐个尝试):
- New API / One API 家族的 /api/usage/token/
- OpenAI 兼容计费的 /v1/dashboard/billing/subscription + /usage
- DeepSeek 的 /user/balance
- Moonshot / Kimi 的 /v1/users/me/balance
- ChatGPT 账号的 /backend-api/wham/usage(参考 codex2api,
额度单位按 500000 quota = 1 USD 归一化为美元;金额保留上游原生币种,
不做汇率换算,避免展示值和上游账单对不上。
接口:
- POST /api/v1/providers/balances:batch 批量查询,默认走缓存
- POST /api/v1/providers/{id}/balance:refresh 单供应商强制刷新
界面:
- 列表视图是带刷新入口的余额方块,服务商视图是单行内联小标签
- 进入视口才登记查询(懒加载),登记的供应商合并成批次顺序发出
- 每 10 分钟自动重新读取一次,与服务端快照缓存期一致
- 查询失败只做弱化展示,不标红
其他:
- 探测请求统一 8 秒超时、响应体 256KB 上限
- 快照缓存键带上地址与密钥指纹,换密钥后旧快照立即失效
- 新增 6 个测试文件覆盖解析、规划、并发、探测、批量存储与展示逻辑
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(providers): address balance review findings
CodeRabbit 的 5 条评审意见,逐条修复:
1. resolveSourceBaseUrl 按来源解析基地址。官方钱包端点(DeepSeek、
Moonshot、ChatGPT)挂在站点根路径下,供应商被配置成
`https://api.deepseek.com/v1` 时不再拼出 `/v1/user/balance`;
网关兼容端点保留子路径挂载,只去掉末尾的 API 版本后缀,
避免 OpenAI 计费端点拼出 `/v1/v1/dashboard/...`。
2. ProviderBalanceProvider 在 effect 内创建并释放自建的 store。
StrictMode 的 setup - cleanup - setup 会复用 useMemo 的实例,
清理后 dispose 过的 store 会让所有余额停留在 idle。
3. load() 记录每个供应商的请求代号,代号已更新的旧批次不再写回,
手动刷新先完成时不会被后完成的旧请求覆盖。
4. revalidateTracked 逐批捕获失败,一批失败不再中断后续批次,
也避免定时器产生未处理的 Promise rejection;refreshAll 同样
继续处理其余批次,并把首个失败交给调用方提示。
5. 非 2xx 响应在抛出前取消响应体。fetchWithDispatcher 走 undici,
不消费响应体,不取消会一直占用连接直到超时。
测试:新增 4 个用例覆盖基地址解析、旧批次丢弃、逐批失败隔离与
刷新全部的部分失败;把断言重复路径的旧用例改成断言正确 URL。
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(providers): cover the balance provider under StrictMode
用 react-dom/client 真实跑一遍 StrictMode 的 setup - cleanup - setup,
直接验证上一提交的修复:effect 重跑后余额仍能加载,外部注入的 store
不会被卸载释放,卸载后定时器不再访问已释放的 store。
dev 模式下供应商列表本身加载不出来(HMR websocket 握手失败,与本功能
无关),改用这组测试给出确定性的证据。
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(providers): keep forced balance refreshes from being overwritten
第二轮评审的 3 条意见:
1. 定时重读会顶掉正在进行的强制刷新结果。revalidateTracked 读到的是
服务端旧快照,却拿到更新的请求代号,于是后完成的强制刷新被判为过期
而写不回去,界面会停留在旧余额直到下一个周期。现在强制刷新期间给
供应商打标记,定时重读跳过它们;强制刷新结束后恢复覆盖。
2. 组件测试里不要在异步 act 内轮询 DOM。批量加载由 setTimeout 调度,
请求可能在 renderStrict 返回后才开始,而 act 内的 commit 时机不确定,
轮询读到的可能是旧 DOM。改成在 act 内等 fetchBalances 被调用,
再在 act 外断言渲染结果。
3. 释放顺序用例只该统计卸载触发的 dispose。StrictMode 初次挂载已经跑过
一次 cleanup,直接断言 toHaveBeenCalled 在旧实现下也会通过。现在先清空
spy 调用记录,再断言卸载恰好新增一次调用。
另外核对过第 1 条用例确实能拦住回归:把生命周期改回复用已释放实例的
写法时,「effect 重跑后余额仍能加载」会失败。
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…mpty stream (#1491) (#1493) * fix(proxy): 🐛 treat Anthropic refusal as a deliverable result, not an empty stream Anthropic can end a stream with message_delta.delta.stop_reason=refusal and no content blocks. The stream content gate classified that as terminal-before-content (empty_stream), synthesized a local 502, retried/failed over and counted it toward the provider circuit breaker, so a single refusing session could open the breaker for the whole provider. - frame-classifier: message_delta with stop_reason=refusal is a content signal (shared by the enforce gate, incremental frame probe, Discovery/hedge validity, protocol observer and client-abort metering), matching the existing refusal content rules for openai-chat / openai-responses. - forwarder non-stream empty-content check and fake-streaming validator accept a content-less refusal instead of raising missing_content / no_deliverable. - StreamPrecommitError now carries the real upstream HTTP status (upstream_status_code in the local error body); the provider-error log marks gate errors as stream_gate_local with circuitBreakerAccountable. Empty streams without an explicit refusal keep the existing empty_stream policy. Closes #1491 Co-authored-by: Wine Fox <fox@ling.plus> * docs(changelog): 📝 note Anthropic refusal stream gate fix (#1491) Co-authored-by: Wine Fox <fox@ling.plus> * fix(proxy): 🐛 align circuitBreakerAccounted log with actual accounting The provider-error log reported circuitBreakerAccountable from the request-scoped gate exemption alone, so probe requests, endpoints without circuit-breaker accounting and attempts that will still retry were logged as accountable although recordFailure was never called. Compute a single recordsCircuitFailure predicate (exhausted retries, non-probe, endpoint allows accounting, not request-scoped) and use it for both the log field (circuitBreakerAccounted, now on every provider error) and the recordFailure branch so the two cannot drift. Refs #1491 Co-authored-by: Wine Fox <fox@ling.plus> --------- Co-authored-by: Wine Fox <fox@ling.plus>
* fix(providers): detect Sub2API balances and read New API account balance via access token Sub2API 供应商此前全部被判定为「不支持」:探测顺序只有 New API 的 /api/usage/token/ 与 OpenAI 兼容计费端点,两者在 Sub2API 上都是 404, 而 Sub2API 的余额端点是 GET /v1/usage(供 CC Switch 使用,密钥认证)。 Sub2API: - 通用兼容端点顺序改为 New API 令牌额度、Sub2API 用量、OpenAI 兼容计费, 与 all-api-hub 的自动遥测顺序一致 - 按 isValid 布尔字段识别 Sub2API 响应,其他网关同一路径的内容不会被误判 - 钱包模式读取账户钱包余额;quota_limited 读取密钥额度的剩余、上限与已用; 订阅模式读取各周期最小剩余额度与到期时间,-1 判定为不限量; 兼容没有 mode 字段的旧版本 New API 系统访问令牌: - 供应商新增 new_api_access_token 与 new_api_user_id 两个可选字段, 新增与编辑页面的「余额查询」卡片可填写、替换或清除令牌 - 配置令牌后只查询 /api/user/self 的账户余额(quota / used_quota), 不再回落到密钥额度 - 用户 ID 按 New-Api-User 及各分支的头名发送;v1.0.0-rc.21 等旧版本 缺少该头时返回 401,令牌无效时返回 HTTP 200 与 success:false, 两者都判定为令牌被拒绝 - 令牌只以掩码形式返回给界面与 REST 接口,审计日志中脱敏,余额缓存键 包含令牌与用户 ID 的指纹 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(providers): restore balance credentials on undo and reject masked tokens 评审意见的修复: - 单个编辑的撤销数据包含系统访问令牌与用户 ID,撤销时一并恢复; 审计日志中撤销数据的 newApiAccessToken 同样脱敏 - REST 接口拒绝把列表返回的掩码令牌(前 4 位 + •••••• + 后 4 位) 当作真实令牌保存,新增与编辑都返回 422 - 令牌输入框、显示切换按钮与用户 ID 输入框补充可访问名称与 aria-pressed - 额度耗尽的 Sub2API 响应样例改为与上游一致的 isValid:true, 另补停用密钥 isValid:false 的用例 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: ding113 <gabriel@lpupe.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(settings): add a memory admission switch, disabled by default Memory admission (introduced in v0.9.6) now follows the system setting enableMemoryAdmission, which defaults to off. When off, the governor only accounts leases and never refuses or queues them, workers do not request credits from the primary, request bodies and stream gate prefixes stay in memory without spilling to disk, the STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP sub-limit and the pre-materialization heap check are skipped, and no local 429 local_capacity_exceeded is produced. When on, every existing admission mechanism applies unchanged. The proxy handler reads cached system settings before reading the inbound body and syncs the switch into the process governor; saving the setting invalidates the cache across processes as usual. - schema: system_settings.enable_memory_admission boolean not null default false (migration 0124) - settings page: toggle next to high-concurrency mode, i18n for 5 locales - v1 API/OpenAPI, server action, legacy admin route, validation, cache defaults, column downgrade ladder Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(memory): count cgroup usage as working set and harden the admission switch The resource snapshot subtracted raw cgroup usage (memory.current, memory.usage_in_bytes, memsw usage) from the limit. That usage includes reclaimable page cache from spooled bodies, loaded code and logs, so a memory-limited container drifted to zero headroom and every request hit the local 429 after running for a while. Usage is now the working set, memory.current - inactive_file on v2 and usage - total_inactive_file on v1 (memsw included), matching kubelet/cAdvisor. In a 1 GiB cgroup with 704 MiB of written file cache the available estimate goes from 279 MiB to 983 MiB, while real anonymous usage still drives the budget to zero. Review fixes for the admission switch: - only sync the switch from the settings object held in the process cache, so a stale query finishing after invalidation or the fallback object returned on a failed read cannot flip it for the whole process - notify listeners when the switch changes and drain stream gate sub-limit waiters immediately when it turns off Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(memory): shrink the request body lease to its retained size after parsing The body_materialize lease was sized for the parse peak, 8192 + 8x bytes + structure, and held for the whole request lifetime, including long streaming responses and background owners. Measured on Claude Code style bodies the lease was 3.2x-4.4x the memory the request actually keeps (buffer plus parsed object), so a backlog of long streams booked leases far above real usage; once usedBytes passed the limit, which follows real free memory, every new request was rejected at body_intake with a local 429 (reported: usedBytes 4.06 GB, limitBytes 2.94 GB, waiting 0). The parse peak stays reserved while the body is decoded and parsed. After parsing, the lease shrinks to the retained estimate 8192 + 4x bytes + structure (raw bytes, UTF-16 worst-case strings, object structure and one outbound serialized copy) on the JSON, non-JSON and multipart paths. For a 3.69 MiB body the held lease goes from 31.9 MiB to 17.2 MiB. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (52)
Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review. 📝 WalkthroughWalkthrough新增默认关闭的内存准入系统设置。开启后,代理会按内存额度处理请求正文和流式门控前缀,并在解析完成后调整请求正文租约。资源快照会从 cgroup 用量中扣除非活动文件缓存。 Changes内存准入设置
代理内存准入
Priority: ⚪ Not assessed Estimated code review effort: 3 (Moderate) | ~25 minutes Merge Risk: ⚪ Minimal · up to The default-off setting, retained-memory accounting, and cgroup adjustments have no established blocking defect. The change is mergeable subject to normal checks. Security Architecture ReviewSecurity architecture risk: 🟠 High · up to Upgrading disables memory protection by default, allowing request bodies to consume unrestricted application memory before the proxy’s authentication checks. Administrator authorization remains intact, but database compatibility handling can also discard an attempt to enable protection. Retained concerns
Security review detailsSecurity Blast Radius
Security Findings and Attack Paths
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 2 | ❓ 3❌ Failed checks (3 inconclusive)
✅ Passed checks (2 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 41 files. (10 skipped: 9 unsupported, 1 too large.) ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
| const hotBytes = | ||
| this.options.hotBytes ?? (getMemoryGovernor().enabled ? HOT_BYTES : Number.POSITIVE_INFINITY); | ||
| if (!this.file && capacity <= hotBytes && (await this.tryGrow(this.scratchBytes + capacity))) { |
There was a problem hiding this comment.
Unbounded pre-auth body retention When memory admission is disabled, as it is by default, this branch keeps uncompressed request bodies in memory instead of spilling them to disk. The body reader checks size only for compressed input, and parsing happens before authentication. An unauthenticated client can therefore send a large body that exhausts the worker instead of receiving a bounded capacity response. Preserve an independent size limit or disk fallback. How this was verified: Uncompressed bytes reach the in-memory body store before the authentication guard, while the body-size checks apply only to compressed input.
Knowledge Base Used: Memory-aware request admission
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/lib/body-store/byte-store.ts
Line: 97-99
Comment:
**Unbounded pre-auth body retention** When memory admission is disabled, as it is by default, this branch keeps uncompressed request bodies in memory instead of spilling them to disk. The body reader checks size only for compressed input, and parsing happens before authentication. An unauthenticated client can therefore send a large body that exhausts the worker instead of receiving a bounded capacity response. Preserve an independent size limit or disk fallback. **How this was verified:** Uncompressed bytes reach the in-memory body store before the authentication guard, while the body-size checks apply only to compressed input.
**Knowledge Base Used:** [Memory-aware request admission](https://app.greptile.com/ygxz/-/custom-context/knowledge-base/ding113/claude-code-hub/-/docs/memory-aware-admission.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| () => | ||
| process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP | ||
| governor.enabled && process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP | ||
| ? resolveStreamGateGlobalPrebufferByteCap() | ||
| : Number.MAX_SAFE_INTEGER, |
There was a problem hiding this comment.
Aggregate prefix cap bypassed When memory admission is disabled, this resolver ignores a configured
STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP. Each enforcing stream still has its own prefix limit, but concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. That can exhaust the worker under concurrent streaming traffic. Keep an aggregate bound independent of the admission toggle.
Knowledge Base Used: Request sessions and streaming
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/app/v1/_lib/proxy/stream-gate/prebuffer-budget.ts
Line: 283-286
Comment:
**Aggregate prefix cap bypassed** When memory admission is disabled, this resolver ignores a configured `STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP`. Each enforcing stream still has its own prefix limit, but concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. That can exhaust the worker under concurrent streaming traffic. Keep an aggregate bound independent of the admission toggle.
**Knowledge Base Used:** [Request sessions and streaming](https://app.greptile.com/ygxz/-/custom-context/knowledge-base/ding113/claude-code-hub/-/docs/request-sessions-and-streaming.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
🧪 测试结果
总体结果: ✅ 所有测试通过 |
| ) { | ||
| // 内存准入关闭时正文全部留在内存,不落盘。 | ||
| const hotBytes = | ||
| this.options.hotBytes ?? (getMemoryGovernor().enabled ? HOT_BYTES : Number.POSITIVE_INFINITY); |
There was a problem hiding this comment.
[CRITICAL] [ERROR-SILENT] Memory admission bypass allows unbounded pre-authentication memory retention
Why this is a problem:
When memory admission is disabled (the default), line 98 sets hotBytes to Number.POSITIVE_INFINITY, which means ALL request bodies are kept in memory regardless of size. Since body parsing happens BEFORE authentication (the body is read in the proxy handler before the auth guard runs), an unauthenticated attacker can send arbitrarily large request bodies to exhaust worker memory.
The original code had a fixed HOT_BYTES limit (256 KiB) that applied unconditionally. This PR removes that protection when enabled: false.
Evidence from codebase:
- Body reading happens in proxy-handler.ts before authentication guards
- The
ByteStoreonly checks compressed input size, not uncompressed bytes - No independent size limit exists when
hotBytes = Infinity
Suggested fix:
// Line 96-99
// Preserve an independent size limit even when memory admission is disabled
const hotBytes = this.options.hotBytes ??
(getMemoryGovernor().enabled ? HOT_BYTES : 10 * 1024 * 1024); // 10 MiB max when disabledThis preserves a reasonable upper bound (10 MiB) while still being much larger than the original 256 KiB, preventing the pre-auth DoS vector while maintaining the intended behavior of keeping normal bodies in memory.
| globalState[STREAM_GATE_PREBUFFER_BUDGET_SYMBOL] ??= new StreamGatePrebufferBudget( | ||
| () => | ||
| process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP | ||
| governor.enabled && process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP |
There was a problem hiding this comment.
[HIGH] [LOGIC-BUG] Aggregate stream prefix cap bypassed when memory admission disabled
Why this is a problem:
When memory admission is disabled (the default), line 284 changes the resolver to return Number.MAX_SAFE_INTEGER regardless of the STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP environment variable. This means the configured aggregate cap (default 256 MiB) is completely ignored.
While each individual stream still has its own prefix limit, concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. Under high concurrent streaming traffic, this can exhaust worker memory.
Original behavior: STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP was always respected as a hard aggregate limit.
New behavior: When enabled: false (default), the aggregate cap is ignored entirely.
Evidence from code:
- Line 284:
governor.enabled && process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAPreturns false when disabled - The resolver then returns
Number.MAX_SAFE_INTEGER(no limit) - Each stream's individual limit still applies, but no aggregate bound exists
Suggested fix:
// Line 283-286
globalState[STREAM_GATE_PREBUFFER_BUDGET_SYMBOL] ??= new StreamGatePrebufferBudget(
() =>
process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP
? resolveStreamGateGlobalPrebufferByteCap()
: (governor.enabled ? Number.MAX_SAFE_INTEGER : 1024 * 1024 * 1024), // 1 GiB aggregate when disabled
governor
);This preserves an aggregate bound (1 GiB) independent of memory admission, preventing concurrent streams from collectively exhausting memory while still being generous for the disabled mode.
There was a problem hiding this comment.
Code Review Summary
This PR introduces a memory admission toggle (default: OFF) to address TTFB and memory management issues. While the implementation is comprehensive with excellent test coverage, two critical security/stability issues must be fixed before merge.
PR Size: XL
- Lines changed: 7197 (7112 additions, 85 deletions)
- Files changed: 52
Split Recommendation: This XL PR would benefit from splitting:
- Phase 1: Memory admission toggle infrastructure + cgroup accounting fixes (migrations, settings cache, UI)
- Phase 2: ByteStore and stream gate integration with the toggle
- Phase 3: Provider balance features (independent feature)
However, given the PR is already in review with comprehensive testing, the split is optional. The two critical issues below MUST be fixed regardless.
Issues Found
| Category | Critical | High | Medium | Low |
|---|---|---|---|---|
| Security | 1 | 0 | 0 | 0 |
| Logic/Bugs | 0 | 1 | 0 | 0 |
| Error Handling | 0 | 0 | 0 | 0 |
| Types | 0 | 0 | 0 | 0 |
| Comments/Docs | 0 | 0 | 0 | 0 |
| Tests | 0 | 0 | 0 | 0 |
| Simplification | 0 | 0 | 0 | 0 |
Critical Issues (Must Fix)
1. [CRITICAL] Unbounded pre-auth body retention (byte-store.ts:98)
Severity: Critical
Category: Security - DoS vector
Confidence: 95/100
When memory admission is disabled (the default), hotBytes is set to Number.POSITIVE_INFINITY. This allows unauthenticated attackers to send arbitrarily large request bodies that are kept entirely in memory, exhausting the worker. Body parsing occurs BEFORE authentication guards run in the proxy pipeline.
Impact: Production DoS vulnerability in default configuration.
High Priority Issues (Should Fix)
1. [HIGH] Aggregate prefix cap bypass (prebuffer-budget.ts:284)
Severity: High
Category: Logic bug - memory exhaustion
Confidence: 90/100
When memory admission is disabled, the configured STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP (default 256 MiB) is completely ignored. Concurrent streaming requests can collectively retain unlimited memory in their prefixes, potentially exhausting the worker under high concurrent load.
Impact: Worker instability under concurrent streaming traffic in default configuration.
Review Coverage
- Logic and correctness - 2 critical issues found
- Security (OWASP Top 10) - 1 pre-auth DoS vector
- Error handling - Clean
- Type safety - Clean
- Documentation accuracy - Good (comprehensive Chinese comments)
- Test coverage - Excellent (136 memory tests, 93.40% coverage)
- Code clarity - Good
Positive Observations (Not Commenting on Good Code)
The following aspects are well-executed:
- Comprehensive test coverage with 136 memory-specific tests
- Proper cgroup v1/v2 working set calculation (inactive_file exclusion)
- Event-driven cache invalidation with Redis pub/sub
- 5-language i18n support for all UI changes
- Database migrations properly versioned (0123, 0124)
- OpenAPI schema synchronization maintained
Recommendation
Status: ❌ NOT SAFE TO MERGE until critical issues are addressed.
The two issues identified create real security and stability risks in the default configuration (memory admission disabled). Both can be fixed with small, surgical changes to add bounded fallback limits instead of Infinity / MAX_SAFE_INTEGER.
Once fixed, this PR represents a significant improvement to memory management and operational visibility.
Automated review by Claude AI Code Review
Summary
This release (v0.9.7) merges the development branch into main, containing 33 commits with critical memory management improvements, provider balance features, proxy fixes, and comprehensive bug fixes across multiple areas of the codebase.
Related Issues:
Problem
Critical Performance & Memory Issues
v0.9.5 TTFB Bottleneck (#1473): The stream gate's fixed 40 MiB reservation per request meant each of the 4 default workers could only handle 1 request awaiting first content simultaneously (64 MiB budget / 40 MiB = 1). Slow upstream requests blocked all other requests in the same worker, causing:
Memory Admission Issues (#1497): v0.9.6's memory admission caused widespread 429 errors due to:
Kernel OOM (#1466): Despite v0.9.4 fixes, production still experienced kernel-level OOM (73-82 GB anon-rss on 125 GB host), with crash intervals accelerating (100 min → 31 min).
Provider & Upstream Issues
x-opencode-sessionheader on 2026-09-06, breaking all OpenCode providersstop_reason=refusal) were misclassified asempty_stream502 errors, incorrectly counting toward circuit breaker thresholds and causing 30-minute provider outagesMissing Features
Solution
Memory Management Overhaul (#1474, #1497)
Memory Admission Toggle (Default: OFF):
system_settings.enable_memory_admission(database migration 0124)Incremental Stream Gate (#1474):
429 local_capacity_exceededwithout penalizing provider healthAccurate Memory Accounting (#1497):
memory.current - inactive_filefor cgroups v2, consistent with kubelet/cAdvisor)8192 + 8 × bytes + structure8192 + 4 × bytes + structureAutomatic Budget Planning:
Provider Balance Display (#1492, #1496)
Upstream Balance Queries:
POST /api/v1/providers/balances:batch(cached by default)POST /api/v1/providers/{id}/balance:refreshSub2API Balance Recognition (#1496):
isValidboolean in/v1/usageresponseNew API System Access Token (#1496):
new_api_access_token+new_api_user_idfields (migration 0123)/api/user/selfwhen configuredUI Features:
Proxy & Upstream Fixes
OpenCode Go Integration (#1480):
x-opencode-sessionheader for OpenCode providersAnthropic Refusal Handling (#1491):
stop_reason=refusalbut no content blocks)Langfuse Enhancements (#1478):
HTTP/2 & Protocol Handling:
Database & Infrastructure
Migration Improvements:
Log Improvements:
Changes
Core Changes
system_settings.enable_memory_admission; Migration 0123 adds provider balance credentialsSupporting Changes
Testing
Automated Tests
bun run lint,bun run typecheck,bun run test,bun run test:v1bun run openapi:check,bun run openapi:lintManual Testing
Breaking Changes
None. This release maintains backward compatibility:
Deployment Notes
Configuration
PROVIDER_BALANCE_CACHE_TTL(default 600s)CCH_MEMORY_SPILL_DIRwhen memory admission enabledMigration Path
Performance Impact
Expected Improvements:
Monitoring:
Documentation
.env.examplewith memory admission configuration notesdocs/research/ttft-memory-aware-admission.mdwith detailed memory admission designdocs/multicore-gateway.mdfor 32-worker scaling formulaChecklist
Co-Authored-By: Claude Opus 4.6 noreply@anthropic.com
🤖 Generated with Claude Code
The PR is not safe to merge while default-off admission permits unbounded pre-auth body retention and bypasses configured aggregate stream-prefix capacity.
Findings
Fix with agent prompt
Summary
This release adds a default-off system setting for memory admission, wires it through administration and persistence, changes request-body and stream-prefix storage when admission is off, and revises cgroup working-set accounting. The default-off path removes the prior disk fallback for uncompressed bodies, and it bypasses configured aggregate stream-prefix capacity.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR Client[Inbound request] --> Settings[Load cached setting] Settings --> Session[Read and parse body] Session --> Auth[Authentication guards] Session --> Body{Memory admission enabled?} Body -- Yes --> Bounded[Budget and disk fallback] Body -- No --> RAM[Retain body in memory] Auth --> Upstream[Upstream stream] Upstream --> Gate[Precontent gate] Gate --> Aggregate{Memory admission enabled?} Aggregate -- Yes --> Cap[Configured aggregate cap] Aggregate -- No --> Unlimited[No aggregate cap]Reviews (1) · Last reviewed commit: "fix(memory): 增加内存准入开关(默认关闭),修复容器内存识别与正文额..."