Skip to content

release v0.9.7 - #1498

Merged
ding113 merged 34 commits into
mainfrom
dev
Sep 30, 2026
Merged

ding113 merged 34 commits into
mainfrom
dev

Conversation

@ding113

@ding113 ding113 commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Summary

This release (v0.9.7) merges the development branch into main, containing 33 commits with critical memory management improvements, provider balance features, proxy fixes, and comprehensive bug fixes across multiple areas of the codebase.

Related Issues:

Problem

Critical Performance & Memory Issues

v0.9.5 TTFB Bottleneck (#1473): The stream gate's fixed 40 MiB reservation per request meant each of the 4 default workers could only handle 1 request awaiting first content simultaneously (64 MiB budget / 40 MiB = 1). Slow upstream requests blocked all other requests in the same worker, causing:

  • TTFT P95 exceeded 110 seconds in production
  • Cross-request, cross-provider head-of-line blocking
  • 87-88 requests queued across workers during observation

Memory Admission Issues (#1497): v0.9.6's memory admission caused widespread 429 errors due to:

  • Container page cache incorrectly counted as used memory (available capacity compressed to 0)
  • Request body quotas held at parsing peak (3-4x actual resident memory) throughout entire request lifecycle, causing quota exhaustion during long streaming

Kernel OOM (#1466): Despite v0.9.4 fixes, production still experienced kernel-level OOM (73-82 GB anon-rss on 125 GB host), with crash intervals accelerating (100 min → 31 min).

Provider & Upstream Issues

Missing Features

Solution

Memory Management Overhaul (#1474, #1497)

Memory Admission Toggle (Default: OFF):

  • Added system_settings.enable_memory_admission (database migration 0124)
  • When disabled: no memory limits, no local 429s, no disk spilling (backward compatible behavior)
  • When enabled: full memory-aware admission as described below
  • UI toggle in System Settings with 5-language i18n support

Incremental Stream Gate (#1474):

  • Start with 128 KiB working set, grow incrementally based on actual usage
  • Disk-backed prefix buffering when memory unavailable (real disk, not tmpfs)
  • Two-tier admission with max 20s wait, returns 429 local_capacity_exceeded without penalizing provider health
  • Replay buffered prefix in original byte order after gate commits

Accurate Memory Accounting (#1497):

  • Container memory calculation: Use working set (memory.current - inactive_file for cgroups v2, consistent with kubelet/cAdvisor)
    • Before: 704 MiB cache → 279 MiB available → 14 MiB budget
    • After: 704 MiB cache → 983 MiB available → 436 MiB budget
  • Request body quota: Shrink to holding amount after parsing completes
    • Peak during parsing: 8192 + 8 × bytes + structure
    • Holding during response: 8192 + 4 × bytes + structure
    • Measured 3.2-4.4x reduction in quota usage

Automatic Budget Planning:

  • Calculate from available RAM + 0.5 × swap
  • Respect cgroup v1/v2 limits, memory.high, swap disabled, memsw joint limits
  • Multi-process coordination with idempotent IPC credit allocation
  • Auto worker scaling up to 32 (from 4-8 previously)

Provider Balance Display (#1492, #1496)

Upstream Balance Queries:

  • Batch API: POST /api/v1/providers/balances:batch (cached by default)
  • Single refresh: POST /api/v1/providers/{id}/balance:refresh
  • Probe order: New API token quota → Sub2API usage → OpenAI-compatible billing → DeepSeek/Moonshot wallets → ChatGPT account quota

Sub2API Balance Recognition (#1496):

  • Identify via isValid boolean in /v1/usage response
  • Support wallet mode (balance + usage), quota_limited mode (key quota), subscription mode (period limits)

New API System Access Token (#1496):

  • Optional new_api_access_token + new_api_user_id fields (migration 0123)
  • Query account balance via /api/user/self when configured
  • Token masking in UI/REST/audit logs
  • Cache keys include token fingerprint for instant invalidation on change

UI Features:

  • Lazy loading: only query providers near viewport
  • Batch loading: merge queries in same window (max 20 per batch)
  • Auto-refresh: every 10 minutes (matches server cache TTL)
  • Manual refresh: click any balance or toolbar "Refresh Balances"
  • Visual indicators: yellow for low balance (<10% or <$1), red for expired keys

Proxy & Upstream Fixes

OpenCode Go Integration (#1480):

  • Auto-attach x-opencode-session header for OpenCode providers
  • Support regex capture templates in custom headers
  • Send masked session IDs to OpenCode (not raw session IDs)

Anthropic Refusal Handling (#1491):

  • Recognize valid refusal responses (HTTP 200 with stop_reason=refusal but no content blocks)
  • Preserve refusal semantics instead of converting to 502
  • Exclude request-level refusals from circuit breaker counting
  • Prevent single session's refusals from blocking other sessions

Langfuse Enhancements (#1478):

  • Reconstruct streamed outputs for trace enrichment
  • Preserve streamed output and redact original headers
  • Fix parameterless tools and data-only Anthropic streams finalization
  • Send Responses usage as exclusive Langfuse buckets
  • Don't flag reconstructed streams as missing responses

HTTP/2 & Protocol Handling:

Database & Infrastructure

Migration Improvements:

  • Refuse blocking prefix index builds on large tables
  • Create session prefix indexes outside runMigrations
  • Scope 0122 guard per table and list concurrent steps
  • Keep 4 database connections per auto-mode worker

Log Improvements:

Changes

Core Changes

  • Database Schema: Migration 0124 adds system_settings.enable_memory_admission; Migration 0123 adds provider balance credentials
  • Memory Management: Complete rewrite with admission control, disk spilling, incremental gate, cgroup-aware budgeting
  • Provider Balance: Full probing infrastructure for New API, Sub2API, OpenAI-compatible, DeepSeek, Moonshot, ChatGPT
  • Proxy Pipeline: Fixed stream gate classification, HTTP/2 handling, refusal semantics, header templates
  • i18n: All new features have complete translations (zh-CN, zh-TW, en, ja, ru)

Supporting Changes

  • Enhanced error handling and routing trace preservation
  • Improved multi-core worker coordination and startup synchronization
  • Extended OpenAPI schemas for new system settings and provider fields
  • Comprehensive test coverage (136 memory-specific tests at 93.40% line coverage)

Testing

Automated Tests

  • ✅ All quality checks pass: bun run lint, bun run typecheck, bun run test, bun run test:v1
  • ✅ OpenAPI validation: bun run openapi:check, bun run openapi:lint
  • ✅ Memory-aware tests: 136 tests passing, 93.40% line coverage
  • ✅ Full test suite: 892 files, 8905 tests passing, 25 skipped, 0 failures
  • ✅ Integration tests: memory admission toggle, balance probing, disk spilling, IPC coordination

Manual Testing

  • Verified memory admission toggle in production build with real PostgreSQL + Redis
  • Tested balance queries against real Sub2API 0.2.8, New API v1.0.0-rc.40, and rc.21
  • Deployed to self-hosted test instance with 1 GiB isolated container
  • Validated 4 concurrent real streaming requests (TTFT 1.50-2.13s)
  • Confirmed 429 local capacity responses with correct Retry-After headers

Breaking Changes

None. This release maintains backward compatibility:

  • Memory admission defaults to OFF (pre-v0.9.6 behavior)
  • Provider balance features are optional additions
  • All proxy fixes preserve existing semantics for valid requests

Deployment Notes

Configuration

  • Memory Admission: Disabled by default, enable via System Settings UI or database
  • Provider Balance: Optional fields, existing providers continue working without changes
  • Environment Variables: New PROVIDER_BALANCE_CACHE_TTL (default 600s)
  • Disk Spilling: Requires writable CCH_MEMORY_SPILL_DIR when memory admission enabled

Migration Path

  1. Deploy v0.9.7 (auto-migration applies 0123, 0124)
  2. Monitor system with memory admission OFF (default)
  3. Optionally enable memory admission in System Settings after observing baseline behavior
  4. Optionally configure provider balance tokens for enhanced visibility

Performance Impact

Expected Improvements:

  • TTFT P95 should drop significantly from 110s (incremental gate eliminates serialization)
  • Memory usage more accurate (page cache exclusion, quota shrinkage)
  • Fewer false 429s from memory admission when properly tuned

Monitoring:

  • Watch for local 429s if enabling memory admission (indicates capacity tuning needed)
  • Balance queries cached 10 minutes (minimal overhead)
  • Disk spilling only occurs when memory unavailable

Documentation

  • Updated .env.example with memory admission configuration notes
  • Added docs/research/ttft-memory-aware-admission.md with detailed memory admission design
  • Updated docs/multicore-gateway.md for 32-worker scaling formula
  • All UI strings have complete 5-language translations

Checklist

  • Code follows project conventions
  • Self-review completed
  • Tests pass locally (892 files, 8905 tests)
  • Database migrations tested (auto-migrate verified)
  • Documentation updated
  • i18n complete (5 languages)
  • OpenAPI schemas synchronized
  • Breaking changes assessed (none)

Co-Authored-By: Claude Opus 4.6 noreply@anthropic.com

🤖 Generated with Claude Code

RetriggerConfidence Score: 0/5

The PR is not safe to merge while default-off admission permits unbounded pre-auth body retention and bypasses configured aggregate stream-prefix capacity.

Findings

  1. P0 Security Unbounded pre-auth body retention ▶
  2. P1 Aggregate prefix cap bypassed ▶
Fix with agent prompt
### Issue 1
src/lib/body-store/byte-store.ts:97-99
When memory admission is disabled, as it is by default, this branch keeps uncompressed request bodies in memory instead of spilling them to disk. The body reader checks size only for compressed input, and parsing happens before authentication. An unauthenticated client can therefore send a large body that exhausts the worker instead of receiving a bounded capacity response. Preserve an independent size limit or disk fallback. **How this was verified:** Uncompressed bytes reach the in-memory body store before the authentication guard, while the body-size checks apply only to compressed input.

### Issue 2
src/app/v1/_lib/proxy/stream-gate/prebuffer-budget.ts:283-286
When memory admission is disabled, this resolver ignores a configured `STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP`. Each enforcing stream still has its own prefix limit, but concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. That can exhaust the worker under concurrent streaming traffic. Keep an aggregate bound independent of the admission toggle.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
Summary

This release adds a default-off system setting for memory admission, wires it through administration and persistence, changes request-body and stream-prefix storage when admission is off, and revises cgroup working-set accounting. The default-off path removes the prior disk fallback for uncompressed bodies, and it bypasses configured aggregate stream-prefix capacity.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Client[Inbound request] --> Settings[Load cached setting]
  Settings --> Session[Read and parse body]
  Session --> Auth[Authentication guards]
  Session --> Body{Memory admission enabled?}
  Body -- Yes --> Bounded[Budget and disk fallback]
  Body -- No --> RAM[Retain body in memory]
  Auth --> Upstream[Upstream stream]
  Upstream --> Gate[Precontent gate]
  Gate --> Aggregate{Memory admission enabled?}
  Aggregate -- Yes --> Cap[Configured aggregate cap]
  Aggregate -- No --> Unlimited[No aggregate cap]
Loading

Reviews (1) · Last reviewed commit: "fix(memory): 增加内存准入开关(默认关闭),修复容器内存识别与正文额..."

ding113 and others added 30 commits September 4, 2026 03:43
Squash merge after successful product CI and fallback review. The primary Codex action failed due hosted runner communication loss; no code findings were reported.
…1471)

Legacy hedge racing did not emit attempt lifecycle events or summary
snapshots to the proxy session routing trace, causing the decision-chain
view to show no provider attempts for hedged streaming requests.

Instrument sendStreamingWithHedge with DiscoveryRequestMetrics and
routing trace events for attempt startup, settlement, retries, and
winner commitment, then capture final trace summaries on completion.
…ure groups in model redirects (#1475)

* feat(providers): support regex capture groups in model redirects

Allow provider regex redirection rules to reference capture groups
($1, $&, $<name>) in the target model. Expand matched groups while
preserving full model replacement semantics, update the tester UI and
proxy redirector to resolve targets dynamically, and add regex
guidance across locale messages.

Fixes #1454

* fix(proxy): inject session header for OpenCode upstreams

OpenCode Zen requires the x-opencode-session header on all requests for
routing and prompt cache affinity, returning a 503 error when missing.

Derive a stable UUID-formatted hash from the session ID to preserve
cache affinity without exposing client identity values to upstream
providers. Inject the header in proxy forwarder requests and provider
connectivity test probes when targeting opencode.ai hosts, while
preserving explicit custom headers.

Fixes #1470

* fix(proxy): normalize opencode session headers and preserve shadow seed

Ensure x-opencode-session is only injected after merging custom headers
and performing a case-insensitive check to avoid duplicate header
values. Retain upstreamSessionSeed on ProxySession so hedge shadow
sessions maintain cache affinity with the parent request.
Node 24 global fetch rejects npm undici ProxyAgent/socksDispatcher
(UND_ERR_INVALID_ARG invalid onRequestStart) in ~12ms without opening
the proxy socket. Provider tests and webhook sends now use the same
undici fetch that created the dispatcher.
* fix(logs): harden session suggestion indexes

* chore: resolve session suggestion PR conflicts

* fix(logs): resolve session suggestion PR conflicts
Restore complete JSON generation outputs from Claude, OpenAI Chat,
Responses, and Gemini streams instead of storing raw SSE text.
Name traces as user:shortModel, record mashed original client headers
in client_metadata, and pass LANGFUSE_TRACING_ENVIRONMENT/RELEASE
into LangfuseSpanProcessor. Create observations inside
propagateAttributes for the JS SDK v5 model, snapshot and cap bodies
at 1 MiB before async export, and add coverage for stream finalization
and trace metadata.
* fix(proxy): 修复首内容串行排队并按可用内存分配预算

* fix(proxy): 修复额度恢复与裸 JSON 流兼容性

* 修复请求额度确定性回收并隔离容器压力与流式准入范围

* 修复子限额排队预占并完整计量 Discovery 前缀生命周期

* 保留实时观测刷新与清理任务的请求内存所有权

* 等待全部 worker 就绪后建立内存准入基线

* 修复启动和门控清理阻塞并补齐失败请求回收

---------

Co-authored-by: tesgth032 <254944484+tesgth032@users.noreply.github.com>
Co-authored-by: tesgth032 <tesgth032@users.noreply.github.com>
* feat(providers): resolve custom header templates at request time

Provider customHeaders values can copy inbound request headers and session IDs via {{header.Name}}, {{session.id}}, and {{session.client_id}}. Missing sources skip that outbound header. Sensitive auth/cookie sources and protected destination names are rejected; empty static values stay empty. Shadow sessions resolve {{session.id}} from upstreamSessionSeed.

* feat(providers): optional OpenCode Go session-header adapter

Saving an official https://opencode.ai provider can prompt to persist x-opencode-session as {{session.id}}. Skip the dialog to keep the hashed auto-inject from #1475.
parseCustomHeaderExpr fell through without a return value, which failed
tsgo (TS2366) and broke the build after #1480.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… shrinking limit (#1487)

* fix(memory): clamp governor limit by real headroom instead of re-discounting

sample() recomputed 0.6 * (remaining RAM - reserve) every second, so the
process's own RSS growth (Next base heap, off-heap buffers) lowered the
lease limit even with zero leases. On a 2C/4G host the limit fell from
1405 MiB at boot to 1110 MiB at idle.

The 0.6 factor already lives in the boot ceiling. At runtime the limit is
now min(ceiling, used + headroom) where headroom is remaining RAM minus
the reserve, so it only tightens when real headroom is short or under
memory pressure. Same change in the multicore coordinator.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(memory): add tagged lease ledger to governor snapshot

The governor only kept a scalar used counter, so a lost release() left no
trace of who held the bytes. Leases now carry an optional tag and are kept
in a ledger; snapshot().leases reports count, bytes and oldest age per tag
(body_read, body_decode, body_materialize, gate) in the 30s
worker_memory_stats log. Body-store acquire sites are tagged; the
stream-gate site is tagged with the gate backstop change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(memory): bound request memory lifetime against stuck background owners

RequestMemoryLifetime freed a request's leases only after every background
retainer settled. AsyncTaskManager tasks and un-timed Redis/DB promises
that never settle therefore pinned the multi-MB body lease and the body
buffers forever, leaving used == limit with zero connections.

Once the response has ended, background owners now get a grace window
(REQUEST_MEMORY_BACKGROUND_GRACE_MS, default 150s, never shorter than the
hedge loser drain timeout + 30s). The window is idle-based: task touch()
refreshes it, so detached streams that keep draining survive. On expiry
the lifetime force-releases its leases, drops the session's body
references, and logs the stuck owner labels. The timer never runs while
the response is still streaming.

Retainers are labelled, attaching to an ended lifetime returns false
instead of throwing (stranding the lease), and worker_memory_stats gains
requestMemory (draining count, oldest age, labels, forcedTotal).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(proxy): back stream-gate leases with the request lifetime

Gate leases were passed hand to hand outside the request lifetime. After
commit() the only owner of the prefix lease (and its spill file) was the
buildBufferedPrefixStream closure, released only on pull-past-prefix or
cancel, so a dropped unread stream leaked it permanently.

Attach the raw and committed gate leases to the request lifetime as a
backstop; explicit release stays the primary path and release is
idempotent. Tag the gate lease in the ledger, and release the incoming
lease before DiscoveryPrebuffer.attachLease throws on a double attach.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(proxy): record local capacity 429s in usage logs

Capacity rejections happen before the message_request row exists (the body
lease is taken in ProxySession.fromContext, ahead of auth), and
logErrorToDatabase silently returned without a messageContext, so the
dashboard never showed them.

Add recordLocalCapacityRejection, following the sensitive-word/warmup
blocked-request convention (providerId 0, statusCode 429, blockedBy
local_capacity, not billed). Post-auth rejections without a messageContext
write a row from the error handler. Pre-auth rejections resolve the key
from headers on the rejection path only, and always emit a structured warn
log. Rows are throttled to one per key per 10s with the suppressed count
carried into the next row so retry loops cannot flood the DB.

Add the local_capacity label to the log detail dialog in all 5 locales.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(memory): document headroom clamp, bounded lifetime and diagnostics

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…1481)

Bumps the npm_and_yarn group with 1 update in the / directory: [next](https://github.com/vercel/next.js).


Updates `next` from 16.3.2 to 16.3.3
- [Release notes](https://github.com/vercel/next.js/releases)
- [Commits](vercel/next.js@v16.3.2...v16.3.3)

---
updated-dependencies:
- dependency-name: next
  dependency-version: 16.3.3
  dependency-type: direct:production
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…1483)

* fix(responses-ws): fall back to HTTP for oversized upstream requests

Recognize request-size errors before accepting the first upstream event
and reuse the existing same-provider HTTP fallback. Keep these failures
out of the unsupported cache and discard the rejected persistent socket.

Cover classification, forwarding, session reuse, and cleanup races while
preserving the current handling of errors after streaming starts.

* fix(responses-ws): recognize size codes without synthesizing messages

Normalize underscore and hyphen separators before matching upstream size
signals. Keep missing upstream messages optional so payload fallback does not
introduce a hardcoded display message. Cover code-only errors in both the
classifier and real WebSocket adapter tests.

* fix(responses-ws): inspect top-level upstream size errors

Read size signals from top-level Responses error fields as well as nested
error objects. Preserve upstream messages from either form and keep the
explicit status and first-event gates. Cover top-level and mixed envelopes
in classifier tests and the real WebSocket adapter regression suite.
Migration 0122 only ran SELECT 1, so databases migrated outside
runMigrations() never got the session-prefix indexes. The migration now
creates them with IF NOT EXISTS and stamps the preflight marker, and
runMigrations() builds them concurrently before migrating any database
older than 0122, so the in-migration CREATE INDEX is skipped there.

The concurrent index lock and statement timeouts are configurable through
MIGRATION_INDEX_LOCK_TIMEOUT_MS and MIGRATION_INDEX_STATEMENT_TIMEOUT_MS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The OpenCode Go save-time adapter wrote x-opencode-session: {{session.id}}
into provider custom headers. A configured header suppresses the hashed
header injected for opencode.ai hosts, so upstream received the raw
session ID, which can originate from the client's metadata.user_id.

Remove the adapter and its confirmation dialog. Requests to opencode.ai
keep the per-session stable hashed header from #1475, which serves the
same routing and prompt-cache purpose.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pre-auth capacity rejections recorded raw request keys in the throttle
table before validating them, so a flood of forged keys filled the
4096-entry table and cleared it, dropping the throttle state of real
keys and letting their retries write rows again.

Only validated keys now enter the table, pre-auth key lookups are capped
at 20 per second per process, and a full table evicts its oldest entry
instead of clearing everything.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Regex redirect targets can now carry captured text from the client's
model name. The Gemini path rewrite interpolated that value into a
String.replace template, so $&, $1 and $` in the name were expanded
again. Use a replacement function instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Streaming traces no longer carry responseText, so a stream with a
reconstructed final output and an error message was reported with
responseMissing: true. Treat a final reconstructed output as a response.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…treams

A tool without parameters streams a single input_json_delta with an empty
partial_json; JSON.parse("") failed and the whole trace output became
malformed_frame. Keep the input from content_block_start in that case.

SSE frames without an event: line are parsed with the default event name
"message", so relays that only send data: lines never reached the
Anthropic event handlers. Prefer data.type, as the OpenAI finalizer does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Responses API usage was forwarded as-is: nested *_tokens_details objects
that flat Langfuse usage_details do not map, and input/output totals that
already include cached and reasoning tokens. Langfuse stores flat keys
verbatim and prices each key as a separate bucket, so cached and reasoning
tokens were lost or double counted.

Map the usage to input, input_cached_tokens, output,
output_reasoning_tokens and total, subtracting the details from their
inclusive totals.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the auto worker ceiling at 32, the default DB_POOL_MAX=20 allowed up
to 20 workers with one connection each, collapsing the data, control and
writer lanes onto a single one-connection pool. Auto mode now also caps
the worker count at floor(DB_POOL_MAX / 4). Explicit worker counts keep
the one-connection minimum.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The governor test called requestCredits a second time while the first
request was still pending, so it received the already resolved promise and
never exercised the IPC callback error. Advance fake timers to the
one-second retry and assert the same request ID is resent.

The hedge lifecycle test now fails with an explicit error when the
occupied gate lease was never acquired, instead of dereferencing a
possibly undefined lease.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Migration 0122 applied outside runMigrations() built its four indexes with
a plain CREATE INDEX, which blocks writes to message_request and
usage_ledger until every build finishes. The migration now raises an
error with a hint when any of the indexes is missing and either table is
larger than 64 MiB, pointing to runMigrations(), which builds them with
CREATE INDEX CONCURRENTLY before the migration. Small databases still get
the indexes from the migration itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tatements

The guard stopped the migration whenever any prefix index was missing and
either table was large, so a small usage_ledger missing its indexes was
blocked by a large, already indexed message_request. It now checks the
size of each missing index's own table.

The error detail lists the CREATE INDEX CONCURRENTLY statements for the
blocked indexes, so deployments that apply migrations outside
runMigrations() can build them without blocking writes and rerun the
migration.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(logs): preserve upstream errors in routing decision chains

* fix(logs): keep hedge setup attempts and circuit state accurate
Legacy Hedge rectifier retries share one participant sequence, so
legacy-hedge-2-1 and legacy-hedge-2-2 both resolved to the same
decision-chain entry and showed the wrong upstream error. The retry entry
also used the request attempt count as attemptNumber, which made the
terminal failure of an initial-provider retry look like a duplicate and
dropped it from the chain.

Legacy Hedge and Discovery now write routingAttemptId and routingRound on
attempt-level chain entries, the chain keeps entries whose
routingAttemptId differs, and the request detail matches routed entries by
attempt ID only. Reconstructed Discovery attempts use the recorded round;
entries without one are grouped under a new "Round not recorded" heading
in all five locales, instead of being counted in the winner's round.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(providers): show upstream balance for each provider key

供应商管理页现在默认展示每个供应商用自己的密钥查询到的上游余额。

余额来源(参考 all-api-hub 的遥测适配器做法,按域名排序后逐个尝试):
- New API / One API 家族的 /api/usage/token/
- OpenAI 兼容计费的 /v1/dashboard/billing/subscription + /usage
- DeepSeek 的 /user/balance
- Moonshot / Kimi 的 /v1/users/me/balance
- ChatGPT 账号的 /backend-api/wham/usage(参考 codex2api,

额度单位按 500000 quota = 1 USD 归一化为美元;金额保留上游原生币种,
不做汇率换算,避免展示值和上游账单对不上。

接口:
- POST /api/v1/providers/balances:batch 批量查询,默认走缓存
- POST /api/v1/providers/{id}/balance:refresh 单供应商强制刷新

界面:
- 列表视图是带刷新入口的余额方块,服务商视图是单行内联小标签
- 进入视口才登记查询(懒加载),登记的供应商合并成批次顺序发出
- 每 10 分钟自动重新读取一次,与服务端快照缓存期一致
- 查询失败只做弱化展示,不标红

其他:
- 探测请求统一 8 秒超时、响应体 256KB 上限
- 快照缓存键带上地址与密钥指纹,换密钥后旧快照立即失效
- 新增 6 个测试文件覆盖解析、规划、并发、探测、批量存储与展示逻辑

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(providers): address balance review findings

CodeRabbit 的 5 条评审意见,逐条修复:

1. resolveSourceBaseUrl 按来源解析基地址。官方钱包端点(DeepSeek、
   Moonshot、ChatGPT)挂在站点根路径下,供应商被配置成
   `https://api.deepseek.com/v1` 时不再拼出 `/v1/user/balance`;
   网关兼容端点保留子路径挂载,只去掉末尾的 API 版本后缀,
   避免 OpenAI 计费端点拼出 `/v1/v1/dashboard/...`。

2. ProviderBalanceProvider 在 effect 内创建并释放自建的 store。
   StrictMode 的 setup - cleanup - setup 会复用 useMemo 的实例,
   清理后 dispose 过的 store 会让所有余额停留在 idle。

3. load() 记录每个供应商的请求代号,代号已更新的旧批次不再写回,
   手动刷新先完成时不会被后完成的旧请求覆盖。

4. revalidateTracked 逐批捕获失败,一批失败不再中断后续批次,
   也避免定时器产生未处理的 Promise rejection;refreshAll 同样
   继续处理其余批次,并把首个失败交给调用方提示。

5. 非 2xx 响应在抛出前取消响应体。fetchWithDispatcher 走 undici,
   不消费响应体,不取消会一直占用连接直到超时。

测试:新增 4 个用例覆盖基地址解析、旧批次丢弃、逐批失败隔离与
刷新全部的部分失败;把断言重复路径的旧用例改成断言正确 URL。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(providers): cover the balance provider under StrictMode

用 react-dom/client 真实跑一遍 StrictMode 的 setup - cleanup - setup,
直接验证上一提交的修复:effect 重跑后余额仍能加载,外部注入的 store
不会被卸载释放,卸载后定时器不再访问已释放的 store。

dev 模式下供应商列表本身加载不出来(HMR websocket 握手失败,与本功能
无关),改用这组测试给出确定性的证据。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(providers): keep forced balance refreshes from being overwritten

第二轮评审的 3 条意见:

1. 定时重读会顶掉正在进行的强制刷新结果。revalidateTracked 读到的是
   服务端旧快照,却拿到更新的请求代号,于是后完成的强制刷新被判为过期
   而写不回去,界面会停留在旧余额直到下一个周期。现在强制刷新期间给
   供应商打标记,定时重读跳过它们;强制刷新结束后恢复覆盖。

2. 组件测试里不要在异步 act 内轮询 DOM。批量加载由 setTimeout 调度,
   请求可能在 renderStrict 返回后才开始,而 act 内的 commit 时机不确定,
   轮询读到的可能是旧 DOM。改成在 act 内等 fetchBalances 被调用,
   再在 act 外断言渲染结果。

3. 释放顺序用例只该统计卸载触发的 dispose。StrictMode 初次挂载已经跑过
   一次 cleanup,直接断言 toHaveBeenCalled 在旧实现下也会通过。现在先清空
   spy 调用记录,再断言卸载恰好新增一次调用。

另外核对过第 1 条用例确实能拦住回归:把生命周期改回复用已释放实例的
写法时,「effect 重跑后余额仍能加载」会失败。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lynricsy and others added 4 commits September 25, 2026 15:39
…mpty stream (#1491) (#1493)

* fix(proxy): 🐛 treat Anthropic refusal as a deliverable result, not an empty stream

Anthropic can end a stream with message_delta.delta.stop_reason=refusal and no
content blocks. The stream content gate classified that as terminal-before-content
(empty_stream), synthesized a local 502, retried/failed over and counted it toward
the provider circuit breaker, so a single refusing session could open the breaker
for the whole provider.

- frame-classifier: message_delta with stop_reason=refusal is a content signal
  (shared by the enforce gate, incremental frame probe, Discovery/hedge validity,
  protocol observer and client-abort metering), matching the existing refusal
  content rules for openai-chat / openai-responses.
- forwarder non-stream empty-content check and fake-streaming validator accept
  a content-less refusal instead of raising missing_content / no_deliverable.
- StreamPrecommitError now carries the real upstream HTTP status
  (upstream_status_code in the local error body); the provider-error log marks
  gate errors as stream_gate_local with circuitBreakerAccountable.

Empty streams without an explicit refusal keep the existing empty_stream policy.

Closes #1491

Co-authored-by: Wine Fox <fox@ling.plus>

* docs(changelog): 📝 note Anthropic refusal stream gate fix (#1491)

Co-authored-by: Wine Fox <fox@ling.plus>

* fix(proxy): 🐛 align circuitBreakerAccounted log with actual accounting

The provider-error log reported circuitBreakerAccountable from the
request-scoped gate exemption alone, so probe requests, endpoints without
circuit-breaker accounting and attempts that will still retry were logged as
accountable although recordFailure was never called. Compute a single
recordsCircuitFailure predicate (exhausted retries, non-probe, endpoint allows
accounting, not request-scoped) and use it for both the log field
(circuitBreakerAccounted, now on every provider error) and the recordFailure
branch so the two cannot drift.

Refs #1491

Co-authored-by: Wine Fox <fox@ling.plus>

---------

Co-authored-by: Wine Fox <fox@ling.plus>
* fix(providers): detect Sub2API balances and read New API account balance via access token

Sub2API 供应商此前全部被判定为「不支持」:探测顺序只有 New API 的
/api/usage/token/ 与 OpenAI 兼容计费端点,两者在 Sub2API 上都是 404,
而 Sub2API 的余额端点是 GET /v1/usage(供 CC Switch 使用,密钥认证)。

Sub2API:
- 通用兼容端点顺序改为 New API 令牌额度、Sub2API 用量、OpenAI 兼容计费,
  与 all-api-hub 的自动遥测顺序一致
- 按 isValid 布尔字段识别 Sub2API 响应,其他网关同一路径的内容不会被误判
- 钱包模式读取账户钱包余额;quota_limited 读取密钥额度的剩余、上限与已用;
  订阅模式读取各周期最小剩余额度与到期时间,-1 判定为不限量;
  兼容没有 mode 字段的旧版本

New API 系统访问令牌:
- 供应商新增 new_api_access_token 与 new_api_user_id 两个可选字段,
  新增与编辑页面的「余额查询」卡片可填写、替换或清除令牌
- 配置令牌后只查询 /api/user/self 的账户余额(quota / used_quota),
  不再回落到密钥额度
- 用户 ID 按 New-Api-User 及各分支的头名发送;v1.0.0-rc.21 等旧版本
  缺少该头时返回 401,令牌无效时返回 HTTP 200 与 success:false,
  两者都判定为令牌被拒绝
- 令牌只以掩码形式返回给界面与 REST 接口,审计日志中脱敏,余额缓存键
  包含令牌与用户 ID 的指纹

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(providers): restore balance credentials on undo and reject masked tokens

评审意见的修复:

- 单个编辑的撤销数据包含系统访问令牌与用户 ID,撤销时一并恢复;
  审计日志中撤销数据的 newApiAccessToken 同样脱敏
- REST 接口拒绝把列表返回的掩码令牌(前 4 位 + •••••• + 后 4 位)
  当作真实令牌保存,新增与编辑都返回 422
- 令牌输入框、显示切换按钮与用户 ID 输入框补充可访问名称与 aria-pressed
- 额度耗尽的 Sub2API 响应样例改为与上游一致的 isValid:true,
  另补停用密钥 isValid:false 的用例

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ding113 <gabriel@lpupe.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(settings): add a memory admission switch, disabled by default

Memory admission (introduced in v0.9.6) now follows the system setting
enableMemoryAdmission, which defaults to off. When off, the governor only
accounts leases and never refuses or queues them, workers do not request
credits from the primary, request bodies and stream gate prefixes stay in
memory without spilling to disk, the STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP
sub-limit and the pre-materialization heap check are skipped, and no local
429 local_capacity_exceeded is produced. When on, every existing admission
mechanism applies unchanged.

The proxy handler reads cached system settings before reading the inbound
body and syncs the switch into the process governor; saving the setting
invalidates the cache across processes as usual.

- schema: system_settings.enable_memory_admission boolean not null default false (migration 0124)
- settings page: toggle next to high-concurrency mode, i18n for 5 locales
- v1 API/OpenAPI, server action, legacy admin route, validation, cache defaults, column downgrade ladder

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(memory): count cgroup usage as working set and harden the admission switch

The resource snapshot subtracted raw cgroup usage (memory.current,
memory.usage_in_bytes, memsw usage) from the limit. That usage includes
reclaimable page cache from spooled bodies, loaded code and logs, so a
memory-limited container drifted to zero headroom and every request hit
the local 429 after running for a while. Usage is now the working set,
memory.current - inactive_file on v2 and usage - total_inactive_file on
v1 (memsw included), matching kubelet/cAdvisor. In a 1 GiB cgroup with
704 MiB of written file cache the available estimate goes from 279 MiB to
983 MiB, while real anonymous usage still drives the budget to zero.

Review fixes for the admission switch:
- only sync the switch from the settings object held in the process
  cache, so a stale query finishing after invalidation or the fallback
  object returned on a failed read cannot flip it for the whole process
- notify listeners when the switch changes and drain stream gate
  sub-limit waiters immediately when it turns off

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(memory): shrink the request body lease to its retained size after parsing

The body_materialize lease was sized for the parse peak, 8192 + 8x bytes +
structure, and held for the whole request lifetime, including long
streaming responses and background owners. Measured on Claude Code style
bodies the lease was 3.2x-4.4x the memory the request actually keeps
(buffer plus parsed object), so a backlog of long streams booked leases
far above real usage; once usedBytes passed the limit, which follows real
free memory, every new request was rejected at body_intake with a local
429 (reported: usedBytes 4.06 GB, limitBytes 2.94 GB, waiting 0).

The parse peak stays reserved while the body is decoded and parsed. After
parsing, the lease shrinks to the retained estimate 8192 + 4x bytes +
structure (raw bytes, UTF-16 worst-case strings, object structure and one
outbound serialized copy) on the JSON, non-JSON and multipart paths.
For a 3.69 MiB body the held lease goes from 31.9 MiB to 17.2 MiB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T05:50:03.591161Z 69dd8e0 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: aa33e206-8010-434a-84ad-d2c1ba751bcc

📥 Commits

Reviewing files that changed from the base of the PR and between 5dc1fd4 and 69dd8e0.

📒 Files selected for processing (52)
  • .env.example
  • docs/research/ttft-memory-aware-admission.md
  • drizzle/0124_system_settings_memory_admission.sql
  • drizzle/meta/0124_snapshot.json
  • drizzle/meta/_journal.json
  • messages/en/settings/config.json
  • messages/ja/settings/config.json
  • messages/ru/settings/config.json
  • messages/zh-CN/settings/config.json
  • messages/zh-TW/settings/config.json
  • server-lib/memory-governor.js
  • server-lib/resource-snapshot.js
  • src/actions/system-config.ts
  • src/app/[locale]/settings/config/_components/system-settings-form.tsx
  • src/app/[locale]/settings/config/page.tsx
  • src/app/api/admin/system-config/route.ts
  • src/app/v1/_lib/proxy-handler.ts
  • src/app/v1/_lib/proxy/session.ts
  • src/app/v1/_lib/proxy/stream-gate/prebuffer-budget.ts
  • src/drizzle/schema.ts
  • src/lib/api-client/v1/openapi-types.gen.ts
  • src/lib/api/v1/schemas/system-config.ts
  • src/lib/body-store/allocation-estimate.ts
  • src/lib/body-store/byte-store.ts
  • src/lib/body-store/request-body-store.ts
  • src/lib/config/system-settings-cache.ts
  • src/lib/memory/governor.ts
  • src/lib/validation/schemas.ts
  • src/repository/_shared/transformers.ts
  • src/repository/system-config.ts
  • src/types/system-config.ts
  • tests/integration/proxy-hedge-lifecycle.test.ts
  • tests/unit/actions/system-config-memory-admission-settings.test.ts
  • tests/unit/actions/system-config-save.test.ts
  • tests/unit/proxy/memory-aware-body-store.test.ts
  • tests/unit/proxy/memory-aware-contention.test.ts
  • tests/unit/proxy/memory-aware-discovery.test.ts
  • tests/unit/proxy/memory-aware-disk-timeout.test.ts
  • tests/unit/proxy/memory-aware-lifetime.test.ts
  • tests/unit/proxy/memory-aware-retained-lease.test.ts
  • tests/unit/proxy/memory-aware-switch.test.ts
  • tests/unit/proxy/proxy-handler-public-errors.test.ts
  • tests/unit/proxy/routing-trace.test.ts
  • tests/unit/proxy/stream-gate-content-gate.test.ts
  • tests/unit/proxy/stream-gate-forwarder-integration.test.ts
  • tests/unit/proxy/stream-gate-ttft-regression.test.ts
  • tests/unit/repository/system-config-degradation-ladder.test.ts
  • tests/unit/repository/system-config-update-missing-columns.test.ts
  • tests/unit/server-memory-governor.test.ts
  • tests/unit/server-memory-growth.test.ts
  • tests/unit/server-memory-resources.test.ts
  • tests/unit/settings/system-settings-form-memory-admission-toggle.test.tsx

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

新增默认关闭的内存准入系统设置。开启后,代理会按内存额度处理请求正文和流式门控前缀,并在解析完成后调整请求正文租约。资源快照会从 cgroup 用量中扣除非活动文件缓存。

Changes

内存准入设置

Layer / File(s) Summary
设置持久化与管理
drizzle/*, src/drizzle/schema.ts, src/types/system-config.ts, src/repository/*, src/lib/api/*, src/actions/system-config.ts, src/app/[locale]/settings/config/*, src/app/api/admin/system-config/route.ts, messages/*/settings/config.json, tests/unit/actions/*, tests/unit/repository/*, tests/unit/settings/*
新增默认关闭的 enableMemoryAdmission 设置,并接入数据库、配置 API、设置表单、翻译和相关测试。

代理内存准入

Layer / File(s) Summary
开关同步与额度控制
.env.example, docs/research/ttft-memory-aware-admission.md, server-lib/memory-governor.js, src/lib/memory/governor.ts, src/app/v1/_lib/proxy-handler.ts, src/app/v1/_lib/proxy/stream-gate/prebuffer-budget.ts, src/lib/config/system-settings-cache.ts, tests/unit/proxy/memory-aware-switch.test.ts, tests/unit/proxy/proxy-handler-public-errors.test.ts, tests/unit/server-memory-governor.test.ts, tests/unit/server-memory-growth.test.ts, tests/unit/proxy/stream-gate-*, tests/integration/proxy-hedge-lifecycle.test.ts
内存治理器新增启用状态、状态监听和快照字段。代理入口根据当前缓存设置同步开关;流式门控预算随开关调整。测试覆盖开关和既有内存治理场景。
请求正文额度与持有周期
src/lib/body-store/*, src/app/v1/_lib/proxy/session.ts, tests/unit/proxy/memory-aware-*.test.ts
正文存储根据准入状态选择内存保留或磁盘溢写。解析完成后,租约调整为估算的持有额度;相关测试覆盖 JSON、原始文本和 multipart 请求。
cgroup 可用内存计算
server-lib/resource-snapshot.js, tests/unit/server-memory-resources.test.ts, docs/research/ttft-memory-aware-admission.md
cgroup v2 使用 inactive_file、v1 使用 total_inactive_file 扣减内存用量;v1 memsw 用量也执行扣减。测试覆盖统计缺失和边界情况。

Priority: ⚪ Not assessed

Estimated code review effort: 3 (Moderate) | ~25 minutes

Merge Risk: ⚪ Minimal · up to 69dd8

The default-off setting, retained-memory accounting, and cgroup adjustments have no established blocking defect. The change is mergeable subject to normal checks.

Security Architecture Review

Security architecture risk: 🟠 High · up to 69dd8

Upgrading disables memory protection by default, allowing request bodies to consume unrestricted application memory before the proxy’s authentication checks. Administrator authorization remains intact, but database compatibility handling can also discard an attempt to enable protection.

Retained concerns

  • High · security · inferred: The upgrade removes default memory-exhaustion containment from the pre-authentication body-processing path. With admission disabled, incoming bodies remain in memory without the normal spill threshold, lease acquisition and growth bypass capacity limits, and materialization bypasses its heap-headroom check. Large or concurrent requests can therefore exhaust shared application memory before the proxy authentication and rate-limit pipeline runs. Actual external exploitability depends on ingress controls that were not established.
  • Medium · security · observed: When a settings mutation encounters a missing-column error, the compatibility ladder removes enableMemoryAdmission from the update before retrying. A successful retry can therefore acknowledge the request without persisting the administrator’s requested protection state. This is conditional on a degraded schema; the migration and API response containing the resulting settings are counterevidence against treating every successful save as misleading.
Security review details

Security Blast Radius

  • inferred — The principal exposure is availability of the reachable proxy process and potentially its shared worker deployment. Requests consume memory before per-session authorization and rate limiting, so other users sharing that runtime can be affected without the attacker obtaining configuration privileges. Wider host or cross-service impact was not established.

Security Findings and Attack Paths

  • inferred — An attacker who can reach the body-processing path can send large uncompressed bodies or concurrent requests while admission is disabled. The body store grows in memory, capacity admission and heap-headroom rejection are bypassed, and authentication occurs afterward. Compressed-input and decoded-size checks remain countercontrols, but do not provide a general aggregate memory bound for this path.

Trust Boundaries and Controls

  • observed — Configuration authority remains administrator-only. The security change is instead removal of default resource containment on an already body-before-authentication path; intact administrative authorization does not constrain memory consumed by incoming request bodies.

Resilience and Maintainability Implications

  • observed — Temporary body-processing leases have finally-based cleanup. The retained lease transfers to request lifetime ownership and shrinks after parsing; response owners, background retainers and forced-end disposal supply terminal cleanup paths. These mechanisms contain retained-memory leaks but cannot prevent exhaustion during unrestricted allocation.

Hardening Proposals

  • proposed — Separate optional performance-oriented admission from non-disableable safety bounds: retain pre-authentication byte and aggregate-memory limits, disk-spill containment and heap-headroom rejection even when admission queuing is disabled. Reject or explicitly report an admission-policy mutation that cannot be persisted during schema degradation.
🚥 Pre-merge checks | ✅ 2 | ❓ 3

❌ Failed checks (3 inconclusive)

Check name Status Explanation Resolution
Title check ❓ Inconclusive 标题“release v0.9.7”仅表示版本发布,未说明本次主要的内存准入功能变更。标题信息不足,无法明确概括变更内容。 将标题改为描述主要变更的单句,例如“Add configurable memory admission control”。
Description check ❓ Inconclusive Pull request 未提供描述。当前信息不足,无法确认作者对内存准入开关、资源计算和相关测试的变更意图。 补充简短描述,说明新增内存准入配置、默认行为、资源准入逻辑、数据库迁移和测试覆盖范围。
Docstring Coverage ❓ Inconclusive Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 41 files. (10 skipped… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 41 files. (10 skipped: 9 unsupported, 1 too large.)

✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment on lines +97 to +99
const hotBytes =
this.options.hotBytes ?? (getMemoryGovernor().enabled ? HOT_BYTES : Number.POSITIVE_INFINITY);
if (!this.file && capacity <= hotBytes && (await this.tryGrow(this.scratchBytes + capacity))) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P0 security Unbounded pre-auth body retention When memory admission is disabled, as it is by default, this branch keeps uncompressed request bodies in memory instead of spilling them to disk. The body reader checks size only for compressed input, and parsing happens before authentication. An unauthenticated client can therefore send a large body that exhausts the worker instead of receiving a bounded capacity response. Preserve an independent size limit or disk fallback. How this was verified: Uncompressed bytes reach the in-memory body store before the authentication guard, while the body-size checks apply only to compressed input.

Knowledge Base Used: Memory-aware request admission

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/lib/body-store/byte-store.ts
Line: 97-99

Comment:
**Unbounded pre-auth body retention** When memory admission is disabled, as it is by default, this branch keeps uncompressed request bodies in memory instead of spilling them to disk. The body reader checks size only for compressed input, and parsing happens before authentication. An unauthenticated client can therefore send a large body that exhausts the worker instead of receiving a bounded capacity response. Preserve an independent size limit or disk fallback. **How this was verified:** Uncompressed bytes reach the in-memory body store before the authentication guard, while the body-size checks apply only to compressed input.

**Knowledge Base Used:** [Memory-aware request admission](https://app.greptile.com/ygxz/-/custom-context/knowledge-base/ding113/claude-code-hub/-/docs/memory-aware-admission.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Comment on lines 283 to 286
() =>
process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP
governor.enabled && process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP
? resolveStreamGateGlobalPrebufferByteCap()
: Number.MAX_SAFE_INTEGER,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Aggregate prefix cap bypassed When memory admission is disabled, this resolver ignores a configured STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP. Each enforcing stream still has its own prefix limit, but concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. That can exhaust the worker under concurrent streaming traffic. Keep an aggregate bound independent of the admission toggle.

Knowledge Base Used: Request sessions and streaming

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/app/v1/_lib/proxy/stream-gate/prebuffer-budget.ts
Line: 283-286

Comment:
**Aggregate prefix cap bypassed** When memory admission is disabled, this resolver ignores a configured `STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP`. Each enforcing stream still has its own prefix limit, but concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. That can exhaust the worker under concurrent streaming traffic. Keep an aggregate bound independent of the admission toggle.

**Knowledge Base Used:** [Request sessions and streaming](https://app.greptile.com/ygxz/-/custom-context/knowledge-base/ding113/claude-code-hub/-/docs/request-sessions-and-streaming.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

@ding113
ding113 merged commit 5f32d8f into main Sep 30, 2026
19 of 21 checks passed
@github-actions

Copy link
Copy Markdown
Contributor

🧪 测试结果

测试类型 状态
代码质量 ✅
单元测试 ✅
集成测试 ✅
API 测试 ✅

总体结果: ✅ 所有测试通过

@github-actions github-actions Bot added the size/XL Extra Large PR (> 1000 lines) label Sep 30, 2026
) {
// 内存准入关闭时正文全部留在内存,不落盘。
const hotBytes =
this.options.hotBytes ?? (getMemoryGovernor().enabled ? HOT_BYTES : Number.POSITIVE_INFINITY);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[CRITICAL] [ERROR-SILENT] Memory admission bypass allows unbounded pre-authentication memory retention

Why this is a problem:
When memory admission is disabled (the default), line 98 sets hotBytes to Number.POSITIVE_INFINITY, which means ALL request bodies are kept in memory regardless of size. Since body parsing happens BEFORE authentication (the body is read in the proxy handler before the auth guard runs), an unauthenticated attacker can send arbitrarily large request bodies to exhaust worker memory.

The original code had a fixed HOT_BYTES limit (256 KiB) that applied unconditionally. This PR removes that protection when enabled: false.

Evidence from codebase:

  • Body reading happens in proxy-handler.ts before authentication guards
  • The ByteStore only checks compressed input size, not uncompressed bytes
  • No independent size limit exists when hotBytes = Infinity

Suggested fix:

// Line 96-99
// Preserve an independent size limit even when memory admission is disabled
const hotBytes = this.options.hotBytes ?? 
  (getMemoryGovernor().enabled ? HOT_BYTES : 10 * 1024 * 1024); // 10 MiB max when disabled

This preserves a reasonable upper bound (10 MiB) while still being much larger than the original 256 KiB, preventing the pre-auth DoS vector while maintaining the intended behavior of keeping normal bodies in memory.

globalState[STREAM_GATE_PREBUFFER_BUDGET_SYMBOL] ??= new StreamGatePrebufferBudget(
() =>
process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP
governor.enabled && process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[HIGH] [LOGIC-BUG] Aggregate stream prefix cap bypassed when memory admission disabled

Why this is a problem:
When memory admission is disabled (the default), line 284 changes the resolver to return Number.MAX_SAFE_INTEGER regardless of the STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP environment variable. This means the configured aggregate cap (default 256 MiB) is completely ignored.

While each individual stream still has its own prefix limit, concurrent streams can collectively retain far more than the configured process-wide cap because their prefixes stay in memory rather than spilling. Under high concurrent streaming traffic, this can exhaust worker memory.

Original behavior: STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP was always respected as a hard aggregate limit.

New behavior: When enabled: false (default), the aggregate cap is ignored entirely.

Evidence from code:

  • Line 284: governor.enabled && process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP returns false when disabled
  • The resolver then returns Number.MAX_SAFE_INTEGER (no limit)
  • Each stream's individual limit still applies, but no aggregate bound exists

Suggested fix:

// Line 283-286
globalState[STREAM_GATE_PREBUFFER_BUDGET_SYMBOL] ??= new StreamGatePrebufferBudget(
  () =>
    process.env.STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP
      ? resolveStreamGateGlobalPrebufferByteCap()
      : (governor.enabled ? Number.MAX_SAFE_INTEGER : 1024 * 1024 * 1024), // 1 GiB aggregate when disabled
  governor
);

This preserves an aggregate bound (1 GiB) independent of memory admission, preventing concurrent streams from collectively exhausting memory while still being generous for the disabled mode.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review Summary

This PR introduces a memory admission toggle (default: OFF) to address TTFB and memory management issues. While the implementation is comprehensive with excellent test coverage, two critical security/stability issues must be fixed before merge.

PR Size: XL

  • Lines changed: 7197 (7112 additions, 85 deletions)
  • Files changed: 52

Split Recommendation: This XL PR would benefit from splitting:

  1. Phase 1: Memory admission toggle infrastructure + cgroup accounting fixes (migrations, settings cache, UI)
  2. Phase 2: ByteStore and stream gate integration with the toggle
  3. Phase 3: Provider balance features (independent feature)

However, given the PR is already in review with comprehensive testing, the split is optional. The two critical issues below MUST be fixed regardless.


Issues Found

Category Critical High Medium Low
Security 1 0 0 0
Logic/Bugs 0 1 0 0
Error Handling 0 0 0 0
Types 0 0 0 0
Comments/Docs 0 0 0 0
Tests 0 0 0 0
Simplification 0 0 0 0

Critical Issues (Must Fix)

1. [CRITICAL] Unbounded pre-auth body retention (byte-store.ts:98)

Severity: Critical
Category: Security - DoS vector
Confidence: 95/100

When memory admission is disabled (the default), hotBytes is set to Number.POSITIVE_INFINITY. This allows unauthenticated attackers to send arbitrarily large request bodies that are kept entirely in memory, exhausting the worker. Body parsing occurs BEFORE authentication guards run in the proxy pipeline.

Impact: Production DoS vulnerability in default configuration.


High Priority Issues (Should Fix)

1. [HIGH] Aggregate prefix cap bypass (prebuffer-budget.ts:284)

Severity: High
Category: Logic bug - memory exhaustion
Confidence: 90/100

When memory admission is disabled, the configured STREAM_GATE_GLOBAL_PREBUFFER_BYTE_CAP (default 256 MiB) is completely ignored. Concurrent streaming requests can collectively retain unlimited memory in their prefixes, potentially exhausting the worker under high concurrent load.

Impact: Worker instability under concurrent streaming traffic in default configuration.


Review Coverage

  • Logic and correctness - 2 critical issues found
  • Security (OWASP Top 10) - 1 pre-auth DoS vector
  • Error handling - Clean
  • Type safety - Clean
  • Documentation accuracy - Good (comprehensive Chinese comments)
  • Test coverage - Excellent (136 memory tests, 93.40% coverage)
  • Code clarity - Good

Positive Observations (Not Commenting on Good Code)

The following aspects are well-executed:

  • Comprehensive test coverage with 136 memory-specific tests
  • Proper cgroup v1/v2 working set calculation (inactive_file exclusion)
  • Event-driven cache invalidation with Redis pub/sub
  • 5-language i18n support for all UI changes
  • Database migrations properly versioned (0123, 0124)
  • OpenAPI schema synchronization maintained

Recommendation

Status: ❌ NOT SAFE TO MERGE until critical issues are addressed.

The two issues identified create real security and stability risks in the default configuration (memory admission disabled). Both can be fixed with small, surgical changes to add bounded fallback limits instead of Infinity / MAX_SAFE_INTEGER.

Once fixed, this PR represents a significant improvement to memory management and operational visibility.


Automated review by Claude AI Code Review

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core area:session enhancement New feature or request size/XL Extra Large PR (> 1000 lines)

Projects

Status: Done

6 participants