Skip to content

fix: hold streamed requests to the concurrency bound and define engine failure - #33

Merged
github-actions[bot] merged 1 commit into
mainfrom
feat/20-engine-failure-behaviour
Oct 3, 2026
Merged

github-actions[bot] merged 1 commit into
mainfrom
feat/20-engine-failure-behaviour

Conversation

@Yash-Chindam

Copy link
Copy Markdown
Owner

Sixth PR closing gaps between the design spec and the implementation. This one starts with a bug.

Bug fixed

A streamed request returned its StreamingResponse from inside async with admission.slot(). Leaving that block frees the slot, and the body is generated afterwards, so streamed requests never counted against ROUTER_MAX_CONCURRENCY. With stream: true, the bounded-queue behaviour in §9 and §13 did not apply. A stream now holds its slot until it ends. The slot is released by the response object as well as the generator, because a generator that is never iterated (client gone before the first chunk) never runs its own cleanup.

Gap closed

§13: "Define behavior for GPU out-of-memory and node loss", "Stop routing to unhealthy replicas", "graceful shutdown".

  • Out of memory: 503 engine_out_of_memory, with a message saying to shorten the prompt or lower max_tokens.
  • Node loss: after ROUTER_ENGINE_FAILURE_THRESHOLD consecutive failures a circuit opens. Requests get 503 engine_unavailable immediately, with Retry-After set to the remaining cooldown, and /readyz fails.
  • Recovery: a healthy probe half-opens the circuit and admits one trial request; only its success closes it. A failed trial restarts the cooldown. A cancelled or abandoned trial frees the trial slot.
  • Shutdown: readiness fails first, then admitted requests get ROUTER_SHUTDOWN_GRACE_SECONDS before the engine client closes.
  • The engine's error body is inspected for an out-of-memory report and discarded, never forwarded.
  • New metric router_engine_circuit_open; README gains a failure-behaviour table and a note on cold start for the scale-to-zero tier.

Decision to check

A failed request is not retried on another model. All local models share the one engine a gateway faces, so a retry would hit the same failure. If you run one gateway across several engines, that reasoning no longer holds.

Test plan

  • 21 new tests, including a live stream holding the only slot while a second request is rejected overloaded, and the slot returning when the stream ends or fails
  • ruff format --check ., ruff check ., mypy clean
  • pytest tests/unit tests/integration: 283 passed, coverage 98%
  • Playwright end-to-end suite: 9 passed locally
  • Out-of-memory detection matches on the text vLLM and PyTorch use for CUDA OOM; it has not been triggered on a real GPU

🤖 Generated with Claude Code

…e failure

A streamed request returned its response from inside the admission slot,
so the slot was freed before generation began and streaming escaped the
concurrency bound that backpressure depends on. A stream now holds its
slot until it ends, including when the client disconnects before the
first chunk or the engine fails mid-stream.

Section 13 also requires behaviour to be defined for GPU out-of-memory
and node loss, and none was: every engine failure was one generic 502,
each waiting out its own timeout.

Out of memory now has its own error type and tells the caller what to
change, since the same request at the same size cannot succeed. After
consecutive failures a circuit opens: requests fail immediately with
retry guidance instead of timing out against a node that is gone, and
readiness fails so the orchestrator stops routing here. A healthy probe
lets exactly one trial request through, and only its success closes the
circuit. A cancelled or abandoned trial does not wedge it.

An engine error body is checked for an out-of-memory report and then
discarded rather than forwarded, since it can echo the prompt.

Shutdown fails readiness first, then gives admitted requests a grace
period before the engine client closes underneath them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation area/api area/tests labels Oct 3, 2026
@github-actions
github-actions Bot merged commit ce7d772 into main Oct 3, 2026
6 checks passed
@github-actions
github-actions Bot deleted the feat/20-engine-failure-behaviour branch October 3, 2026 14:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/api area/tests documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant