fix: hold streamed requests to the concurrency bound and define engine failure - #33
Merged
Merged
Conversation
…e failure A streamed request returned its response from inside the admission slot, so the slot was freed before generation began and streaming escaped the concurrency bound that backpressure depends on. A stream now holds its slot until it ends, including when the client disconnects before the first chunk or the engine fails mid-stream. Section 13 also requires behaviour to be defined for GPU out-of-memory and node loss, and none was: every engine failure was one generic 502, each waiting out its own timeout. Out of memory now has its own error type and tells the caller what to change, since the same request at the same size cannot succeed. After consecutive failures a circuit opens: requests fail immediately with retry guidance instead of timing out against a node that is gone, and readiness fails so the orchestrator stops routing here. A healthy probe lets exactly one trial request through, and only its success closes the circuit. A cancelled or abandoned trial does not wedge it. An engine error body is checked for an out-of-memory report and then discarded rather than forwarded, since it can echo the prompt. Shutdown fails readiness first, then gives admitted requests a grace period before the engine client closes underneath them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sixth PR closing gaps between the design spec and the implementation. This one starts with a bug.
Bug fixed
A streamed request returned its
StreamingResponsefrom insideasync with admission.slot(). Leaving that block frees the slot, and the body is generated afterwards, so streamed requests never counted againstROUTER_MAX_CONCURRENCY. Withstream: true, the bounded-queue behaviour in §9 and §13 did not apply. A stream now holds its slot until it ends. The slot is released by the response object as well as the generator, because a generator that is never iterated (client gone before the first chunk) never runs its own cleanup.Gap closed
§13: "Define behavior for GPU out-of-memory and node loss", "Stop routing to unhealthy replicas", "graceful shutdown".
503 engine_out_of_memory, with a message saying to shorten the prompt or lowermax_tokens.ROUTER_ENGINE_FAILURE_THRESHOLDconsecutive failures a circuit opens. Requests get503 engine_unavailableimmediately, withRetry-Afterset to the remaining cooldown, and/readyzfails.ROUTER_SHUTDOWN_GRACE_SECONDSbefore the engine client closes.router_engine_circuit_open; README gains a failure-behaviour table and a note on cold start for the scale-to-zero tier.Decision to check
A failed request is not retried on another model. All local models share the one engine a gateway faces, so a retry would hit the same failure. If you run one gateway across several engines, that reasoning no longer holds.
Test plan
overloaded, and the slot returning when the stream ends or failsruff format --check .,ruff check .,mypycleanpytest tests/unit tests/integration: 283 passed, coverage 98%🤖 Generated with Claude Code