Found while reviewing #1089 (the orchestration_wait_agents cap for #827). The cap shortens the exposure window; it does not close the hole, and this is the hole.
The mechanism
BuiltInMcpHttpHost builds the POST transport as new StreamableHTTPServerTransport({ sessionIdGenerator: undefined }) — no eventStore, no enableJsonResponse. Three consequences compose into a permanent loss:
- Tool replies are SSE streams. With
enableJsonResponse unset the SDK streams the response rather than returning a JSON body.
- That stream carries zero bytes for the whole call. The server transport has no keep-alive, and the host's own comment records that Agent Code sends no progress notifications. A wait of any length is an idle connection.
- It is not resumable. With no
eventStore there is nothing to replay from, so when a client exhausts its SSE reconnect attempts (maxRetries: 2 in the SDK client), the pending callTool() rejects and the reply is gone — even though the work completed.
So the failure in #827 is drop-triggered, not clock-triggered. Any long tool call is exposed, not only orchestration_wait_agents; that one is simply the only blocking wait in src/mcp/ today.
Two candidate fixes, neither verified
- Emit progress notifications from long-running handlers. This is the MCP-native answer: the stream stops being idle, and the SDK resets its request timeout on progress (Claude Code already wires
onprogress at its callTool site). It also gives callers real feedback during a wait.
enableJsonResponse: true on the POST transport, which removes the held-open stream entirely. Viable precisely because we never send server notifications on it — but it changes the response shape for every tool, so it needs its own verification.
An eventStore would make the stream resumable, but it is the heaviest of the three and only helps after a drop has already happened.
Why this is filed rather than fixed in #1089
#1089 is a bound on one tool's wait, justified by numbers that are checkable (the SDK's DEFAULT_REQUEST_TIMEOUT_MSEC = 60000, this repo's existing 30 s long-poll bound). This issue is a transport change that affects every built-in MCP tool and needs its own evidence — ideally a recorded reproduction of the drop before anything is changed.
Found while reviewing #1089 (the
orchestration_wait_agentscap for #827). The cap shortens the exposure window; it does not close the hole, and this is the hole.The mechanism
BuiltInMcpHttpHostbuilds the POST transport asnew StreamableHTTPServerTransport({ sessionIdGenerator: undefined })— noeventStore, noenableJsonResponse. Three consequences compose into a permanent loss:enableJsonResponseunset the SDK streams the response rather than returning a JSON body.eventStorethere is nothing to replay from, so when a client exhausts its SSE reconnect attempts (maxRetries: 2in the SDK client), the pendingcallTool()rejects and the reply is gone — even though the work completed.So the failure in #827 is drop-triggered, not clock-triggered. Any long tool call is exposed, not only
orchestration_wait_agents; that one is simply the only blocking wait insrc/mcp/today.Two candidate fixes, neither verified
onprogressat itscallToolsite). It also gives callers real feedback during a wait.enableJsonResponse: trueon the POST transport, which removes the held-open stream entirely. Viable precisely because we never send server notifications on it — but it changes the response shape for every tool, so it needs its own verification.An
eventStorewould make the stream resumable, but it is the heaviest of the three and only helps after a drop has already happened.Why this is filed rather than fixed in #1089
#1089 is a bound on one tool's wait, justified by numbers that are checkable (the SDK's
DEFAULT_REQUEST_TIMEOUT_MSEC = 60000, this repo's existing 30 s long-poll bound). This issue is a transport change that affects every built-in MCP tool and needs its own evidence — ideally a recorded reproduction of the drop before anything is changed.