Skip to content

Built-in MCP replies ride a silent, non-resumable SSE stream — a dropped connection loses the reply for good #1091

Description

@Juliusolsson05

Found while reviewing #1089 (the orchestration_wait_agents cap for #827). The cap shortens the exposure window; it does not close the hole, and this is the hole.

The mechanism

BuiltInMcpHttpHost builds the POST transport as new StreamableHTTPServerTransport({ sessionIdGenerator: undefined }) — no eventStore, no enableJsonResponse. Three consequences compose into a permanent loss:

  1. Tool replies are SSE streams. With enableJsonResponse unset the SDK streams the response rather than returning a JSON body.
  2. That stream carries zero bytes for the whole call. The server transport has no keep-alive, and the host's own comment records that Agent Code sends no progress notifications. A wait of any length is an idle connection.
  3. It is not resumable. With no eventStore there is nothing to replay from, so when a client exhausts its SSE reconnect attempts (maxRetries: 2 in the SDK client), the pending callTool() rejects and the reply is gone — even though the work completed.

So the failure in #827 is drop-triggered, not clock-triggered. Any long tool call is exposed, not only orchestration_wait_agents; that one is simply the only blocking wait in src/mcp/ today.

Two candidate fixes, neither verified

  • Emit progress notifications from long-running handlers. This is the MCP-native answer: the stream stops being idle, and the SDK resets its request timeout on progress (Claude Code already wires onprogress at its callTool site). It also gives callers real feedback during a wait.
  • enableJsonResponse: true on the POST transport, which removes the held-open stream entirely. Viable precisely because we never send server notifications on it — but it changes the response shape for every tool, so it needs its own verification.

An eventStore would make the stream resumable, but it is the heaviest of the three and only helps after a drop has already happened.

Why this is filed rather than fixed in #1089

#1089 is a bound on one tool's wait, justified by numbers that are checkable (the SDK's DEFAULT_REQUEST_TIMEOUT_MSEC = 60000, this repo's existing 30 s long-poll bound). This issue is a transport change that affects every built-in MCP tool and needs its own evidence — ideally a recorded reproduction of the drop before anything is changed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    class:C3-silent-failureThe app knows it failed and does not sayneeds-evidenceNeeds a recording/real use to settlesev:P2Real bug with a workaroundtype:bugSomething works wrong

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions