Repository navigation
Streamable HTTP: ASGI application returns before the full SSE body is sent #3494
Description
Activity
- added a commit that references this issue
on Sep 11, 2026 - addedv2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)v1Affects the v1.x maintenance lineAffects the v1.x maintenance line
on Sep 11, 2026 Local investigation suggests a restart-specific dependency mechanism worth separating from the final-send race hypothesis.
On SDK main
9972c21/ mcp 2.2.0, Windows, Python 3.11.5, AnyIO 4.10.0, Starlette 1.6.0, h11 0.16.0, httpx2/httpcore2 2.5.0, using real Uvicorn HTTP connections andprotocolVersion=2025-06-18:sse-starlette Uvicorn Second initialization after stopping the first server 3.4.11 0.52.4 RemoteProtocolError, incomplete chunked response3.0.2 0.52.4 Succeeds 3.4.11 0.52.4, diagnostic reset of AppStatus.should_exitSucceeds The 3.4.11 watcher captures the first Uvicorn instance, then copies its shutdown state into global
AppStatus.should_exit. That flag remains true for the second server, whose SSE exit listener immediately cancels the response. The probe records the actual captured server and waits for that watcher's shutdown broadcast; it does not simulate disconnects. Resetting the flag is only a causal control, not a recommended fix.Which
sse-starletteversion was installed in the reported runs? This reproduces the restart symptom, but does not yet explain the fresh-process CPU-load failures. Also, Uvicorn 0.52.4's h11receive()waits on an event while the response is incomplete, so a receiver repeatedly returning empty requests would not reproduce that implementation.AI disclosure: Codex investigated the code, wrote and ran the diagnostic, and drafted this comment.
Reproducer and commands
Save as
reproduce_restart.pyin an SDK checkout. With uv >=0.9.5:uv run --frozen --with sse-starlette==3.4.11 --with uvicorn==0.52.4 python reproduce_restart.py uv run --frozen --with sse-starlette==3.4.11 --with uvicorn==0.52.4 python reproduce_restart.py --reset-between uv run --frozen --with uvicorn==0.52.4 python reproduce_restart.py
"""Diagnose MCP initialization after programmatic Uvicorn restarts. The private SSE exit listener observes the dependency's background watcher; it does not set the shutdown flag. --reset-between is a diagnostic control, not a proposed production workaround. """ import argparse import json import platform import socket from functools import partial from importlib.metadata import version from unittest.mock import patch import anyio import httpx2 import sse_starlette.sse as sse import uvicorn from mcp.server.mcpserver import MCPServer class ReadyServer(uvicorn.Server): def __init__(self, config: uvicorn.Config) -> None: super().__init__(config) self.ready = anyio.Event() async def startup(self, sockets: list[socket.socket] | None = None) -> None: await super().startup(sockets) self.ready.set() async def attempt(label: str) -> tuple[ReadyServer, dict[str, str | bool]]: app = MCPServer("restart-probe").streamable_http_app() server = ReadyServer(uvicorn.Config(app, log_level="error", lifespan="on")) with socket.socket() as listener, anyio.fail_after(5): listener.bind(("127.0.0.1", 0)) port = listener.getsockname()[1] async with anyio.create_task_group() as tasks: tasks.start_soon(partial(server.serve, sockets=[listener])) with anyio.fail_after(5): await server.ready.wait() observation: dict[str, str | bool] = {"phase": label, "exit_before": sse.AppStatus.should_exit} try: async with httpx2.AsyncClient(trust_env=False, timeout=5) as client: response = await client.post( f"http://127.0.0.1:{port}/mcp", headers={"Accept": "application/json, text/event-stream"}, json={ "jsonrpc": "2.0", "id": 1, "method": "initialize", "params": { "protocolVersion": "2025-06-18", "capabilities": {}, "clientInfo": {"name": "probe", "version": "1"}, }, }, ) response.raise_for_status() data = next(line[6:] for line in response.text.splitlines() if line.startswith("data: ")) result = json.loads(data) assert result["id"] == 1 assert result["result"]["protocolVersion"] == "2025-06-18" observation["result"] = "initialize succeeded; complete HTTP response" except httpx2.HTTPError as error: observation["result"] = f"{type(error).__name__}: {error}" finally: server.should_exit = True print(json.dumps(observation), flush=True) return server, observation async def main(reset_between: bool) -> None: packages = ("mcp", "mcp-types", "sse-starlette", "uvicorn", "anyio", "starlette", "h11", "httpx2", "httpcore2") print( json.dumps( { "python": platform.python_version(), "platform": platform.system(), "architecture": platform.machine(), "packages": {name: version(name) for name in packages}, } ), flush=True, ) assert not sse.AppStatus.should_exit captured: list[object] = [] capture_ready = anyio.Event() original_lookup = getattr(sse, "_get_uvicorn_server", None) def record_lookup() -> object: result = original_lookup() captured.append(result) capture_ready.set() return result # Observe which live server the existing watcher captures; preserve its return value. with patch.object(sse, "_get_uvicorn_server", record_lookup, create=True): first, first_result = await attempt("first") assert first_result["result"] == "initialize succeeded; complete HTTP response" if original_lookup is not None: with anyio.fail_after(5): await capture_ready.wait() assert captured == [first] print(json.dumps({"watcher_captured_first_server": True, "exit_after_first": sse.AppStatus.should_exit})) # The watcher has already captured server #1. Observe its shutdown broadcast. with anyio.fail_after(5): await sse.EventSourceResponse._listen_for_exit_signal() assert sse.AppStatus.should_exit else: print(json.dumps({"watcher_available": False, "exit_after_first": sse.AppStatus.should_exit})) if reset_between: sse.AppStatus.should_exit = False print(json.dumps({"diagnostic_reset": True})) await attempt("second") if __name__ == "__main__": parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("--reset-between", action="store_true") args = parser.parse_args() anyio.run(main, args.reset_between)
YS-OH-CORE commented
on Sep 12, 2026 More actionsI isolated the final-send ordering hypothesis on sse-starlette current main
6754ef387da97cf6cfbcd1bd5c216b533937b304, separately from the restart/global-state mechanism in the preceding comment.Probes, exact sources and candidate diff · Executed run34703822255
The trigger is explicit scheduling instrumentation: wait to deliver the first ordinary receive event until the final ASGI send begins, and delay that send by20ms. The minimal probe returns one
http.requestthen blocks. A second probe uses a real Uvicorn/h11 + HTTPX localhost connection and forwards real server receive events through the same barrier. No repeated empty-receive loop is used.Same probe Baseline Move active=Falseafter final awaited sendNormal response Complete Complete Delayed final send, normal request event Final send cancelled Final send completes Supplied disconnect event during final send Cancellation and close callback once Same behavior Instrumented localhost HTTP RemoteProtocolError; Uvicorn logs incomplete responseHTTP200, exact Korean SSE body The listener returns because
activeis already false;cancel_on_finishthen cancels the pending final send. Moving the completion flag after the send, inside the existing lock, fixes this selected interleaving without shielding cancellation. The original59 cases intest_sse.py,test_issue167.py, andtest_event.pypass unchanged on both variants.All eight probe executions use fresh processes. Both HTTP cases have
AppStatus.should_exit=Falsebefore stopping their only server, so this case does not require a stale global shutdown flag. Versions: sse-starlette3.4.11 source, AnyIO4.15.1, Starlette1.6.0, Uvicorn0.52.4, HTTPX0.28.1, hosted Linux/Python3.12.Scope: this demonstrates a separate final-send failure under controlled scheduling. It does not yet show that this interleaving explains the original fresh-process CPU-load failures or the reported restart frequency. No actual MCP initialize request is used. The SSE application payload may already be sent; the missing message here is the final HTTP completion. The supplied-disconnect control is an ASGI event test, not a physical socket-close test.
The downloaded artifact10301616324 contains actual observations, client/server logs, original-test XMLs and the one-ordering patch; SHA256
d6a1108f4287b546bd57c0adbcd7505c45763ee2749ad0acb95d40f58f17d50d. Original issue/hypothesis credit stays with jonpspri and the preceding restart analysis with skulitom. AI disclosure: prepared and executed with Zero (ChatGPT).
Follow-up, 2026-09-13 KST: this candidate does not repair the tested MCP restart failure.
I crossed the same unchanged patch with actual
mcp==2.2.0initialization over Uvicorn/h11 HTTP, now forwarding receives immediately and unchanged (only final send delayed20ms). Countercheck and exact conditions · Run34705469328.Baseline and candidate outcomes are identical: first initialization succeeds; after observing the real watcher carry server1's shutdown state forward, server2 fails with
RemoteProtocolError; resetting only that flag as a diagnostic restores initialization on both. The two restart failures occur before any final-send attempt, withAppStatus.should_exit=True. All ten requests consume the ordinary POST body before response start.This supports skulitom's separate restart explanation and limits what my earlier plain-SSE result establishes. The final-send ordering patch is not a solution for that restart path. The flag reset remains diagnostic, not a recommended fix. These six fresh-process scenarios do not reproduce the original CPU-contention frequency. Both failed requests remain failures in the evidence; the completed comparison is not an all-passing-initializations claim. I appended this here rather than creating another issue or repeating the original result.
Independent confirmation of the restart mechanism from the first comment, on the 1.x line: mcp 1.30.0 (and 1.28.1), sse-starlette 3.4.5, uvicorn 0.50.0, Python 3.12, Linux. Our test suite starts a real uvicorn server per test in one process. The first
initializePOST after an earlier server had stopped hung withincomplete chunked read, 4 of 6 runs on 1.28.1 and 6 of 6 on 1.30.0. At the hang,sse_starlette.sse.AppStatus.should_exitwasTrue.A standalone reproducer without MCP bisects the regression to sse-starlette 3.1.1 (PR sysid/sse-starlette#151). 3.1.0 and earlier are fine, and 3.1.1 through 3.4.11 fail on every run. I've filed it upstream as sysid/sse-starlette#211 with the script.
For test harnesses that start several servers in one process, sse-starlette's public
AppStatus.disable_automatic_graceful_drain()removes the latch. Called once before the first server starts, it made our suite pass deterministically (20 consecutive runs of the full suite and of the end-to-end file). We also have a regression test that forces the watcher to observe a shutdown and fails without it. This doesn't address the separate final-send race from the second comment, which does not need a restart.
Release line: 2.x (current stable) — also reproducible on the 1.x maintenance line, see below.
Description
On mcp 1.x and 2.x, a streamable-HTTP POST response can stop before the server sends the full SSE body. uvicorn writes
ASGI callable returned without completing responseand closes the connection. The client receiveshttpx.RemoteProtocolError: peer closed connection without sending complete message body. The failure occurs on the firstinitializePOST of a session. Session setup fails.We observed this behavior in GitHub Actions and in controlled reproduction environments.
The failure rate grows with two factors: CPU contention, and repeated uvicorn server start/stop cycles in one process.
Verified results, each with a fresh server per attempt:
In the 20-server test, the first two or three servers pass. All later servers fail. One reused server never fails. This pattern points at state that accumulates across server restarts on one event loop.
Expected behavior
The server must always send the last body chunk of the response. The client must always receive the complete stream for a request that has an answer.
Suspected cause (inference)
This section is an inference from code reading.
EventSourceResponseobject.active = Falsebefore it sends the last body chunk.receive()function in uvicorn returns immediately after it consumes the request body. This lets the disconnect task run in a loop.Example Code
Run it in a CPU-limited container:
docker run --rm --cpus=1 -it python:3.13-slim bash.pip install fastmcp==4.0.3 uvicorn==0.52.4 httpx anyio.Expected result: 20 of 20 attempts pass. Actual result: approximately 17 of 20 attempts fail from the third or fourth server onwards.
Python & MCP Python SDK
Python 3.13; mcp 2.2.0 (newest 2.x) and mcp 1.30.0 (newest 1.x); fastmcp 4.0.3 / 3.4.7; uvicorn 0.52.4 (h11); starlette 1.6.0; anyio 4.14.2; httpx 0.28.1; Linux (Ubuntu 24.04, Debian slim; x86_64 and arm64). The closest existing issues (#2150, #3441) describe different defects.