Repository navigation
[v2] Expose the SSE max_event_size setting in Streamable HTTP clients #3332
Description
Activity
We hit exactly this in production, and it is not limited to
tools/call: atools/listresponse above 1 MiB is enough to make a whole upstream server disappear behind a Streamable HTTP client.Real-world occurrence
Our MCP gateway (FastMCP 4.0.0, which pulls
mcp==2.1.1andhttpx2==2.12.0) proxies a hosted MCP server (Composio,backend.composio.dev) that exposes 1030 tools. That server answerstools/listwith HTTP 200text/event-streamand delivers the whole result as ONEdata:event of 2,140,787 bytes. Verified with a raw replay of the identical request bodies the SDK sends (with and withoutparams._meta, protocol version 2025-11-25): the server returns the complete, valid response every time.With
mcp1.x (FastMCP 3.4.5,httpx0.28 +httpx-sse) the client lists all 1030 tools. Withmcp2.1.1 the same request ends inMCPError: SSE stream ended without a responseand a proxy built on top of it reports 0 tools for that upstream, without any hint that a size limit was involved. The actual
SSEError("Server-sent event exceeded the 1048576 byte limit.")is only visible at DEBUG level.Where it happens (mcp 2.1.1)
src/mcp/client/streamable_http.py:426:event_source = EventSource(response)- nomax_event_size, sohttpx2._config.DEFAULT_MAX_EVENT_SIZE_BYTES = 1024 * 1024applies.src/mcp/client/streamable_http.py:411-460(_handle_sse_response): theSSEErrorraised by the httpx2 decoder is caught byexcept Exception: logger.debug("SSE stream ended", exc_info=True)and the request is resolved via_resolve_abandoned_request(..., "SSE stream ended without a response"), so the caller gets a genericCONNECTION_CLOSEDerror that points in the wrong direction (network) instead of the real cause (client-side size cap).- Neither
streamable_http_client()norStreamableHTTPTransportexpose the setting, and FastMCP (checked 4.0.0 and 4.0.1, packagefastmcp-slim) has no way to pass it through either.
Minimal reproduction (no external services)
A plain Starlette app that answers every JSON-RPC request with
text/event-streamcontaining exactly onedata:event, like several hosted servers do (stateless, nomcp-session-id). Run with:uv run --no-project --with "mcp==2.1.1" --with uvicorn --with starlette python repro_sse_event_limit.pyimport asyncio, json, socket, threading import uvicorn from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response, StreamingResponse from starlette.routing import Route from mcp import ClientSession from mcp.client.streamable_http import streamable_http_client def make_app(total_kib: int): n = 10 desc = "x" * (total_kib * 1024 // n) tools = [{"name": f"tool_{i}", "description": desc, "inputSchema": {"type": "object", "properties": {}}} for i in range(n)] async def endpoint(request: Request): body = await request.json() method = body.get("method") if method == "initialize": result = {"protocolVersion": "2025-11-25", "capabilities": {"tools": {"listChanged": True}}, "serverInfo": {"name": "one-big-sse-event", "version": "0.1.0"}} elif method == "tools/list": result = {"tools": tools} elif "id" not in body: # notifications return Response(status_code=202) else: # e.g. server/discover from 2.x clients return Response(json.dumps({"jsonrpc": "2.0", "id": body["id"], "error": {"code": -32601, "message": "Method not found"}}), status_code=400, media_type="application/json") payload = json.dumps({"jsonrpc": "2.0", "id": body["id"], "result": result}) async def gen(): yield f"event: message\ndata: {payload}\n\n".encode() return StreamingResponse(gen(), media_type="text/event-stream", headers={"cache-control": "no-cache"}) return Starlette(routes=[Route("/mcp", endpoint, methods=["POST", "GET", "DELETE"])]) def free_port(): s = socket.socket(); s.bind(("127.0.0.1", 0)); p = s.getsockname()[1]; s.close(); return p async def probe(total_kib: int) -> str: port = free_port() server = uvicorn.Server(uvicorn.Config(make_app(total_kib), host="127.0.0.1", port=port, log_level="error")) th = threading.Thread(target=server.run, daemon=True); th.start() while not server.started: await asyncio.sleep(0.05) try: async with streamable_http_client(f"http://127.0.0.1:{port}/mcp") as streams: async with ClientSession(streams[0], streams[1]) as session: await session.initialize() result = await session.list_tools() out = f"OK, {len(result.tools)} tools" except Exception as exc: inner = exc while isinstance(inner, BaseExceptionGroup) and inner.exceptions: inner = inner.exceptions[0] out = f"FAIL: {type(inner).__name__}: {inner}" server.should_exit = True; th.join(timeout=5) return out async def main(): for kib in (512, 1000, 1023, 1030, 1100, 2200): print(f"tools/list as ONE SSE event of ~{kib:>5} KiB -> {await probe(kib)}", flush=True) asyncio.run(main())
Output (macOS, Python 3.12,
mcp2.1.1,httpx22.12.0):tools/list as ONE SSE event of ~ 512 KiB -> OK, 10 tools tools/list as ONE SSE event of ~ 1000 KiB -> OK, 10 tools tools/list as ONE SSE event of ~ 1023 KiB -> OK, 10 tools tools/list as ONE SSE event of ~ 1030 KiB -> FAIL: MCPError: SSE stream ended without a response tools/list as ONE SSE event of ~ 1100 KiB -> FAIL: MCPError: SSE stream ended without a response tools/list as ONE SSE event of ~ 2200 KiB -> FAIL: MCPError: SSE stream ended without a responseThe cut is exactly at the 1 MiB default. Note that a FastMCP 2.x/4.x test server does NOT reproduce this because it answers with
application/json; the trigger is a server that streams the result as a single SSE event, which the spec allows.What would help
- Expose the limit: a
max_event_sizeparameter onstreamable_http_client()andStreamableHTTPTransport(httpx2 terminology), applied consistently to the POST response path (EventSource(response, max_event_size=...)), the GET stream and the reconnection path (AsyncClient.sse(...)currently has no transport-level setting either). - Surface the cause: when the decoder raises
SSEError, resolve the pending request with an error that carries theSSEErrortext (and log at WARNING, not DEBUG). "SSE stream ended without a response" sends people hunting for network problems; it took a wire trace to find the size cap. - Consider a larger default than 1 MiB for the POST response path, since
tools/listresults of large hosted servers routinely exceed it. The bound itself is valuable (it protects the client from a misbehaving server), it just needs to be configurable.
Workaround we use in the meantime
Process-local, at import time, before any client connects:
import functools, httpx2 import mcp.client.streamable_http as sh sh.EventSource = functools.partial(httpx2.EventSource, max_event_size=32 * 1024 * 1024)
That keeps a bound in place (memory protection against a broken upstream) while letting the real-world
tools/listthrough. Patchinghttpx2itself does not work because the default is bound at function definition time.Happy to test a PR against the real server that triggered this.
Reacted by Antoine Balliet, Thomas Delayen and Charles FrancisThe idea of the limit in HTTPX2 was to prevent an attack, but we can change that limit if proven that a bigger value doesn't bring back a vulnerability.
Thanks — I agree that the safe default should remain unchanged.
My proposal is not to globally increase or remove the limit, but to expose
max_event_sizeas an explicit client option for applications that legitimately receive larger SSE events. The existing value would remain the default, so current users retain the same protection.We can also document that increasing it raises memory/DoS risk and should only be done when the server and payload sizes are trusted or otherwise controlled. Would that address the security concern?
@Zhangs-11 are you openclaw or hermes?
Neither — I’m using OpenAI Codex to help investigate and communicate the issue.
The limit protects a client from a hostile or broken server exhausting its memory, and that protection should stay the default. Exposing
max_event_sizeas an explicit, opt-in parameter onstreamable_http_client()andStreamableHTTPTransportdoes not reintroduce the risk: the default path is untouched, and only an operator who trusts one specific server raises the bound for that one client, the same trust decision they already make withtimeout. In our gateway we run exactly that today via a workaround: a 32 MiB bound for one hosted upstream whosetools/listis a single 2.1 MB event (1030 tools), everything else keeps the 1 MiB default.Independent of the limit itself: when the decoder raises
SSEError, resolving the pending request with that message instead of the generic "SSE stream ended without a response" has no security cost and would have saved us a wire trace to find the cause.We hit this exact failure in production through FastMCP 4 / MCP SDK 2 and httpx2 2.12.0. Two successful HTTP 200 tool responses were approximately 1.83 MB and 2.15 MB, so each deterministically exceeded httpx2's 1,048,576-byte per-event limit. The caller only received
SSE stream ended without a response.Our downstream compatibility fix deliberately keeps the 1 MiB safety cap. It catches
SSEErrorbefore the SDK's catch-all, logs the exception type/message withexc_info, resolves only the affected request with an actionable message, and raises the caller-facing error from the originalSSEError. A separate regression confirms a plain premature EOF still receives the existing generic error.The key upstream requirement for us is therefore request-scoped exception propagation, not
max_event_size=None: the original parser exception (or an equivalent structured error retaining its details) should reach the call boundary, siblings should remain usable, and the default cap should stay intact.I would be happy to contribute the smaller upstream change once maintainers confirm the preferred exception-plumbing shape, especially how the original cause should cross the stream/dispatcher boundary.
What happened?
MCP Python SDK v2.0.0 constructs
httpx2.EventSource(response)directly when parsing a Streamable HTTP POST response. Since HTTPX2 2.10, a single SSE event is limited to 1 MiB by default.When a valid
tools/callresult is larger than 1 MiB and is returned as one SSE event, HTTPX2 raisesSSEError. The SDK catches that transport error and the caller receives the generic MCP error:There is currently no public Streamable HTTP setting that lets callers raise the SSE event-size limit. The transport also creates SSE readers in multiple places: the POST response path constructs
EventSource(response)directly, while the GET and reconnection paths callAsyncClient.sse()without a transport-level event-size setting.What did you expect?
I expected the Streamable HTTP client/transport to expose an SSE event-size setting and apply it consistently to:
One possible API shape would be a
max_event_sizeargument onstreamable_http_client()andStreamableHTTPTransport, matching HTTPX2 terminology. The exact public API is open for maintainer direction.If an event exceeds the configured limit, the original JSON-RPC request should receive a clear request-scoped error. The already-sent POST should not be replayed, and sibling requests sharing the session should remain usable.
Reproduction
Use an MCP Streamable HTTP server whose tool returns more than 1 MiB of text in a single SSE event, then call that tool with the v2 client:
With a 2 MiB single-event response, the call fails with
SSE stream ended without a response. The same response succeeds when the underlyingEventSourceis constructed with a largermax_event_size.I can contribute an implementation and exact boundary tests if maintainers agree with exposing this setting.
Environment
Reference
max_event_sizeto cap SSE event buffering pydantic/httpx2#1071This issue was prepared with AI assistance and reviewed by the reporter.