Repository navigation
Streamable HTTP clean EOF reconnects can exceed the request retry budget #3307
Description
Activity
- addedbugSomething isn't workingSomething isn't workingP2Moderate issues affecting some users, edge cases, potentially valuable featureModerate issues affecting some users, edge cases, potentially valuable featureneeds confirmationNeeds confirmation that the PR is actually required or needed.Needs confirmation that the PR is actually required or needed.
on Aug 14, 2026 I plan to prepare a small regression fix for this issue: count a clean EOF without a JSON-RPC response toward the request-scoped reconnection budget, with a focused transport test. I will keep the change limited to the existing reconnection path.
I'd like to work on this — I've reproduced it and have a fix + regression tests ready. Could a maintainer assign this to me?
AI disclosure: this investigation and fix were prepared with AI assistance (Claude Code). Posting here first, before opening a PR, so a maintainer can weigh in given the issue is currently labeled
needs confirmation.Root cause: confirmed on current
main(src/mcp/client/streamable_http.py,_handle_reconnection). The clean-EOF branch always recurses withattempt=0:# Stream ended again without response - reconnect again (reset attempt counter) await self._handle_reconnection(ctx, reconnect_last_event_id, reconnect_retry_ms, 0)
while the exception branch correctly increments (
attempt + 1). A server that keeps reopening the resumable stream, sends only a bare id-bearing priming event, and closes again without ever producing a response can therefore reconnect forever instead of giving up afterMAX_RECONNECTION_ATTEMPTS.Fix: track whether a reconnect actually delivered a real event (any event with non-empty
data— a notification, response, or error) before its EOF. A reconnect that made real progress still earns a fresh budget for the next one; a reconnect that saw nothing but a bare priming event increments the counter, same as the exception path.The first version I tried was the simpler "always increment on clean EOF" (matching the issue's literal wording), but that broke
tests/shared/test_streamable_http.py::test_streamable_http_multiple_reconnections, which relies on the budget resetting across reconnects that each deliver a real notification (a legitimate long-lived tool call with 3 intentional stream closes). The progress-aware version above keeps that test passing unchanged while fixing the no-progress infinite-reconnect case from this issue.Verification:
- New regression test in
tests/client/test_streamable_http.py: drives_handle_reconnectiondirectly with a mock transport that returns a bare priming-event-then-EOF response on every reconnect; fails (hangs/times out) on currentmain, passes after the fix (gives up after exactlyMAX_RECONNECTION_ATTEMPTSreconnects and resolves the waiter withCONNECTION_CLOSED). uv run pytest tests/client/test_streamable_http.py— 29 passeduv run pytest tests/shared/test_streamable_http.py::test_streamable_http_multiple_reconnections— passed (no regression on the multi-reconnect-with-progress scenario)uv run pytest(full suite) — 5621 passed, 16 skipped, 1 xfailed; the only failure (test_safe_join_rejects_symlink_escape) is a pre-existing Windows-privilege limitation for symlink creation, confirmed identical onmainbefore this changeuv run ruff check ./uv run ruff format --check ./uv run pyright— clean on the changed files
Happy to open the PR once assigned, or share the diff here first if that's preferred.
- New regression test in
Submitted a fix: #3323
- added a commit that references this issue
on Aug 17, 2026 - addedv1Affects the v1.x maintenance lineAffects the v1.x maintenance linev2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)
on Aug 18, 2026 I'd like to take on this issue.
Proposed Solution:
- In
_convert_to_content(result)withinsrc/mcp/server/fastmcp/utilities/func_metadata.py, detect when an empty sequence (listortuple) is returned and serialize it to[types.TextContent(type="text", text="[]")](or"()"). - Preserve existing
Contentinstances (e.g.TextContent,ImageContent,EmbeddedResource) and handle nested collections without destructive flattening. - Add regression tests in
tests/server/fastmcp/test_tools.pycovering empty lists, empty tuples, nested empty lists, and non-empty sequences.
Could you please assign this to me? Thanks!
- In
I'd like to work on this. I have a fix ready — the reconnection loop passes the initial attempt count (0) on every retry instead of incrementing it, so the budget is never consumed.
Kludex commented
on Oct 10, 2026 MemberMore actionsBoth reports identify
_handle_reconnectionresetting its attempt count after a clean SSE EOF without a response, allowing unbounded request retries. This is tracked in #2393, so I’m closing this as a duplicate. AI-assisted triage; I reviewed both reports.
Description
The streamable HTTP client can exceed its per-request SSE reconnection budget when each reconnect opens successfully, emits only an id-bearing priming event, and then reaches EOF without a JSON-RPC response.
In that case
_handle_reconnection()currently recurses withattempt=0after the clean EOF path. The exception path increments the counter, but the normal EOF-without-response path resets it, so a no-timeout request such assubscriptions/listencan keep reconnecting instead of resolving the waiter withCONNECTION_CLOSEDafterMAX_RECONNECTION_ATTEMPTS.Reproduction
Drive
StreamableHTTPTransport._handle_reconnection()with a mock HTTP transport that returns these per-request SSE responses:id: evt-1with emptydata, then EOF.id: evt-2with emptydata, then EOF.Starting from
Last-Event-ID: evt-0andretry_interval_ms=0, currentmainmakes the third HTTP request and delivers the success response. I expected the client to stop after the two configured reconnect attempts, emit aJSONRPCErrorfor the original request withCONNECTION_CLOSED, and only sendLast-Event-ID: evt-0andLast-Event-ID: evt-1.Expected Behavior
Each per-request reconnect that reaches EOF without delivering a JSON-RPC response should consume the reconnect budget. That keeps request-scoped SSE drops consistent whether they end by transport exception or by clean EOF, and prevents no-timeout callers from staying parked forever when a server repeatedly closes resumable streams without producing the response.
Local Verification
I have a local regression test that fails on current
mainbecause the third reconnect response is accepted, then passes when the clean EOF path recurses withattempt + 1.Commands run locally:
uv run --frozen pytest tests/client/test_streamable_http.py::test_empty_resumable_sse_reconnects_count_toward_the_request_budget -quv run --frozen pytest tests/client/test_streamable_http.py -quv run --frozen ruff check src/mcp/client/streamable_http.py tests/client/test_streamable_http.pyuv run --frozen ruff format --check src/mcp/client/streamable_http.py tests/client/test_streamable_http.pyuv run --frozen pyright src/mcp/client/streamable_http.py tests/client/test_streamable_http.pyUV_FROZEN=1 uv run --frozen strict-no-coverDisclosure: I used AI assistance to help prepare this report and a local patch; I reviewed the reproduction, root cause, and test results.