Repository navigation
Streamable HTTP client treats 404 as terminal instead of re-initializing, as the spec requires #3556
Description
Activity
- addedv2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)v1Affects the v1.x maintenance lineAffects the v1.x maintenance line
on Sep 21, 2026 I was able to reproduce this on
main(f1b65890).For the repro, I used two in-process
StreamableHTTPSessionManagerinstances and switched the client's HTTP transport from the first one to the second, basically simulating what the client sees after a server restart.After the switch, the first request gets a
404with:{"error": {"code": -32600, "message": "Session not found"}}The client surfaces this as an
MCPError, but it never sends a newinitializerequest. Because of that, every request after that continues using the stale session and fails in the same way.This isn't limited to deployments/restarts either.
StreamableHTTPSessionManagerremoves idle sessions aftersession_idle_timeout, which defaults to 30 minutes. Its docstring also mentions that once the session is removed, the server responds with a 404 and the client is expected to initialize a new session. So an SDK client that simply stays idle for 30 minutes against an SDK server can hit the same issue.I have a working fix on a branch, limited to
src/mcp/client/streamable_http.py, and I'd be happy to open a PR if you're okay with assigning this to me.The approach I took is:
- The transport already observes the original
initializerequest andnotifications/initialized, so it keeps enough information to replay the handshake. - If a request with a session ID gets a
404, the transport retries the handshake without the stale session ID or protocol-version header, stores the newly returned session ID, and retries the original request once. - If the retry also gets a
404, or re-initialization itself fails, the error is surfaced normally. - Concurrent requests that hit the stale session at the same time share a single re-initialization using a lock and a check to see whether another request has already replaced the session ID.
- The GET stream is tied to the session it was created for. Once the session changes, the old stream stops reconnecting and a new stream is created for the new session.
- The replayed
initializeresult is not forwarded back to the session, so the client keeps the capabilities/protocol state it negotiated originally.
That last point is the main thing I'd like feedback on. If the restarted server reports different capabilities, the client won't notice them. I went with this because it seems closest to the "re-initialize transparently" behavior described in the issue, but I'm happy to adjust that if there's a better expected behavior.
The recovery only happens when the transport performed the handshake itself. If the session ID was provided through caller-supplied headers, behavior stays unchanged.
tests/shared/test_streamable_http.py::test_streamable_http_client_session_terminationstill passes for that case.I also added five tests in
tests/client/test_streamable_http.pycovering:- recovery after a real server restart
- concurrent requests hitting an expired session
- a second
404after recovery - failed re-initialization
- GET stream handoff to the new session
All five fail without the change. The full test suite passes with the fix, and coverage for the changed source is 100%.
- The transport already observes the original
I independently reproduced the stale-session 404 path on the two published SDK lines I could test:
mcp 1.28.1/ Python 3.11.15mcp 2.2.0/ Python 3.14.7
I used a deterministic transport-level fixture: seed
Mcp-Session-Id: stale-session-xbstack, make two consecutivetools/listrequests receive HTTP 404, then record the client state and outbound requests.Both versions behaved the same way:
request 1 -> Session terminated request 2 -> Session terminated session id after both 404s -> stale-session-xbstack outbound methods -> tools/list, tools/list automatic initialize -> noneSo this also reproduces on the published 1.x and 2.x versions, not just the current main-branch reproduction above. The fixture intentionally isolates the transport's 404 recovery branch; it does not attempt to model every deploy/proxy/idle-timeout condition.
Repro + raw logs + version matrix:
https://github.com/xbstack/mcp-streamable-http-lifecycle-regressionsRelated XBSTACK deployment guide with the tested boundary and recovery notes:
https://www.xbstack.com/en/ai/mcp-streamable-http-deployment/?utm_source=github&utm_medium=referral&utm_campaign=mcp_stale_session_404_reinitialize&utm_content=issue_comment&ref=githubKludex commented
on Oct 10, 2026 MemberMore actionsBoth reports describe the streamable HTTP client treating a session-not-found
404as terminal instead of starting a new session withInitializeRequest. This is tracked in #1676, so I’m closing this as a duplicate. AI-assisted triage; I reviewed both reports.
Summary
The 2025-06-18 Streamable HTTP transport specification says that a client which receives
HTTP 404in response to a request carrying anMcp-Session-Idheader MUST start a newsession by sending a new
InitializeRequestwithout a session id.StreamableHTTPTransportnever does this. Insrc/mcp/client/streamable_http.pya 404 isturned into a terminal error and the transport is left unusable:
_send_session_terminated_erroremitsJSONRPCError(code=32600, message="Session terminated")and returns. There is no re-initialization path anywhere in the transport: the stored
session id is never dropped, no new
InitializeRequestis sent, and the failed request isnever retried.
Why this matters
A server that restarts — an ordinary deploy — legitimately answers 404 for every session
id it no longer knows. Per the specification that is the correct server behaviour, and the
client is supposed to recover transparently. Because it does not, every connected client
is permanently broken by any server restart until a human reconnects it.
We hit this twice in one day on routine deploys of a hosted MCP server. From the server
side it is not fixable: adopting an unknown session id would mean fabricating a handshake,
because
ServerSessionstarts inNotInitializedand the first request then raisesReceived request before initialization was complete(src/mcp/shared/session.py), afterwhich the session manager tears the session down again.
Expected behaviour
On a 404 for a request that carried a session id: discard the stored session id,
re-initialize transparently, and retry the request once. A second 404 on the retry can
reasonably remain terminal.
Actual behaviour
The request fails with
Session terminatedand the transport stays dead for the rest ofthe client's lifetime.
Version
mcp1.28.1, Python 3.11.