Skip to content

FastMCP/stdio: in-flight tool responses dropped on stdin EOF when input is bash-redirected from a file #2678

Description

@theone139344

1. Initial Checks

2. Description

Triage note: checked jlowin/fastmcp (now PrefectHQ fork) — FastMCP's server-side stdio delegates to mcp.server.stdio via a transport mixin, so the bug belongs here in the SDK rather than in the FastMCP wrapper.

When driving a FastMCP stdio server with a file-redirected stdin (e.g. python -m my_server < payload.jsonl > response.jsonl), in-flight tool-call responses can be dropped if their response writer hasn't been scheduled when stdin EOF arrives.

The stdio read loop appears to treat stdin EOF as an immediate-shutdown signal, cancelling pending writer tasks before they flush their JSON-RPC responses to stdout. The failure is silent — no traceback, no log line, the response is simply absent from stdout.

Expected: All responses for processed requests appear on stdout before the server exits.
Actual: Responses for the last-issued requests can be missing entirely.

3. Example Code (minimal reproducible)

# server.py
from mcp.server.fastmcp import FastMCP
import asyncio

mcp = FastMCP("repro")

@mcp.tool()
async def slow_echo(text: str) -> str:
    await asyncio.sleep(0.05)  # guaranteed yield point so writer scheduling is observable
    return text

if __name__ == "__main__":
    mcp.run(transport="stdio")
# payload.jsonl
{"jsonrpc":"2.0","id":0,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"repro","version":"0.1"}}}
{"jsonrpc":"2.0","method":"notifications/initialized","params":{}}
{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"slow_echo","arguments":{"text":"first"}}}
{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"slow_echo","arguments":{"text":"second"}}}
python server.py < payload.jsonl > response.jsonl
# Observed: response.jsonl contains id=0 + id=1 results, id=2 is absent.

The race is timing-sensitive; if id=2 surfaces on the first run, increase the asyncio.sleep delay or run repeatedly. We saw it deterministically in production where the tool body did a real HTTP call (~50-200ms).

Diagnostic fingerprint: the difference between this transport race and a quality bug in the tool itself is absence-of-response vs response-with-empty-result. If you ever see id=N missing entirely (not {"id":N,"result":[]}), suspect this race.

4. Python and MCP Python SDK

  • Python: 3.12 (python:3.12-slim Docker base)
  • MCP SDK: 1.27.x (mcp>=1.27,<2)
  • OS: Linux (Debian Bookworm; reproduced on Synology DSM 7.2 host)
  • Transport: stdio via mcp.run(transport="stdio")

Workaround we shipped (in case it helps the fix design): wrote a Python driver that owns both pipes via subprocess.Popen(stdin=PIPE, stdout=PIPE) and refuses to close stdin until the response for the last-issued id is observed on stdout. ~140 lines stdlib-only.

Suggested fix: before exiting on stdin EOF, await any pending writer tasks (with a small timeout) so their JSON-RPC responses reach stdout before the process terminates.

Activity

  1. Asti1982 commented on May 26, 2026

    @Asti1982

    I think the sharp regression boundary here is: once a JSON-RPC request has been read and accepted, stdin EOF should not be able to cancel the write path before that request's response is flushed.

    A focused test could drive stdio with redirected stdin using:

    • initialize
    • notifications/initialized
    • two tools/call requests where the second tool awaits at least once
    • captured stdout assertions that response ids 0, 1, and 2 all appear before process exit

    A narrow fix shape would be to let stdin EOF close the read side, but keep the stdout writer alive long enough to drain already-queued responses, with a bounded shutdown window so the server cannot hang forever.

    The important part is that a missing id=N response is materially different from an error response: clients and replay harnesses may retry without knowing whether the accepted tool call already ran. I can turn this into a focused fix/test PR if maintainers want it.

  2. added
    triageQueued for automated analysis — bot will process and remove this label
    on May 31, 2026
  3. mcp-claude commented on Jun 1, 2026

    @mcp-claude

    reproduces on main (616476f) and v1.x (6213787): with file-redirected stdin, responses for tool calls that are still mid-await when stdin EOFs are silently dropped — stderr shows the handlers were dispatched, but stdout only contains the initialize response.

    root cause is two-layer. BaseSession._receive_loop (shared/session.py:352) wraps read and write streams in one async with, so read-stream close immediately closes the write stream. then Server.run's finally at server/lowlevel/server.py:410-415 does tg.cancel_scope.cancel(), cancelling in-flight handler tasks before they reach message.respond(). the cancellation path at _handle_request (server.py:503-511) re-raises with no response, and even if a handler raced past it, the write stream is already closed.

    a narrow "wait briefly before cancelling" tweak in Server.run isn't enough on its own because the write stream has already closed by the session layer. fix needs to decouple read/write stream lifecycle in BaseSession._receive_loop and adjust the cancel semantics in Server.run without regressing the streamable_http case that the existing comment was written for. flagging as non-trivial for the fix stage.

    repro.py
    """Repro for issue 2678: stdin EOF cancels in-flight tool calls before responses flush."""
    
    import json
    import subprocess
    import sys
    import tempfile
    from pathlib import Path
    
    SERVER_SRC = '''
    from mcp.server.mcpserver import MCPServer
    import asyncio
    
    mcp = MCPServer("repro")
    
    @mcp.tool()
    async def slow_echo(text: str) -> str:
        await asyncio.sleep(0.2)
        return text
    
    if __name__ == "__main__":
        mcp.run(transport="stdio")
    '''
    
    PAYLOAD = [
        {"jsonrpc": "2.0", "id": 0, "method": "initialize",
         "params": {"protocolVersion": "2024-11-05", "capabilities": {},
                    "clientInfo": {"name": "repro", "version": "0.1"}}},
        {"jsonrpc": "2.0", "method": "notifications/initialized", "params": {}},
        {"jsonrpc": "2.0", "id": 1, "method": "tools/call",
         "params": {"name": "slow_echo", "arguments": {"text": "first"}}},
        {"jsonrpc": "2.0", "id": 2, "method": "tools/call",
         "params": {"name": "slow_echo", "arguments": {"text": "second"}}},
    ]
    
    
    def main():
        workdir = Path(tempfile.mkdtemp(prefix="mcp-2678-"))
        server_py = workdir / "server.py"
        payload_jsonl = workdir / "payload.jsonl"
        response_jsonl = workdir / "response.jsonl"
    
        server_py.write_text(SERVER_SRC)
        payload_jsonl.write_text("\n".join(json.dumps(m) for m in PAYLOAD) + "\n")
    
        with payload_jsonl.open("rb") as stdin, response_jsonl.open("wb") as stdout:
            proc = subprocess.run(
                [sys.executable, str(server_py)],
                stdin=stdin, stdout=stdout, stderr=subprocess.PIPE,
                timeout=20,
            )
    
        lines = [l for l in response_jsonl.read_text().splitlines() if l.strip()]
        received_ids = []
        for line in lines:
            try:
                msg = json.loads(line)
                if "id" in msg:
                    received_ids.append(msg["id"])
            except json.JSONDecodeError:
                pass
    
        print(f"server exit code: {proc.returncode}")
        print(f"received ids: {received_ids}")
        print(f"missing ids: {sorted(set([0,1,2]) - set(received_ids))}")
        print(f"stderr tail: {proc.stderr.decode(errors='replace')[-300:]}")
    
    
    if __name__ == "__main__":
        main()

    (on v1.x, swap from mcp.server.mcpserver import MCPServer → from mcp.server.fastmcp import FastMCP and MCPServer( → FastMCP(.)

    command + output
    $ uv run python repro.py
    server exit code: 0
    received ids: [0]
    missing ids: [1, 2]
    stderr tail:           INFO     Processing request of type            server.py:449
                                 CallToolRequest
                        INFO     Processing request of type            server.py:449
                                 CallToolRequest
    

    both CallToolRequest log lines confirm the handlers were entered; their responses never reached stdout.

    code path

    src/mcp/server/stdio.py:49-62 — stdin_reader exits its async for line in stdin on EOF; async with read_stream_writer exit closes the read stream.

    src/mcp/shared/session.py:352 — async with self._read_stream, self._write_stream: exit closes both streams together when the receive loop sees the read-stream EndOfStream.

    src/mcp/server/lowlevel/server.py:392-415 — async for message in session.incoming_messages exits, and the finally: tg.cancel_scope.cancel() cancels every still-running _handle_message task.

    src/mcp/server/lowlevel/server.py:503-511 — _handle_request catches the cancellation. message.cancelled is false (no client-side notifications/cancelled), so the cancel re-raises and no response is sent. even removing the cancel_scope.cancel() wouldn't help on its own because by that point _write_stream is already closed at session.py:352.

  4. added
    bugSomething isn't working
    ready for workEnough information for someone to start working on
    P2Moderate issues affecting some users, edge cases, potentially valuable feature
    and removed
    triageQueued for automated analysis — bot will process and remove this label
    on Jun 1, 2026
  5. norika1207-lab commented on Jun 2, 2026

    @norika1207-lab
  6. TimeToBuildBob commented on Jul 10, 2026

    @TimeToBuildBob
  7. rudidev08 commented on Aug 10, 2026

    @rudidev08

    Still reproduces on v2.0.0.

    Repro: a minimal MCPServer with one tool, fed initialize + notifications/initialized + tools/list on stdin with immediate EOF (Popen.communicate). Over 5 rounds, the tools/list response (id 2) was missing from stdout in 1 round. The race is timing-sensitive; a short delay before the server imports (time.sleep(0.5)) makes it easier to hit.

    # /// script
    # requires-python = ">=3.11"
    # dependencies = ["mcp==2.0.0"]
    # ///
    import time
    time.sleep(0.5)
    from mcp.server.mcpserver import MCPServer
    mcp = MCPServer("probe")
    @mcp.tool()
    def ping() -> str:
        """ping"""
        return "pong"
    mcp.run()

    Driver: pipe these three lines to the script's stdin and close it, then check stdout for "id":2:

    {"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {"protocolVersion": "2025-11-25", "capabilities": {}, "clientInfo": {"name": "probe", "version": "1.0"}}}
    {"jsonrpc": "2.0", "method": "notifications/initialized"}
    {"jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {}}

    Environment: macOS, Python 3.11+, mcp 2.0.0 in a fresh uv environment.

  8. ace-trump-tech commented on Aug 31, 2026

    @ace-trump-tech

    1. Initial Checks

    2. Description

    Triage note: checked jlowin/fastmcp (now PrefectHQ fork) — FastMCP's server-side stdio delegates to mcp.server.stdio via a transport mixin, so the bug belongs here in the SDK rather than in the FastMCP wrapper.

    When driving a FastMCP stdio server with a file-redirected stdin (e.g. python -m my_server < payload.jsonl > response.jsonl), in-flight tool-call responses can be dropped if their response writer hasn't been scheduled when stdin EOF arrives.

    The stdio read loop appears to treat stdin EOF as an immediate-shutdown signal, cancelling pending writer tasks before they flush their JSON-RPC responses to stdout. The failure is silent — no traceback, no log line, the response is simply absent from stdout.

    Expected: All responses for processed requests appear on stdout before the server exits. Actual: Responses for the last-issued requests can be missing entirely.

    3. Example Code (minimal reproducible)

    server.py

    from mcp.server.fastmcp import FastMCP
    import asyncio

    mcp = FastMCP("repro")

    @mcp.tool()
    async def slow_echo(text: str) -> str:
    await asyncio.sleep(0.05) # guaranteed yield point so writer scheduling is observable
    return text

    if name == "main":
    mcp.run(transport="stdio")

    payload.jsonl

    {"jsonrpc":"2.0","id":0,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"repro","version":"0.1"}}}
    {"jsonrpc":"2.0","method":"notifications/initialized","params":{}}
    {"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"slow_echo","arguments":{"text":"first"}}}
    {"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"slow_echo","arguments":{"text":"second"}}}
    python server.py < payload.jsonl > response.jsonl

    Observed: response.jsonl contains id=0 + id=1 results, id=2 is absent.

    The race is timing-sensitive; if id=2 surfaces on the first run, increase the asyncio.sleep delay or run repeatedly. We saw it deterministically in production where the tool body did a real HTTP call (~50-200ms).

    Diagnostic fingerprint: the difference between this transport race and a quality bug in the tool itself is absence-of-response vs response-with-empty-result. If you ever see id=N missing entirely (not {"id":N,"result":[]}), suspect this race.

    4. Python and MCP Python SDK

    • Python: 3.12 (python:3.12-slim Docker base)
    • MCP SDK: 1.27.x (mcp>=1.27,<2)
    • OS: Linux (Debian Bookworm; reproduced on Synology DSM 7.2 host)
    • Transport: stdio via mcp.run(transport="stdio")

    Workaround we shipped (in case it helps the fix design): wrote a Python driver that owns both pipes via subprocess.Popen(stdin=PIPE, stdout=PIPE) and refuses to close stdin until the response for the last-issued id is observed on stdout. ~140 lines stdlib-only.

    Suggested fix: before exiting on stdin EOF, await any pending writer tasks (with a small timeout) so their JSON-RPC responses reach stdout before the process terminates.

    I reproduced this issue on the current main branch.

    The problem is not the tool result itself. When stdin reaches EOF, the JSON-RPC dispatcher immediately enters shutdown and cancels request handlers that have already been accepted but are still waiting to produce their responses. This can silently drop the final tool responses.

    My proposed fix is to add an opt-in, bounded graceful-shutdown drain:

    # JSONRPCDispatcher.run(...)
    if graceful_shutdown_timeout and self._active_requests:
        with anyio.move_on_after(graceful_shutdown_timeout):
            await self._wait_for_active_requests()
    
    # Requests are tracked until their response write has completed.
    async def _run_request(..., completion: anyio.Event) -> None:
        try:
            await self._handle_request(...)
        finally:
            self._active_requests.discard(completion)
            completion.set()
    The stdio entry point enables the drain window:
    await self._lowlevel_server.run(
        read_stream,
        write_stream,
        self._lowlevel_server.create_initialization_options(),
        graceful_shutdown_timeout=0.5,
    )
    The important design choice is that the default timeout remains 0, so existing behavior for non-stdio transports is unchanged. Only stdio sessions that receive redirected JSONL input get a bounded opportunity to flush accepted responses before handlers are cancelled.
    I also added a regression test that:
    1. starts a slow request;
    2. closes the input stream immediately;
    3. verifies that the response is still emitted;
    4. keeps the timeout bounded so a broken handler cannot hang shutdown.
    The implementation is available here:
    https://github.com/ace-trump-tech/python-sdk/tree/fix/stdio-eof-drain
  9. BoltyBolterson commented on Sep 11, 2026

    @BoltyBolterson

    Adding measurements from a slightly different angle, because two things were not obvious from the repros above: the tool does not need an await to lose its response (a plain synchronous echo loses too), and the loss rate is roughly half of the outstanding calls from two calls upward. It also still reproduces on 2.2.0.

    Environment: Linux (Ubuntu 26.04 on WSL2, kernel 6.6), Python 3.14.4, fresh venvs with mcp==1.27.2 and mcp==2.2.0.

    Repro: a one-tool server with a synchronous echo tool, and a client that writes initialize + notifications/initialized + N tools/call requests through subprocess.run(input=...) (so stdin closes as soon as the payload is written), then counts which request ids got a response on stdout.

    server.py
    try:
        from mcp.server.mcpserver import MCPServer as FastMCP  # mcp >= 2.0
    except ImportError:
        from mcp.server.fastmcp import FastMCP  # mcp 1.x
    
    mcp = FastMCP("echo")
    
    
    @mcp.tool()
    def echo(text: str) -> str:
        return text
    
    
    if __name__ == "__main__":
        mcp.run(transport="stdio")
    client.py
    """usage: python client.py server.py <calls> [runs]"""
    import json
    import subprocess
    import sys
    
    server, n = sys.argv[1], int(sys.argv[2])
    runs = int(sys.argv[3]) if len(sys.argv) > 3 else 20
    
    lossy_runs = lost = total = 0
    for _ in range(runs):
        msgs = [
            {"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {
                "protocolVersion": "2024-11-05", "capabilities": {},
                "clientInfo": {"name": "drain-probe", "version": "1"}}},
            {"jsonrpc": "2.0", "method": "notifications/initialized"},
        ] + [
            {"jsonrpc": "2.0", "id": 10 + i, "method": "tools/call",
             "params": {"name": "echo", "arguments": {"text": f"call {i}"}}}
            for i in range(n)
        ]
        wanted = {m["id"] for m in msgs if "id" in m}
        p = subprocess.run([sys.executable, server], text=True, capture_output=True,
                           input="\n".join(json.dumps(m) for m in msgs) + "\n")
        got = {f["id"] for f in map(json.loads, filter(str.strip, p.stdout.splitlines())) if "id" in f}
        missing = wanted - got
        total += len(wanted)
        lost += len(missing)
        lossy_runs += bool(missing)
        if p.returncode != 0 or p.stderr.strip():
            print(f"  exit={p.returncode} stderr={p.stderr.strip()[:200]!r}")
    
    print(f"calls={n:<3d} runs={runs}  runs_losing_replies={lossy_runs}/{runs}  replies_lost={lost}/{total}")

    Results (python client.py server.py N 20; "replies" = the initialize response plus one per tools/call):

    calls sent mcp 1.27.2: runs losing a reply replies lost mcp 2.2.0: runs losing a reply replies lost
    1 3/20 3/40 6/20 6/40
    2 16/20 16/60 19/20 19/60
    5 20/20 49/120 20/20 54/120
    10 20/20 104/220 20/20 103/220
    20 20/20 216/420 20/20 209/420

    The server exits 0 on every run and stdout is valid JSON-RPC, there are simply fewer frames. Nothing on stderr on 2.2.0; on 1.27.2 with Python 3.14 the only stderr output is an unrelated pydantic_settings warning.

    Waiting before closing stdin does not help; only holding stdin open until every id has been answered does (mcp 1.27.2, 5 calls, 30 runs per row, same payload written through Popen with a reader thread on stdout):

    close strategy runs losing a reply replies lost
    close stdin immediately after writing 30/30 66/180
    sleep 0.05 s, then close 30/30 66/180
    sleep 0.25 s, then close 30/30 67/180
    read until all ids answered, then close 0/30 0/180

    So the trigger is "stdin closed while requests are outstanding", not "closed too soon", which matches the cancel-on-EOF analysis above: the response for anything accepted but not yet written is dropped regardless of how long the client waits.

    For comparison, the same client pattern (write everything, close stdin, count replies) against an equivalent one-tool echo server built on the TypeScript SDK (@modelcontextprotocol/sdk) and on Rust rmcp 3.2.0 lost 0 replies at every batch size in the table above (measured earlier this month with the same payload shape, not re-run today).

  10. sairam0424 commented on Sep 14, 2026

    @sairam0424

    We independently reproduced this while auditing our own server (cost-guard-mcp, a local-stdio-only MCP server whose tool calls can run up to 120 seconds - warehouse EXPLAIN/dry-run calls against BigQuery/Snowflake/Databricks). Confirmed the response drop happens exactly as described here.

    For anyone else finding this: we could not find evidence either way on whether real, persistent MCP clients (Claude Desktop, Claude Code, Cursor) ever close stdin mid-tool-call during normal operation - every reproduction known to us (including our own) used batch/file-redirected stdin with an injected delay. So real-world exposure for well-behaved clients that keep stdin open for the session appears low, but this is still worth fixing given several long-running-tool-call servers (like ours) exist.

    +1 on getting one of the existing fix attempts (#2680/#2682/#2815/#2821/#3026/#3421) merged - happy to help test if useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Moderate issues affecting some users, edge cases, potentially valuable featurebugSomething isn't workingready for workEnough information for someone to start working on

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions