Skip to content

OAuthClientProvider auth lock is permanently poisoned when httpx closes async_auth_flow from a different task (RuntimeError: The current task is not holding this lock) #3382

Description

@flanker-stack

Summary

OAuthClientProvider.async_auth_flow holds self.context.lock (an anyio.Lock) across the entire httpx auth-flow generator, including every yield (src/mcp/client/auth/oauth2.py:582 in v2.0.0; same code is present on v2.1.0 and main). anyio.Lock.release() is bound to the acquiring task. When httpx closes the auth generator from a different task than the one that advanced it — which happens routinely when a request is cancelled mid-flight (network drop, timeout, task-group teardown) — the async with __aexit__ runs in the closing task and raises:

RuntimeError: The current task is not holding this lock

The exception escapes into a fire-and-forget teardown task ("Task exception was never retrieved"), and the lock is left permanently held. Every subsequent request through the same OAuthClientProvider then blocks forever at async with self.context.lock: without sending any HTTP. For a long-lived client that reuses the provider across reconnects, that server is dead until the whole process restarts.

Environment

Observed traceback

Two captures from the same host, one per yield point (authorization-code exchange and the main yield request):

ERROR asyncio: Task exception was never retrieved
future: <Task finished name='Task-598' coro=<<async_generator_athrow without __name__>()> exception=RuntimeError('The current task is not holding this lock')>
Traceback (most recent call last):
  File ".../site-packages/mcp/client/auth/oauth2.py", line 754, in async_auth_flow
    yield request
GeneratorExit

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File ".../site-packages/mcp/client/auth/oauth2.py", line 582, in async_auth_flow
    async with self.context.lock:
  File ".../site-packages/anyio/_core/_synchronization.py", line 173, in __aexit__
    self.release()
  File ".../site-packages/anyio/_backends/_asyncio.py", line 1935, in release
    raise RuntimeError("The current task is not holding this lock")
RuntimeError: The current task is not holding this lock

(The other capture is identical except the GeneratorExit lands at line 746, token_response = yield await self._perform_authorization().)

After this fires, every reconnect attempt for that server times out with no HTTP traffic — the flow generator never gets past line 582.

Minimal reproduction (no network needed)

import asyncio
import httpx2
from mcp.client.auth.oauth2 import OAuthClientProvider
from mcp.shared.auth import OAuthClientMetadata


class MemStorage:
    async def get_tokens(self): return None
    async def set_tokens(self, tokens): pass
    async def get_client_info(self): return None
    async def set_client_info(self, client_info): pass


async def main():
    provider = OAuthClientProvider(
        server_url="https://example.invalid/mcp",
        client_metadata=OAuthClientMetadata(redirect_uris=["http://localhost:1/callback"]),
        storage=MemStorage(),
    )

    # Task A: advance the flow. With no stored tokens it suspends at
    # `response = yield request`, still inside `async with context.lock`.
    flow = provider.async_auth_flow(httpx2.Request("POST", "https://example.invalid/mcp"))
    await flow.asend(None)

    # Task B: httpx tears the stream down from a different task on
    # cancellation, closing the generator there.
    try:
        await asyncio.get_running_loop().create_task(flow.aclose())
    except RuntimeError as exc:
        print(f"aclose raised: {exc!r}")            # <- fires

    print("lock still held:", provider.context.lock.locked())   # True

    # The "reconnect" — blocks forever at line 582:
    flow2 = provider.async_auth_flow(httpx2.Request("POST", "https://example.invalid/mcp"))
    try:
        await asyncio.wait_for(flow2.asend(None), timeout=2)
        print("second flow proceeds")
    except asyncio.TimeoutError:
        print("second flow BLOCKED on poisoned lock")           # <- happens
    finally:
        await flow2.aclose()

asyncio.run(main())

Output on 2.0.0:

aclose raised: RuntimeError('The current task is not holding this lock')
lock still held: True
second flow BLOCKED on poisoned lock

Root cause

anyio.Lock is task-bound by design; httpx makes no guarantee that the auth-flow generator is closed from the task that advanced it (cancellation/teardown commonly runs aclose() from a sibling task). Holding a task-bound lock across the generator's yields therefore poisons the lock on any cross-task close: the release both raises and never happens.

Suggested fix

Serializing the auth flow per-context is still needed (token refresh must not race), but the primitive must allow release from the closing task. Options:

  1. Use anyio.Semaphore(1) instead of anyio.Lock for OAuthContext.lock. anyio semaphores are not owner-bound, so release from the closing task is legal on both asyncio and trio backends. One-line change in the OAuthContext dataclass plus the async with keeps working.
  2. Alternatively, wrap the generator body so GeneratorExit/cross-task unwind releases via a tolerant path (catch the ownership RuntimeError and force the lock back to a released state), though anyio has no public API for that today.

Happy to send a PR for option 1 if that direction is acceptable.

Activity

  1. added
    v2Affects the v2 line (2.x on main)
    v1Affects the v1.x maintenance line
    on Aug 25, 2026
  2. jstar0 commented on Aug 25, 2026

    @jstar0

    I reproduced the failure against current main (4d6f87e8) and confirmed the proposed primitive change isolates the lifecycle problem:

    • A flow suspended inside async with anyio.Lock() raises RuntimeError: The current task is not holding this lock when a different asyncio task calls aclose(), and the lock remains held.
    • The same flow using anyio.Semaphore(1) closes cleanly from the other task and the next flow acquires the semaphore.
    • OAuthContext.lock is only used by OAuthClientProvider.async_auth_flow, which intentionally holds it across the generator's HTTP yields, so the change can stay local to the OAuth context and preserve the existing async with serialization.

    The regression should exercise the generator lifecycle rather than only checking the field type: advance one async_auth_flow from one task until it yields, close it from a second task, then prove a subsequent flow can proceed. The test should use the repository's anyio style and bound the wait with fail_after, covering the cross-task teardown path without wall-clock assertions.

    This looks like a narrowly scoped compatibility fix rather than an API change. Please confirm whether Semaphore(1) is the preferred direction; I will follow the repository's assignment/help-wanted gate before opening any PR.

  3. dportier1021-crypto commented on Aug 27, 2026

    @dportier1021-crypto

    Production field report confirming this exact failure mode (mcp 2.0.0, CPython 3.11, macOS arm64, anyio 4.x asyncio backend).

    A long-lived MCP client process (an ACP server hosting several OAuth streamable-HTTP servers: Notion, Composio, Consensus) hit this on 2026-08-22 03:44 and stayed poisoned for 4.5 days until the process was manually killed.

    Trigger sequence (from logs):

    1. GET stream disconnected, reconnecting in 1000ms... (server closed the standalone GET stream)
    2. Session teardown closed the in-flight async_auth_flow generator from a different task:
    ERROR asyncio: Task exception was never retrieved
    future: <Task finished coro=<<async_generator_athrow without __name__>()>
      exception=RuntimeError('The current task is not holding this lock')>
    Traceback:
      File "mcp/client/auth/oauth2.py", line 601, in async_auth_flow
        response = yield request
    GeneratorExit
    During handling of the above exception, another exception occurred:
      File "mcp/client/auth/oauth2.py", line 582, in async_auth_flow
        async with self.context.lock:
      File "anyio/_core/_synchronization.py", line 166, in __aexit__
        self.release()
      File "anyio/_backends/_asyncio.py", line 1854, in release
        raise RuntimeError("The current task is not holding this lock")
    RuntimeError: The current task is not holding this lock
    

    Post-poison behavior (matches your analysis exactly): context.lock remained held forever. Every subsequent connect attempt for the affected servers hung inside async with self.context.lock — producing zero network traffic (verified via socket inspection: no SYN, only CLOSE_WAIT zombies) — until each attempt hit its 300s connect timeout. Client-side symptom was a perfectly periodic retry metronome (300s backoff + 3×300s hung attempts = 25m07s cycle) that survived restarts of every other client process on the machine, because the poisoned state is in-memory per-process.

    Additional confirmation of the mechanism: the same OAuth servers handshaked fine (<1s initialize+list_tools) from fresh processes throughout the entire incident — tokens, network, and servers were never the problem. Only the poisoned process could not connect.

    Killing the process fully resolved it. Strong +1 for the Semaphore(1) direction (or any release path that doesn't assume the releasing task is the acquirer).

  4. vultrarex21-dev commented on Aug 27, 2026

    @vultrarex21-dev

    I’d like to take this issue if it is still available.

    I’ve reviewed the reproduction and the production report, and
    the cross-task generator teardown behavior looks like the key
    lifecycle issue.

    I can implement the Semaphore(1) change together with the
    cross-task async_auth_flow regression test, following the
    repository’s contribution workflow.

    If this is still available, please assign it to me.

  5. anythingoah commented on Sep 3, 2026

    @anythingoah

    This looks nasty to debug blind did the RuntimeError point you straight at the lock, or did you have to trace back from a hang/timeout first? Also curious if this showed up right after an SDK bump or you'd been running this version a while before it surfaced.

  6. Kludex commented on Oct 10, 2026

    @Kludex
    Member

    Both reports identify OAuthClientProvider.async_auth_flow holding an anyio.Lock across a yield, which can fail on cross-task generator closure and leave the lock held. This is tracked in #2847, so I’m closing this as a duplicate. AI-assisted triage; I reviewed both reports.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingv1Affects the v1.x maintenance linev2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions