Skip to content

Pending outgoing control requests (interrupt, set_permission_mode, etc.) hang for the full 60s timeout after close()/disconnect() instead of failing fast #1094

Description

@mikko-lab

Related Issues / PRs

Description

Query (src/claude_agent_sdk/_internal/query.py) tracks in-flight outgoing control requests (wire subtypes interrupt, set_permission_mode, set_model, rewind_files, mcp_reconnect, mcp_toggle, stop_task, mcp_status, get_context_usage, initialize) in two dicts:

# lines 114-115
self.pending_control_responses: dict[str, anyio.Event] = {}
self.pending_control_results: dict[str, dict[str, Any] | Exception] = {}

_send_control_request (lines 501-546) registers an anyio.Event for the request (line 519) and waits on it under anyio.fail_after(timeout) (default 60s):

# lines 531-546
try:
    with anyio.fail_after(timeout):
        await event.wait()
    result = self.pending_control_results.pop(request_id)
    self.pending_control_responses.pop(request_id, None)
    ...
except TimeoutError as e:
    self.pending_control_responses.pop(request_id, None)
    self.pending_control_results.pop(request_id, None)
    raise Exception(f"Control request timeout: {request.get('subtype')}") from e

There are exactly two places that set the event and unblock this wait:

  1. Normal resolution in _read_messages, when a control_response message arrives (lines 260-268, event.set() at line 268).
  2. The bulk-fail loop added by fix: propagate CLI errors to pending control requests #388, at lines 328-333, inside except Exception as e: (line 328) in _read_messages.

But _read_messages has an earlier, separate handler for cancellation:

# lines 324-327
except anyio.get_cancelled_exc_class():
    # Task was cancelled - this is expected behavior
    logger.debug("Read task cancelled")
    raise  # Re-raise to properly handle cancellation
except Exception as e:
    # Signal all pending control requests so they fail fast instead of timing out
    for request_id, event in list(self.pending_control_responses.items()):
        ...

close()_close_impl() (lines 886-907) cancels the read task directly:

# from _close_impl
if self._read_task is not None and not self._read_task.done():
    self._read_task.cancel()
    await self._read_task.wait()

This makes the read task raise anyio.get_cancelled_exc_class(), which is caught by the except anyio.get_cancelled_exc_class(): branch at line 324 — a BaseException subclass, not caught by except Exception at line 328. That branch only logs and re-raises; it never touches pending_control_responses/pending_control_results. Nothing else in _close_impl signals them either.

Result: if a caller sends a control request (most commonly interrupt(), e.g. as part of a "stop" action) and then calls close()/disconnect() before the CLI has responded, disconnect() itself returns immediately (it doesn't wait on the pending request), but the coroutine awaiting interrupt() — or whichever request is in flight — is not signaled and sits blocked until _send_control_request's own anyio.fail_after(60.0) fires, 60 seconds later, raising Exception("Control request timeout: interrupt"). For a caller that awaits interrupt() and disconnect() together (a natural pattern for a "stop the running turn" action), this looks like a hang, not a clean cancellation.

Reproduction Steps / Example Code (Python)

Driven entirely through the public API (ClaudeSDKClient with a custom Transport, the SDK's own supported extension point — not a direct call into _internal). The fake transport acknowledges initialize immediately (so the client finishes connecting) but never responds to interrupt, simulating "the CLI hasn't answered yet" — exactly the window close()/disconnect() can land in during normal use.

import anyio
from claude_agent_sdk import ClaudeSDKClient
from claude_agent_sdk._internal.transport import Transport


class FakeTransport(Transport):
    """Acks `initialize`, never responds to anything else (e.g. `interrupt`)."""

    def __init__(self):
        self._send, self._recv = anyio.create_memory_object_stream(10)
        self.connected = False

    async def connect(self) -> None:
        self.connected = True

    async def write(self, data: str) -> None:
        import json
        msg = json.loads(data)
        if msg.get("type") == "control_request" and msg["request"].get("subtype") == "initialize":
            await self._send.send({
                "type": "control_response",
                "response": {"subtype": "success", "request_id": msg["request_id"], "response": {}},
            })
        # anything else (e.g. "interrupt") is silently dropped -- no response ever sent

    async def read_messages(self):
        async for msg in self._recv:
            yield msg

    async def close(self) -> None:
        self.connected = False

    def is_ready(self) -> bool:
        return self.connected

    async def end_input(self) -> None:
        pass


async def main():
    client = ClaudeSDKClient(transport=FakeTransport())
    await client.connect()

    async def do_interrupt():
        try:
            with anyio.move_on_after(3.0) as scope:
                await client.interrupt()
            print("interrupt() still pending after 3s deadline:", scope.cancelled_caught)
        except Exception as e:
            print("interrupt() raised:", e)

    async with anyio.create_task_group() as tg:
        tg.start_soon(do_interrupt)
        await anyio.sleep(0.05)  # ensure interrupt() is genuinely in flight
        import time
        t0 = time.monotonic()
        await client.disconnect()
        print(f"disconnect() returned in {time.monotonic() - t0:.4f}s")


anyio.run(main)

Actual output (current main, 3 runs, identical each time)

disconnect() returned in 0.0003s
interrupt() still pending after 3s deadline: True

Run with a longer deadline (65s) instead of 3s to show what eventually happens:

interrupt() raised: Control request timeout: interrupt

— firing at 60.00s, matching _send_control_request's own anyio.fail_after(60.0) at line 532 exactly (confirmed via time.monotonic() around the call: 60.002s elapsed). This is the SDK's own eventual-timeout path, not an artifact of the test — disconnect() had already returned 60 seconds earlier.

Expected: close()/disconnect() should signal any still-pending outgoing control requests immediately (the same pending_control_results[request_id] = <some cancelled/closed exception>; event.set() pattern #388 already added for the crash case, lines 329-333), so interrupt() (or whichever request was in flight) fails fast with a clear "connection closed" error instead of silently hanging for up to 60 seconds after the caller has already moved on.

System Info

claude-agent-sdk-python, commit fdee0adc99f46e65ae9d6d029a6f4fb31bb8cffa (main, verified 2026-07-09/10)
python: 3.14, anyio backend (also reproduced on trio backend, same result)

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions