fix(computer_use): reconnect a dead cua-driver session instead of hanging (#67138)

Bug 1 of #55048: when the MCP connection dropped (driver crash / restart),
_lifecycle_coro exited but left _started=True, so the next list_apps/capture
passed _require_started() and then operated on a None session — hanging
forever instead of reconnecting.

- _lifecycle_coro's finally now resets _started=False on ANY exit, so a dead
  session is re-enterable (idempotent no-op on the normal stop() path; atomic
  bool write, safe from the bridge-loop thread without the lock stop() holds).
- call_tool() re-enters start() when the session isn't active, rebuilding it
  before the call. The start_session/end_session handshake (driven by start()/
  stop() themselves) is exempted so bootstrap doesn't recurse.

Tests: two cases in test_computer_use_delivery_ladder.py — finally resets
_started, and call_tool restarts a dead session exactly once. Full
computer_use suite green (233).

Refs #55048 (Bug 1). Bug 2 (expose foreground dispatch) is covered by the
delivery_mode work in #67123.
This commit is contained in:
Teknium 2026-07-18 15:07:04 -07:00 committed by GitHub
parent c34b29d11a
commit 7a43ab042f
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 86 additions and 0 deletions

View file

@ -725,6 +725,16 @@ class _CuaDriverSession:
# outer context-manager exits AFTER this block, so set to
# None here is fine: stop() has already flipped _started.
self._session = None
# Reset _started so a session that dies for ANY reason (MCP
# connection drop, driver crash, unexpected coro exit) is
# re-enterable: the next start()/call sees _started False and
# rebuilds the session instead of hanging forever on a dead one
# via _require_started(). On the normal stop() path this is a
# harmless idempotent no-op (stop() already set it False). A
# plain bool write is atomic in CPython, so this is safe from
# the bridge-loop thread without taking self._lock (which stop()
# may hold while awaiting this coro's future). See #55048 Bug 1.
self._started = False
async def _populate_capabilities(self, session: Any) -> None:
"""Surface 4: cache per-tool capability sets + capability_version
@ -1060,7 +1070,24 @@ class _CuaDriverSession:
except OSError:
pass
# Lifecycle handshake calls issued BY start()/stop() themselves — these
# must not trigger the auto-restart guard below, or start() would recurse
# into start() when the session-start hasn't flipped _started yet.
_LIFECYCLE_CALLS = frozenset({"start_session", "end_session"})
def call_tool(self, name: str, args: Dict[str, Any], timeout: float = 30.0) -> Dict[str, Any]:
# A prior session may have died (MCP drop / driver crash): its
# lifecycle coro reset _started to False in its finally (#55048
# Bug 1). Re-enter start() so we rebuild the session instead of
# calling _require_started() straight into a "not started" raise or
# a None session. start() is idempotent when already started. Skip
# this for the start_session/end_session handshake, which start()/
# stop() drive directly while _started is still in flux.
if not self._started and name not in self._LIFECYCLE_CALLS:
logger.warning(
"cua-driver session not active on %s; (re)starting before call", name
)
self.start()
self._require_started()
# The cua-driver daemon proxy returns POSIX EAGAIN ("Resource
# temporarily unavailable") for heavier calls like get_window_state when