hermes-agent/tests/gateway/test_moa_one_shot_restore.py
Teknium 39975613b1
test: prune wave 2 + speed fixes — 28,106 → 19,757 test functions, suite wall 315s → 294s
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.

Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
  (DEFAULT_CONFIG smart-approval leaked in) — pinned approval
  mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
  silently probed live provider catalogs (~2s/test) — stubbed
  cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
  shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
  MagicMock agents — interrupt.side_effect now clears _running_agents
  (22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
  (24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
  patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
  5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
  with rationale comment — the watchdog-dump assertion needs the loop
  blocked well past deadline+grace under parallel load (flaked once in
  the 40-worker verification run at a 0.2s margin).

Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
2026-07-29 13:39:40 -07:00

55 lines
1.9 KiB
Python

"""MoA one-shot model override must be restored on both success and failure.
These exercise the real ``GatewayRunner._restore_moa_one_shot`` helper that the
message-handling ``finally`` block calls, so they prove the production logic —
not a re-implementation of it. The bug being guarded: the restore used to live
in the ``try`` block, so a turn that raised skipped it and the MoA override
leaked permanently (every later message silently fanned out through MoA).
"""
from types import SimpleNamespace
from unittest.mock import MagicMock
from gateway.run import GatewayRunner
def _make_runner():
"""Minimal GatewayRunner with only the fields _restore_moa_one_shot reads."""
runner = object.__new__(GatewayRunner)
runner._session_model_overrides = {}
runner._evict_cached_agent = MagicMock()
return runner
def _make_event(moa_disable=False, moa_restore=None):
event = SimpleNamespace()
if moa_disable:
event._moa_disable_after_turn = True
event._moa_restore_override = moa_restore
return event
def test_restore_runs_from_finally_even_when_turn_raises():
"""The whole point of the fix: a raising turn still reverts the override.
Mirrors the real call site — the restore is invoked from a ``finally`` block,
so it fires after an exception propagates out of the turn body.
"""
runner = _make_runner()
key = "agent:main:telegram:dm:999"
runner._session_model_overrides[key] = {"provider": "moa", "model": "default"}
event = _make_event(
moa_disable=True,
moa_restore={"provider": "openrouter", "model": "gpt-4"},
)
with __import__("pytest").raises(RuntimeError):
try:
raise RuntimeError("provider error mid-turn")
finally:
runner._restore_moa_one_shot(event, key)
assert runner._session_model_overrides[key] == {
"provider": "openrouter",
"model": "gpt-4",
}