hermes-agent/tests/agent/test_close_interrupted_tool_sequence.py
Teknium 6b81590c55
test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00

68 lines
2.3 KiB
Python

"""Regression tests for ``close_interrupted_tool_sequence`` (#48879 follow-up).
#48879 closed the tool-call sequence on interrupt inside ``finalize_turn``,
but the retry/backoff/error interrupt aborts in ``conversation_loop`` ``return``
early and never reach it — so they persisted a raw ``tool`` tail. The next user
message then lands as ``... tool → user``, the role-alternation violation that
makes strict providers (Gemini, Claude) hallucinate a continuation and ignore
prior context (what the user perceives as "lost context").
The fix routes every interrupt abort through this one shared helper. These tests
pin the helper's contract and prove the post-interrupt + next-user-message
transcript is alternation-safe.
"""
from agent.message_sanitization import close_interrupted_tool_sequence
def _tool_tail():
return [
{"role": "user", "content": "edit the file"},
{
"role": "assistant",
"content": "",
"tool_calls": [{"id": "c1", "function": {"name": "patch", "arguments": "{}"}}],
},
{"role": "tool", "tool_call_id": "c1", "content": "ok edited"},
]
def _assert_no_tool_then_user(messages):
for i in range(len(messages) - 1):
if messages[i].get("role") == "tool":
assert messages[i + 1].get("role") != "user", (
f"role-alternation violation: tool → user at index {i}"
)
def test_closing_makes_next_user_message_alternation_safe():
"""The whole point: appending a user turn after the close must not
produce the ``tool → user`` shape strict providers choke on."""
messages = _tool_tail()
close_interrupted_tool_sequence(messages, None)
follow_on = messages + [{"role": "user", "content": "they do! increase the timing"}]
_assert_no_tool_then_user(follow_on)
def test_assistant_tail_is_left_untouched():
messages = [
{"role": "user", "content": "hi"},
{"role": "assistant", "content": "partial reply"},
]
before = [dict(m) for m in messages]
assert close_interrupted_tool_sequence(messages, "interrupted") is False
assert messages == before
def test_user_tail_is_left_untouched():
messages = [{"role": "user", "content": "hi"}]
assert close_interrupted_tool_sequence(messages, None) is False
assert len(messages) == 1