hermes-agent/tests/tools/test_delegate_summary_budget.py
Teknium 39975613b1
test: prune wave 2 + speed fixes — 28,106 → 19,757 test functions, suite wall 315s → 294s
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.

Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
  (DEFAULT_CONFIG smart-approval leaked in) — pinned approval
  mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
  silently probed live provider catalogs (~2s/test) — stubbed
  cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
  shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
  MagicMock agents — interrupt.side_effect now clears _running_agents
  (22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
  (24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
  patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
  5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
  with rationale comment — the watchdog-dump assertion needs the loop
  blocked well past deadline+grace under parallel load (flaked once in
  the 40-worker verification run at a 0.2s margin).

Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
2026-07-29 13:39:40 -07:00

78 lines
3.2 KiB
Python

"""Tests for subagent summary budgeting (PR #9126).
delegate_task caps subagent summaries against the parent's remaining context
headroom (split across the batch) before they enter the parent's context, and
spills the full text to disk so nothing is lost. This guards the
compression/429 death spiral that batch fan-out could trigger by returning N
full summaries verbatim into the parent.
"""
import os
import tempfile
import pytest
import tools.delegate_tool as dt
class _FakeCompressor:
def __init__(self, context_length, max_tokens):
self.context_length = context_length
self.max_tokens = max_tokens
class _FakeParent:
def __init__(self, context_length, used_tokens, max_tokens):
self.context_compressor = _FakeCompressor(context_length, max_tokens)
self.session_prompt_tokens = used_tokens
def test_small_summaries_pass_through_untouched():
parent = _FakeParent(context_length=200_000, used_tokens=10_000, max_tokens=8_000)
results = [
{"task_index": 0, "summary": "short result A", "status": "completed"},
{"task_index": 1, "summary": "short result B", "status": "completed"},
]
dt._apply_summary_budget(results, parent)
assert results[0]["summary"] == "short result A"
assert "summary_truncated" not in results[0]
assert "summary_truncated" not in results[1]
def test_batch_overflow_trimmed_and_spilled_losslessly(monkeypatch):
# Isolate spill directory to a temp HERMES_HOME.
with tempfile.TemporaryDirectory() as td:
monkeypatch.setenv("HERMES_HOME", os.path.join(td, ".hermes"))
# Distinct head + tail markers so we can prove the tail survives.
big = "HEAD_MARKER\n" + ("X" * 50_000) + "\nTAIL_MARKER"
# Parent nearly full (120k/131k) → tiny headroom → aggressive trim.
parent = _FakeParent(context_length=131_000, used_tokens=120_000, max_tokens=8_000)
results = [
{"task_index": i, "summary": big, "status": "completed"} for i in range(5)
]
dt._apply_summary_budget(results, parent)
for r in results:
assert r["summary_truncated"] is True
assert len(r["summary"]) < len(big)
# Head+tail window: both ends survive in-context.
assert "HEAD_MARKER" in r["summary"]
assert "TAIL_MARKER" in r["summary"]
path = r.get("summary_full_path")
assert path and os.path.exists(path)
# The spill file holds the FULL original text — nothing is lost.
with open(path, encoding="utf-8") as fh:
assert fh.read() == big
# The footer points the parent at the full version with an offset.
assert "read_file" in r["summary"]
assert "offset=" in r["summary"]
# Spilled into the delegation cache (mounted into remote backends).
assert os.path.join("cache", "delegation") in path
def test_empty_results_is_noop():
# No summaries → nothing to do, must not raise.
dt._apply_summary_budget([], _FakeParent(131_000, 1_000, 8_000))
dt._apply_summary_budget(
[{"task_index": 0, "status": "failed", "summary": None}],
_FakeParent(131_000, 1_000, 8_000),
)