hermes-agent/tests/agent/test_preflight_compression_gate.py
Teknium 6b81590c55
test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00

67 lines
2.3 KiB
Python

"""Regression tests for issue #27405.
The preflight compression gate must trigger when *either* the message
count exceeds the protected ranges OR the cheap char-based token
estimate already crosses the configured threshold. Pre-fix, only the
message-count condition was checked, so a session with a small number
of huge messages would silently skip compression and eventually hit a
hard context-overflow error.
"""
from agent.turn_context import _should_run_preflight_estimate
# Protected-range counts mirror the compressor defaults. THRESHOLD_TOKENS is an
# arbitrary test threshold passed explicitly into the helper — it is NOT the
# live runtime threshold (which is max(0.5*window, MINIMUM_CONTEXT_LENGTH) per
# model); the helper takes the threshold as a parameter so the tests are
# self-contained and independent of model metadata.
PROTECT_FIRST_N = 3
PROTECT_LAST_N = 20
THRESHOLD_TOKENS = 64_000
def _msg(content: str) -> dict:
return {"role": "user", "content": content}
def test_few_messages_huge_content_triggers_gate():
"""The bug from #27405: 8 messages with one massive content blob."""
# ~280K chars in one message ~= 70K tokens at 4 chars/token.
big = "x" * 280_000
messages = [_msg("hi")] * 7 + [_msg(big)]
assert len(messages) <= PROTECT_FIRST_N + PROTECT_LAST_N + 1 # would fail old gate
assert _should_run_preflight_estimate(
messages, PROTECT_FIRST_N, PROTECT_LAST_N, THRESHOLD_TOKENS
) is True
def test_content_above_threshold_triggers():
"""A single message comfortably above the threshold trips branch (b)."""
# ~threshold*4 chars => ~threshold tokens; +1000 tokens of margin so the
# test doesn't depend on per-message dict-wrapping overhead in the
# shared estimator's (chars+3)//4 rounding.
messages = [_msg("x" * ((THRESHOLD_TOKENS + 1000) * 4))]
assert _should_run_preflight_estimate(
messages, PROTECT_FIRST_N, PROTECT_LAST_N, THRESHOLD_TOKENS
) is True
def test_content_below_threshold_does_not_trigger():
"""A single message comfortably below the threshold (and few messages)
must not trigger — the estimator stays under and the count gate is not
tripped."""
messages = [_msg("x" * ((THRESHOLD_TOKENS - 1000) * 4))]
assert _should_run_preflight_estimate(
messages, PROTECT_FIRST_N, PROTECT_LAST_N, THRESHOLD_TOKENS
) is False