hermes-agent/tests/gateway/test_session_store_lock_io.py
Teknium 597615ade4
fix(ci): make tests, workflows, and attribution reliable under load (#66373)
* feat(attribution): conflict-free contributor mappings via contributors/emails/ directory

The AUTHOR_MAP dict in scripts/release.py was a merge-conflict magnet:
every concurrent salvage PR appended entries to the same lines of the
same file, so parallel PRs re-conflicted on every merge to main.

New system: one file per email under contributors/emails/ — filename is
the commit-author email, first non-comment line is the GitHub login.
File additions never conflict, so any number of PRs can add mappings
concurrently.

- scripts/release.py: AUTHOR_MAP is now LEGACY_AUTHOR_MAP (frozen)
  merged with the directory at import time (directory wins). All
  existing consumers (resolve_author, contributor_audit.py) unchanged.
- scripts/add_contributor.py: idempotent CLI to add a mapping; refuses
  conflicting reassignments (incl. against the legacy map), validates
  email/login shapes.
- contributor-check.yml: attribution gate now accepts a mapping file OR
  a legacy entry; failure message prints the exact add_contributor
  command. Also auto-resolves bare <login>@users.noreply.github.com
  emails is intentionally NOT added (kept id+login form only, matching
  previous behavior).
- contributor_audit.py: guidance now points at add_contributor.py.
- tests/scripts/test_contributor_map.py: 12 tests covering loader,
  merge precedence, CLI idempotency/conflict/validation, subprocess E2E.

* feat(ci): one-shot per-file flake retry in the parallel test runner

A failing test FILE is re-run once in a fresh subprocess. Pass-on-retry
counts as green but is loudly reported in a '⚠ FLAKY' summary section
(with both attempts' output preserved) so the flake gets fixed instead
of eating a full-run rerun. Deterministic failures fail both attempts —
regressions cannot be laundered green.

- --file-retries N / HERMES_TEST_FILE_RETRIES (default 1, 0 disables)
- E2E verified: simulated first-run-fail flake goes green with banner;
  deterministic failure still exits 1; retries=0 restores old behavior.

This converts the dominant CI failure mode (one timing-sensitive test
flaking a 4600-test shard, requiring a manual 10-minute rerun and an
agent triage loop) into a self-healing retry that costs one file's
runtime.

* test(approval): loosen wall-clock perf bounds 0.15s -> 2.0s

These guard against catastrophic regex backtracking (seconds-to-minutes
class), but 0.15s is within scheduler-stall noise on loaded shared CI
runners — test_max_accepted_separator_free_input_is_fast failed a CI
shard this week on runner load alone. 2.0s still catches the regression
class with zero flake surface.

* fix(ci): job timeouts everywhere + retries on all network installs

Reliability pass over every workflow:
- timeout-minutes on all 21 jobs that lacked one (a hung job previously
  burned the 6-hour default runner budget)
- ./.github/actions/retry wrapped around every network-fetching install
  that lacked it: pip installs (deploy-site, skills-index), npm ci
  (deploy-site website, upload_to_pypi web + ui-tui), uv sync (docker
  test deps). Deterministic build steps (npm run build) deliberately
  NOT retried — split into separate steps so a real build failure fails
  fast instead of retrying 3x.

* docs(agents): document the file-retry flake policy

* fix(ci): curl retries on deploy hook + skills-index probe

* fix(ci): kill the remaining transient-failure classes in workflows + Dockerfile

From the workflow reliability audit:
- tests.yml: duration-cache restore had NO restore-keys while saves use
  run_id-suffixed keys — the cache never matched once, so LPT slicing
  always ran blind and unbalanced slices pushed heavy files toward the
  per-file timeout. One-line restore-keys fixes slice balancing.
- Label gates (lint ci-reviewed, supply-chain mcp-catalog-reviewed):
  'gh pr view || true' turned an API blip into 'label absent' → false
  BLOCKING failure. Now 3x retry, and API failure is reported as an API
  failure instead of a missing label.
- detect-changes action: compare API retried before failing open (was
  silently running all lanes on any blip).
- uv-lockfile-check: 'uv lock --check' resolves against PyPI — retried
  so registry blips don't read as 'lockfile stale'.
- docker.yml merge job: imagetools create retried (Docker Hub eventual
  consistency on just-pushed digests).
- Dockerfile: apt-get Acquire::Retries=3; s6-overlay ADDs converted to
  curl --retry 3 (ADD cannot retry; checksums still enforced); npm
  --fetch-retries=5; playwright chromium fetch retried 3x.
- Advisory artifact uploads (per-slice durations, ci-timings report)
  get continue-on-error so an artifact-service blip can't fail a green
  test slice.

* fix(tests): kill the two root-cause flakes — leaking pre-warm timer + env-dependent provider list

- test_tui_gateway_server.py: session.create / non-eager session.resume
  arm a 50ms threading.Timer (_schedule_agent_build) that outlives its
  test and fires into the NEXT test's _make_agent mock, racily
  corrupting captured state (the recurring session_resume shard
  failures). Replaced the per-test whack-a-mole stub with a module-wide
  autouse fixture; the 3 worker-lifecycle tests that genuinely need the
  deferred build opt back in via @pytest.mark.real_agent_prewarm (new
  marker in pyproject).
- test_api_key_providers.py: PROVIDER_ENV_VARS is now derived from the
  live PROVIDER_REGISTRY instead of a hand-list that had drifted
  (missing HF_TOKEN / DEEPINFRA_API_KEY) — resolve_provider('auto')
  tests failed on any machine with HF_TOKEN exported. E2E-verified with
  HF_TOKEN/DEEPINFRA_API_KEY set: 42/42 pass.

* test: de-flake 30 timing-sensitive test files for loaded CI runners

Root-cause fixes from the flake audit (session-DB mining + repo sweep):

Event-based sync instead of sleep-sync:
- title_generator: mock sets threading.Event, wait(10) replaces
  sleep(0.3) hoping the daemon thread got scheduled
- docker zombie_reaping / profile_gateway: poll-for-state helpers
  replace fixed 1-3s sleeps (s6 transitions + SIGCHLD reaping are async)
- process_registry tree test: select()-bounded readline replaces an
  unbounded blocking read (parent wedge now fails THIS test with a clear
  message instead of an opaque rc=124 file kill); SIGTERM grace 1s->2s
  (the 1s partition window mid-interpreter-startup is how a child PID
  escaped the live-system guard in CI)

Timeout raises (loaded 8-way-sliced runners see ~5s scheduling floors;
all of these complete in ms-to-1s when healthy so the raises cost
nothing on green runs):
- subprocess/thread waits <= 2s raised to 10-15s across mcp_tool,
  mcp_circuit_breaker, mcp_reconnect_retry_reset, mcp_parked_self_probe,
  mcp_cancelled_error_propagation, registry, clarify_gateway, interrupt,
  voice_cli_integration, docker_environment, session_store_lock_io,
  planned_stop_watcher, cli_interrupt_subagent, thread_scoped_output
  (joins now also assert not is_alive() so stragglers fail loudly)
- wall-clock discrimination ceilings loosened where the guarded hang is
  10x larger: local_background_child_hang 4s->10s, interrupt_cleanup
  setup 5s->20s + pgid-exit 30s->60s, mcp_stability grandchild spinup
  5s->15s, protocol/gil-starvation fast-handler 0.5s->2s,
  iso_certify_seam 1.5s->5s, wait_for_mcp_discovery 0.1s->1s
- narrow assertion windows widened: honcho first-turn wait 0.4..0.65 ->
  0.25..2.0 (property is bounded-not-hung, not an exact wall-clock);
  compression fork-lock TTL 1s->3s (12 refresh chances per lease);
  compression-lock expiry margins symmetric (ttl 0.05->0.5, sleep 1.0)
- telegram hung-DNS bound 1.0->1.4 (fake hang is 1.5s — must stay under)

* fix(tests): repair indentation from de-flake batch edit

* fix(tests): harden env isolation and replace remaining sleep-sync races

The full 42k-test run and complete npm check surfaced three more classes:

- Environment isolation: local ~/.honcho defaultHost and SSH_* variables
  leaked into Python/TUI tests. Pin the default Honcho host in the
  hermetic fixture, isolate the one fallback test from ~/.honcho, and
  blank SSH_* around terminalSetup tests. This flipped 20 false failures
  back to deterministic behavior on developer machines.
- Background-thread sleep-sync: Honcho async writer tests patched
  time.sleep globally, then busy-polled with that same mocked sleep. Under
  full-suite load the poller could starve the writer. Each test now waits
  on an Event emitted by the exact flush/retry transition; 30/30 passed
  under 15-way contention.
- Desktop streaming: the test slept 80ms and assumed a 500ms timer could
  not fire before its assertion. A loaded runner descheduled the test for
  >500ms and both chunks arrived. Producer controls now gate second-chunk
  and completion transitions explicitly.

Also make file-retry observability complete: a self-healed flaky file now
prints BOTH attempts' full output in the FLAKY summary. Two behavioral
runner tests prove pass-on-retry is green+loud+traceback-preserving, while
a deterministic failure remains red.

* refactor(ci): use gh bot pat, better retries

refactor(ci): use retry action for PR label fetch
the retry action now captures stdout as a step output, so it can serve
double duty: retry + output capture for commands like 'gh pr view' whose
result must be consumed by later steps.

Retry action gains:
- 'stdout' output (heredoc-delimited to preserve newlines)
- tee to temp file so stdout still streams to the job log
- step id 'retry' for output reference

Both lint.yml and supply-chain-audit.yml now use the retry action
directly with 'command: gh pr view ...' and read
steps.<id>.outputs.stdout.

ci: use AUTOFIX_BOT_PAT for all gh CLI / GitHub API auth

Replace secrets.GITHUB_TOKEN and github.token with
secrets.AUTOFIX_BOT_PAT across all workflows and composite actions
that use the gh CLI or GitHub API. The PAT has consistent permissions
across fork PRs (where GITHUB_TOKEN is read-only), avoids API rate
limit sharing with the default token, and is already used by
js-autofix.yml for the same reasons.

19 sites swapped across 9 files:
- lint.yml (3): label fetch, comment post/edit, comment update
- supply-chain-audit.yml (5): scan, critical comment, unbounded dep
  comment, label fetch, mcp-catalog comment
- lockfile-diff.yml (1): PR comment post/update
- skills-index-freshness.yml (1): issue creation on degraded probe
- skills-index.yml (2): index build, trigger deploy workflow
- upload_to_pypi.yml (2): release view poll, release upload
- ci.yml (1): timings report
- deploy-site.yml (2): skills index crawl
- detect-changes/action.yml (1): compare API call

---------

Co-authored-by: ethernet <arilotter@gmail.com>
2026-07-17 20:55:24 +00:00

389 lines
13 KiB
Python

"""Regression: blocking I/O must not run while session_store._lock is held.
``get_or_create_session`` previously held the store lock during SQLite
SELECTs (``_is_session_ended_in_db``), a full routing-index rewrite +
``os.fsync`` (``_save``), and a recovery DB query
(``_recover_session_from_db``) -- all on every inbound message.
These tests assert those three I/O calls are invoked *outside* the lock.
They follow the mock-DB idiom from ``test_session_store_runtime_stale_guard``.
"""
import json
import threading
from concurrent.futures import ThreadPoolExecutor
from datetime import datetime, timedelta
from unittest.mock import MagicMock, patch
import pytest
from gateway.config import GatewayConfig, Platform, SessionResetPolicy
from gateway.session import SessionEntry, SessionSource, SessionStore
# ---------------------------------------------------------------------------
# Test helpers
# ---------------------------------------------------------------------------
class _TrackedLock:
"""Drop-in replacement for ``threading.Lock`` that tracks hold state.
Used to assert that blocking I/O runs only when the lock is released.
"""
def __init__(self):
self._lock = threading.Lock()
self._held = False
def acquire(self, *a, **kw):
r = self._lock.acquire(*a, **kw)
if r:
self._held = True
return r
def release(self):
self._held = False
self._lock.release()
def __enter__(self):
self.acquire()
return self
def __exit__(self, *a):
self.release()
@property
def held(self) -> bool:
return self._held
def _db_with_rows(rows: dict) -> MagicMock:
"""Mock SessionDB where ``get_session`` maps session_id -> row dict."""
db = MagicMock()
db.get_session.side_effect = lambda sid: rows.get(sid)
db.find_latest_gateway_session_for_peer.return_value = None
db.reopen_session.return_value = None
db.create_session.return_value = None
# Identity compression tip (no child session).
db.get_compression_tip.side_effect = lambda sid: sid
return db
def _make_store(tmp_path, db_mock=None) -> SessionStore:
"""Build a SessionStore with a ``_TrackedLock``, bypassing disk load."""
config = GatewayConfig(default_reset_policy=SessionResetPolicy(mode="none"))
with patch("gateway.session.SessionStore._ensure_loaded"):
store = SessionStore(sessions_dir=tmp_path, config=config)
if db_mock is not None:
store._db = db_mock
store._loaded = True
store._lock = _TrackedLock()
return store
def _source() -> SessionSource:
return SessionSource(
platform=Platform.TELEGRAM,
chat_id="12345",
chat_type="dm",
user_id="12345",
)
def _seed_entry(store, key, session_id) -> SessionEntry:
now = datetime.now()
entry = SessionEntry(
session_key=key,
session_id=session_id,
created_at=now - timedelta(hours=2),
updated_at=now - timedelta(hours=1),
platform=Platform.TELEGRAM,
chat_type="dm",
)
store._entries[key] = entry
return entry
# ---------------------------------------------------------------------------
# Tests
# ---------------------------------------------------------------------------
class TestStaleCheckOutsideLock:
def test_is_session_ended_not_holding_lock(self, tmp_path):
"""``_is_session_ended_in_db`` must run with the lock released."""
source = _source()
db = _db_with_rows({
"sid_alive": {"end_reason": None, "id": "sid_alive"},
})
store = _make_store(tmp_path, db)
key = store._generate_session_key(source)
_seed_entry(store, key, "sid_alive")
lock = store._lock
calls_under_lock = []
orig = store._is_session_ended_in_db
def tracking(sid):
if lock.held:
calls_under_lock.append(sid)
return orig(sid)
store._is_session_ended_in_db = tracking # type: ignore[method-assign]
store.get_or_create_session(source)
assert not calls_under_lock, (
f"_is_session_ended_in_db called {len(calls_under_lock)} "
f"time(s) while lock was held"
)
class TestSaveOutsideLock:
def test_save_not_holding_lock(self, tmp_path):
"""``_save`` must run with the lock released."""
source = _source()
db = _db_with_rows({})
store = _make_store(tmp_path, db)
lock = store._lock
save_calls_under_lock = []
orig_save = store._save_entries
def tracking_save():
if lock.held:
save_calls_under_lock.append(True)
orig_save()
store._save_entries = tracking_save # type: ignore[method-assign]
# force_new bypasses the existing-entry path, goes straight to create.
store.get_or_create_session(source, force_new=True)
assert not save_calls_under_lock, (
f"_save called {len(save_calls_under_lock)} time(s) "
f"while lock was held"
)
class TestRecoverOutsideLock:
def test_recover_not_holding_lock(self, tmp_path):
"""``_recover_session_from_db`` must run with the lock released."""
source = _source()
db = _db_with_rows({})
db.find_latest_gateway_session_for_peer.return_value = {
"id": "sid_recovered",
"started_at": datetime.now().timestamp(),
}
store = _make_store(tmp_path, db)
# No entry seeded -- forces the recovery path.
lock = store._lock
recover_calls_under_lock = []
orig = store._query_recoverable_session
def tracking(**kw):
if getattr(lock, "held", False):
recover_calls_under_lock.append(True)
return orig(**kw)
store._query_recoverable_session = tracking # type: ignore[method-assign]
store.get_or_create_session(source)
assert not recover_calls_under_lock, (
f"_recover_session_from_db called "
f"{len(recover_calls_under_lock)} time(s) while lock was held"
)
def test_concurrent_same_key_returns_one_published_session(tmp_path):
"""Concurrent first messages for one routing key must converge on one ID."""
source = _source()
db = _db_with_rows({})
store = _make_store(tmp_path, db)
owner_started = threading.Event()
release_owner = threading.Event()
original_query = store._query_recoverable_session
def synchronized_query(**kwargs):
owner_started.set()
assert release_owner.wait(timeout=10)
return original_query(**kwargs)
store._query_recoverable_session = synchronized_query # type: ignore[method-assign]
with ThreadPoolExecutor(max_workers=2) as pool:
owner = pool.submit(store.get_or_create_session, source)
assert owner_started.wait(timeout=10)
follower = pool.submit(store.get_or_create_session, source)
release_owner.set()
entries = [owner.result(timeout=10), follower.result(timeout=10)]
key = store._generate_session_key(source)
assert entries[0] is entries[1]
assert entries[0].session_id == store._entries[key].session_id
created_ids = {call.kwargs["session_id"] for call in db.create_session.call_args_list}
assert created_ids == {entries[0].session_id}
def test_concurrent_force_new_returns_one_published_session(tmp_path):
"""Concurrent /new delivery must not create orphan SQLite sessions."""
source = _source()
db = _db_with_rows({})
store = _make_store(tmp_path, db)
owner_started = threading.Event()
release_owner = threading.Event()
original_impl = store._get_or_create_session_impl
def synchronized_impl(*args, **kwargs):
owner_started.set()
assert release_owner.wait(timeout=10)
return original_impl(*args, **kwargs)
store._get_or_create_session_impl = synchronized_impl # type: ignore[method-assign]
with ThreadPoolExecutor(max_workers=2) as pool:
owner = pool.submit(store.get_or_create_session, source, True)
assert owner_started.wait(timeout=10)
follower = pool.submit(store.get_or_create_session, source, True)
release_owner.set()
entries = [owner.result(timeout=10), follower.result(timeout=10)]
assert entries[0] is entries[1]
created_ids = {call.kwargs["session_id"] for call in db.create_session.call_args_list}
assert created_ids == {entries[0].session_id}
def test_auto_reset_does_not_recover_session_being_ended(tmp_path):
source = _source()
db = _db_with_rows({})
store = _make_store(tmp_path, db)
key = store._generate_session_key(source)
old = _seed_entry(store, key, "old-session")
old.suspended = True
db.find_latest_gateway_session_for_peer.return_value = {
"id": old.session_id,
"session_key": key,
"started_at": old.created_at.timestamp(),
}
entry = store.get_or_create_session(source)
assert entry.session_id != old.session_id
assert entry.was_auto_reset is True
db.reopen_session.assert_not_called()
# Auto-reset now writes through promote_to_session_reset (upgrades
# accidental agent_close/ws_orphan_reap ends) with the specific
# auditable reason — a suspended session resets as "suspended".
db.promote_to_session_reset.assert_called_once_with(
old.session_id, "suspended"
)
db.end_session.assert_not_called()
def test_legacy_and_off_lock_saves_share_one_serialization_lock(tmp_path):
db = _db_with_rows({})
persisted: dict[str, str] = {}
first_write_started = threading.Event()
release_first_write = threading.Event()
write_count = 0
count_lock = threading.Lock()
def replace(entries, *, scope):
nonlocal write_count, persisted
with count_lock:
write_count += 1
call_number = write_count
if call_number == 1:
first_write_started.set()
assert release_first_write.wait(timeout=10)
persisted = dict(entries)
db.replace_gateway_routing_entries.side_effect = replace
store = _make_store(tmp_path, db)
source_a = _source()
source_b = SessionSource(
platform=Platform.TELEGRAM,
chat_id="67890",
chat_type="dm",
user_id="67890",
)
key_a = store._generate_session_key(source_a)
key_b = store._generate_session_key(source_b)
_seed_entry(store, key_a, "sid-a")
with ThreadPoolExecutor(max_workers=2) as pool:
future_a = pool.submit(store._save_entries)
assert first_write_started.wait(timeout=10)
_seed_entry(store, key_b, "sid-b")
future_b = pool.submit(store._save)
release_first_write.set()
future_a.result(timeout=10)
future_b.result(timeout=10)
assert set(persisted) == {key_a, key_b}
def test_save_serialization_snapshots_latest_routing_index(tmp_path):
"""A delayed earlier writer must snapshot the state visible when it writes."""
db = _db_with_rows({})
persisted: dict[str, str] = {}
first_write_started = threading.Event()
release_first_write = threading.Event()
write_count = 0
count_lock = threading.Lock()
def replace(entries, *, scope):
nonlocal write_count, persisted
with count_lock:
write_count += 1
call_number = write_count
if call_number == 1:
first_write_started.set()
assert release_first_write.wait(timeout=10)
persisted = dict(entries)
db.replace_gateway_routing_entries.side_effect = replace
store = _make_store(tmp_path, db)
source_a = _source()
source_b = SessionSource(
platform=Platform.TELEGRAM,
chat_id="67890",
chat_type="dm",
user_id="67890",
)
key_a = store._generate_session_key(source_a)
key_b = store._generate_session_key(source_b)
entry_a = _seed_entry(store, key_a, "sid-a")
with ThreadPoolExecutor(max_workers=2) as pool:
future_a = pool.submit(store._save_entries)
assert first_write_started.wait(timeout=10)
entry_b = _seed_entry(store, key_b, "sid-b")
future_b = pool.submit(store._save_entries)
release_first_write.set()
future_a.result(timeout=10)
future_b.result(timeout=10)
assert set(store._entries) == {key_a, key_b}
assert set(persisted) == {key_a, key_b}
assert json.loads(persisted[key_a])["session_id"] == entry_a.session_id
assert json.loads(persisted[key_b])["session_id"] == entry_b.session_id
def test_recovery_rejects_other_profile_row(tmp_path, monkeypatch):
"""The lock-free recovery path must retain the canonical profile guard."""
source = _source()
db = _db_with_rows({})
db.find_latest_gateway_session_for_peer.return_value = {
"id": "foreign-session",
"session_key": "agent:other:telegram:dm:12345",
"started_at": datetime.now().timestamp(),
}
store = _make_store(tmp_path, db)
monkeypatch.setattr(store, "_active_profile_name", lambda: "default")
entry = store.get_or_create_session(source)
assert entry.session_id != "foreign-session"
db.reopen_session.assert_not_called()