Retry of the blocking docker exec approach, this time using
stdout=DEVNULL/stderr=DEVNULL instead of capture_output=True.
The previous attempt regressed on CI (+135s) with capture_output.
Hypothesis: the docker daemon stdio pipe management for long-lived
exec sessions added overhead under 8-way parallel test execution.
DEVNULL should be lighter — the daemon never has to buffer output
or manage pipe read ends.
Also uses user="root" (not "hermes") to avoid deadlocking on
PUID/PGID remap tests where the hermes user UID is being changed
during cont-init.
The probe runs an until/sleep loop inside the container, replacing
N separate docker exec polls with a single blocking exec. On CI
where docker exec costs ~0.25s overhead per call, this should save
(N-1) × 0.25s per container where N is the number of polls
(typically 3-7).
Reduce the default interval_s in poll_container, wait_for_docker_logs,
and inline time.sleep() calls in test_dashboard.py and
test_gateway_run_supervised.py from 0.5s to 0.25s.
CI profile with sleep tracking showed 116.2s of time.sleep() across
345 sleep calls — 23% of total wall time. The majority of these
sleeps were in poll_container (0.5s default) and inline polling
loops in the dashboard and gateway-supervised tests. Halving the
interval halves the sleep time for each poll cycle without changing
docker call count or overhead.
Unlike the blocking-exec experiment (which traded fewer calls for
longer per-call duration and regressed on CI), this change only
reduces idle sleep time — the docker exec calls stay the same.
Tried replacing the N-times-docker-exec polling loop with a single
blocking docker exec (in-container until/sleep loop). CI profiling
showed this was a net regression: docker exec connection overhead on
shared runners is ~0.7-1s per call, while the existing poll approach
uses 3-7 quick 0.27s execs. The blocking approach traded fewer calls
for longer per-call duration and lost.
Reverted to the original poll loop, added a docstring note explaining
why so the next person does not re-attempt the same optimization.
Add a module-scoped `shared_container` fixture to tests/docker/conftest.py
that boots one `sleep infinity` container per test module and tears it
down at module exit. Convert read-only tests that previously used
`docker run --rm --entrypoint sh/cat/test/su` (bypassing s6 to check
static image properties) or `docker run -d` + `docker exec` (starting
identical containers per test) to use `docker exec` on the shared
container instead.
Converted files:
test_immutable_install_permissions.py — 2 throwaway runs → 2 execs
test_license_file_present.py — 1 throwaway run → 1 exec
test_tini_compat_shim.py — 1 throwaway run → 1 exec
test_tui_prebuilt_bundle.py — 2 throwaway runs → 2 execs
test_dump_build_sha.py — 2 throwaway runs → 2 execs
test_immutable_install.py — 3 detached runs → 1 shared + 1 isolated
test_dashboard.py — 2 detached runs → 0 (use shared)
Local profiling shows docker run calls in these 7 files dropped from
~25 to 7 (the 7 are shared_container boots per module + the one test
that needs a restart). Each eliminated `docker run` was paying 1-9s
of s6 cont-init startup; the replacement `docker exec` calls average
0.10s — an ~50x speedup per operation.
Tests that mutate state (restarts, config changes, gateway starts)
still use their own containers via `container_name` + `start_container`.
Add a pytest profiling plugin (tests/docker/profiling.py) that
instruments every subprocess.run call targeting docker, collecting
per-test and per-subcommand wall-clock timings. Activated via
HERMES_DOCKER_TEST_PROFILE=1 — zero overhead when unset.
The CI docker.yml workflow now enables profiling and uploads the
JSON report as an artifact (docker-test-profile-<arch>), so we can
see exactly which docker operations (run, exec, restart, polling)
dominate the slow CI runs vs. fast local runs.
Output includes:
- Per-subcommand summary (call count, total/avg time, %)
- Top 10 slowest tests by docker operation time
- Top 10 slowest individual docker calls
- Full JSON report for offline analysis
PR #30136 review surfaced two issues, both rooted in the same audit gap:
docker integration tests were running as root, not the unprivileged
`hermes` user (UID 10000) that the runtime actually uses via
`s6-setuidgid hermes`. Anything that probed PID-1 state or wrote to
the s6 control surface worked as root in the tests but was inert in
production.
Fixes:
1. `_s6_running()` previously called `Path("/proc/1/exe").resolve()`,
which is root-only readable. For UID 10000 the symlink yields
PermissionError, `resolve()` silently returns the unresolved path,
and `exe.name == "exe"` — so detection always returned False, the
service-manager runtime-registration path was inert, and every
`hermes profile create` / `hermes -p X gateway start` silently
skipped the s6 hook. Replace with `/proc/1/comm` (world-readable)
+ `/run/s6/basedir` (s6-overlay-specific) — both required, fail
closed.
2. `02-reconcile-profiles` now also chowns `/run/service/.s6-svscan/`
{control,lock} to hermes so `s6-svscanctl -a/-an` works without
root. Previously the directory chown stopped at `/run/service`
and the FIFO inside stayed root-owned, so `register_profile_gateway`
from hermes failed at the rescan-trigger step with EACCES — the
wrapper in profiles.py caught the exception and printed a swallowed
warning, so profile creation appeared to succeed while the slot
was rolled back.
Audit changes to flush this class of bug next time:
- Add `docker_exec` / `docker_exec_sh` helpers to `tests/docker/conftest.py`
that default to `-u hermes`. The module docstring explains why and
flags `user="root"` as opt-in only for tests that explicitly need
root (none currently do).
- Refactor every `docker exec` call in tests/docker/ through the new
helpers (test_dashboard.py, test_zombie_reaping.py, test_profile_gateway.py,
test_container_restart.py, test_s6_profile_gateway_integration.py).
- Add 5 unit tests covering `_s6_running` under various probe states
(both signals present; comm wrong; basedir missing; PermissionError
on /proc/1/comm; missing /proc — non-Linux). The PermissionError
test is the explicit regression guard for the original bug.
Known follow-up: the per-service `supervise/control` FIFO inside each
`/run/service/gateway-<profile>/supervise/` is created root-owned by
s6-supervise (which runs as root because s6-svscan is PID 1). `s6-svc
-u/-d/-t` from the hermes user will get EACCES on those. The audit
under `-u hermes` will reveal this in lifecycle tests — surfacing the
issue cleanly so it can be fixed in a focused follow-up (likely via a
small SUID helper or a polling chown loop in cont-init.d). The
detection + svscanctl fixes here are independent and complete on
their own.
The agent-test suite default is 30s; docker test_no_args (the dashboard
spin-up, the container restart) routinely take 60-90s. Without this
they intermittently fail in CI with TimeoutError.
Task 0.1 of the s6-overlay supervision plan. Establishes the test
infrastructure for tests/docker/: skip-on-missing-Docker collection
hook, session-scoped image-build fixture (overridable via the
HERMES_TEST_IMAGE env var for faster local iteration), and a
container_name fixture that ensures cleanup on test exit.
Refs: docs/plans/2026-05-07-s6-overlay-dynamic-subagent-gateways.md