mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-30 19:09:28 +00:00
Replaces the half-duplex per-playback barge monitors with ONE listener that runs for the entire agent turn in continuous voice mode: armed at utterance-submit, disarmed when the turn is fully done (response + TTS finished). Fixes Teknium's live report that voice interruption never works: (a) not while the LLM is generating, (b) not while TTS plays. Root causes: - HALF-DUPLEX GAP: the barge monitor only spawned when TTS playback STARTED (cli.py streaming/whole-file paths, gateway _tts_stream_begin). During LLM generation there was NO microphone listener at all. - PLAYBACK DEAFNESS: the monitor calibrated its VAD noise floor WHILE the speaker was blasting TTS (speaker bleed baked into the floor), then multiplied it by 8x with a 1s strictly-consecutive block requirement — normal speech could rarely reach the trigger, and the 2s grace swallowed early interjections. New model — tools/voice_mode.full_duplex_listen(): - Pre-playback calibration: quiet-room noise floor established at turn start and HELD through playback (never recalibrated against bleed). - Phase-aware trigger: generation = floor x voice.barge_in_threshold_multiplier (new config, default 3.0, justified by synthetic-frame tests); playback = additionally clamped to a 1500-RMS minimum so bleed alone can't trip; 4000-RMS ceiling keeps speech always reachable. - Windowed-majority detection (>=80% of a 300ms window) instead of the strictly-consecutive counter that reset on intra-word energy dips. - Grace on playback ONSET only (voice.barge_in_grace_seconds, default down 2.0 -> 0.5) — suppresses the onset transient, not the mic. - Debug diagnostics at every decision point (calibrated floor, per-window RMS above 50% of trigger, trip/no-trip, grace suppressions) — always logger.debug, mirrored to stderr under HERMES_VOICE_DEBUG=1. Phase behavior (CLI cli.py + tui_gateway/server.py, same model): - generation: speech interrupts the in-flight turn via the SAME seam the typed/Ctrl+C interrupt uses (agent.interrupt()), cuts any pending TTS pipeline so the stale reply never plays, and submits the captured interjection (pre-roll capture, first syllable kept) as the next turn. - playback: cuts TTS (streaming pipeline stop + fallback speak stop events + file player) and submits the capture. - stop phrase honored in BOTH phases: mid-generation 'stop' interrupts the turn AND ends the voice chat (stop everything). - one listener instance spans generation -> playback (no re-arm race); double-arm refused (CLI _voice_fd_active / gateway _fd_listener_active). Gateway specifics: _arm_full_duplex_listener() at _run_prompt_submit turn start and inside _tts_stream_begin; _speak_text_with_barge registers its (stop, done) pair in _fd_speak_pipelines so fallback speaks are cut and tracked; _tts_stream_barge_in_monitor kept as a shim that arms the new listener. Desktop renderer owns its own mic path (voice-barge-in.ts) and is unaffected; if desktop backend-mic mode is used it inherits via the gateway. Tests: full_duplex_listen synthetic-RMS suite (speech-over-bleed trips, bleed alone doesn't, quiet floor held through playback, grace window, multiplier math 3x vs 8x, windowed-majority dips), CLI listener phase tests (generation interrupt seam, playback cut, lifecycle spans phases, double-arm, config forwarding, stop-phrase-mid-generation), gateway generation-phase interrupt + stop-phrase tests. Generation-interrupt test sabotage-verified. |
||
|---|---|---|
| .. | ||
| egress | ||
| features | ||
| messaging | ||
| secrets | ||
| skills | ||
| _category_.json | ||
| checkpoints-and-rollback.md | ||
| cli.md | ||
| configuration.md | ||
| configuring-models.md | ||
| desktop.md | ||
| docker.md | ||
| git-worktrees.md | ||
| import-from-other-agents.md | ||
| managed-scope.md | ||
| multi-profile-gateways.md | ||
| profile-distributions.md | ||
| profiles.md | ||
| security.md | ||
| sessions.md | ||
| tui.md | ||
| windows-native.md | ||
| windows-wsl-quickstart.md | ||