mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-22 16:25:58 +00:00
Opt-in, fully on-device wake word — the "Hey Siri" pattern — across all three
local surfaces (CLI, TUI, desktop GUI), with one configurable owner. Say the
phrase and Hermes opens a fresh session and starts hands-free voice; talk
back-and-forth; end it and the wake word re-arms. "Hey Hermes" works out of the
box (a trained model ships with Hermes). Off by default — nothing listens until
you turn it on.
Footprint: no new core model tool (config + a CLI command + gateway RPCs);
cache-safe (on wake we hand a transcript to the normal input path, never a
system-prompt/toolset mutation); .env only carries PORCUPINE_ACCESS_KEY.
- tools/wake_word.py: shared, engine-pluggable detector (openWakeWord default,
free/local; Porcupine premium) over the existing 16 kHz sounddevice capture.
Background daemon thread with pause()/resume() so it yields the mic during a
voice turn, and reset() on every (re)start so a resume can't re-fire on stale
audio. Ships a bundled "hey hermes" openWakeWord model (tools/wakewords/,
trained with the openWakeWord pipeline, Apache-2.0) as the default; a built-in
name or a custom .onnx/.tflite path still works. download_models() is called
for any model so a fresh install fetches the shared feature models (else it
crashed on a missing melspectrogram.onnx).
- wake_word.surface ("auto"|"cli"|"tui"|"gui") + wake_surface_enabled() gate so
exactly one surface owns the listener and the session it opens.
- CLI: in-process detector; on wake → new session + single-utterance capture via
the existing voice pipeline, with an idle watchdog that re-arms the mic.
/wake [on|off|status] command.
- TUI + desktop GUI: share the Python tui_gateway, which runs the detector
server-side and exposes wake.start/stop/pause/resume/status + a wake.detected
event (routed back over the same transport that armed it). Desktop arms it on
connect, opens a fresh session + starts voice on wake, and hands the mic
between the detector and its browser voice loop.
- config.yaml wake_word section; [wake] extra + uv.lock; packaging ships the
bundled model in wheel and sdist.
- Also: collapse ElevenLabs voice-list 401 log spam; treat an empty STT
transcript (silence) as a quiet re-listen, not a "transcription failed" toast.
- Tests (mocked, no live audio/network) + packaging guard + feature docs.
|
||
|---|---|---|
| .. | ||
| features | ||
| messaging | ||
| secrets | ||
| skills | ||
| _category_.json | ||
| checkpoints-and-rollback.md | ||
| cli.md | ||
| configuration.md | ||
| configuring-models.md | ||
| desktop.md | ||
| docker.md | ||
| git-worktrees.md | ||
| managed-scope.md | ||
| multi-profile-gateways.md | ||
| profile-distributions.md | ||
| profiles.md | ||
| security.md | ||
| sessions.md | ||
| tui.md | ||
| windows-native.md | ||
| windows-wsl-quickstart.md | ||