hermes-agent/tools/wakewords
Brooklyn Nicholson 07648318ad feat(voice): "Hey Hermes" wake word to start a hands-free session
Opt-in, fully on-device wake word — the "Hey Siri" pattern — across all three
local surfaces (CLI, TUI, desktop GUI), with one configurable owner. Say the
phrase and Hermes opens a fresh session and starts hands-free voice; talk
back-and-forth; end it and the wake word re-arms. "Hey Hermes" works out of the
box (a trained model ships with Hermes). Off by default — nothing listens until
you turn it on.

Footprint: no new core model tool (config + a CLI command + gateway RPCs);
cache-safe (on wake we hand a transcript to the normal input path, never a
system-prompt/toolset mutation); .env only carries PORCUPINE_ACCESS_KEY.

- tools/wake_word.py: shared, engine-pluggable detector (openWakeWord default,
  free/local; Porcupine premium) over the existing 16 kHz sounddevice capture.
  Background daemon thread with pause()/resume() so it yields the mic during a
  voice turn, and reset() on every (re)start so a resume can't re-fire on stale
  audio. Ships a bundled "hey hermes" openWakeWord model (tools/wakewords/,
  trained with the openWakeWord pipeline, Apache-2.0) as the default; a built-in
  name or a custom .onnx/.tflite path still works. download_models() is called
  for any model so a fresh install fetches the shared feature models (else it
  crashed on a missing melspectrogram.onnx).
- wake_word.surface ("auto"|"cli"|"tui"|"gui") + wake_surface_enabled() gate so
  exactly one surface owns the listener and the session it opens.
- CLI: in-process detector; on wake → new session + single-utterance capture via
  the existing voice pipeline, with an idle watchdog that re-arms the mic.
  /wake [on|off|status] command.
- TUI + desktop GUI: share the Python tui_gateway, which runs the detector
  server-side and exposes wake.start/stop/pause/resume/status + a wake.detected
  event (routed back over the same transport that armed it). Desktop arms it on
  connect, opens a fresh session + starts voice on wake, and hands the mic
  between the detector and its browser voice loop.
- config.yaml wake_word section; [wake] extra + uv.lock; packaging ships the
  bundled model in wheel and sdist.
- Also: collapse ElevenLabs voice-list 401 log spam; treat an empty STT
  transcript (silence) as a quiet re-listen, not a "transcription failed" toast.
- Tests (mocked, no live audio/network) + packaging guard + feature docs.
2026-07-22 10:36:37 -05:00
..
hey_hermes.onnx feat(voice): "Hey Hermes" wake word to start a hands-free session 2026-07-22 10:36:37 -05:00
hey_hermes.tflite feat(voice): "Hey Hermes" wake word to start a hands-free session 2026-07-22 10:36:37 -05:00
README.md feat(voice): "Hey Hermes" wake word to start a hands-free session 2026-07-22 10:36:37 -05:00

Bundled wake-word models

hey_hermes.onnx / hey_hermes.tflite — the on-device "Hey Hermes" hotword model. This is the default detector for the wake word feature (see website/docs/user-guide/features/wake-word.md); no training or setup is required to say "hey hermes".

  • Engine: openWakeWord (Apache-2.0).
  • Provenance: trained with the openWakeWord training pipeline (synthetic TTS-generated speech), which produces both the .onnx and .tflite artifacts. Redistribution is permitted under the openWakeWord license.
  • Label: the model registers as hey_hermes (matches the filename).
  • Runtime: openWakeWord's shared feature-extraction models (melspectrogram + embedding) are NOT bundled here — they are fetched once on first use by tools/wake_word.py via openwakeword.utils.download_models().

To use a different phrase, train your own model and point wake_word.openwakeword.model at its path, or set a built-in openWakeWord name (hey_jarvis, alexa, hey_mycroft, …). See the wake-word docs for the training guide.