mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-23 16:36:23 +00:00
Opt-in, fully on-device wake word — the "Hey Siri" pattern — across all three
local surfaces (CLI, TUI, desktop GUI), with one configurable owner. Say the
phrase and Hermes opens a fresh session and starts hands-free voice; talk
back-and-forth; end it and the wake word re-arms. "Hey Hermes" works out of the
box (a trained model ships with Hermes). Off by default — nothing listens until
you turn it on.
Footprint: no new core model tool (config + a CLI command + gateway RPCs);
cache-safe (on wake we hand a transcript to the normal input path, never a
system-prompt/toolset mutation); .env only carries PORCUPINE_ACCESS_KEY.
- tools/wake_word.py: shared, engine-pluggable detector (openWakeWord default,
free/local; Porcupine premium) over the existing 16 kHz sounddevice capture.
Background daemon thread with pause()/resume() so it yields the mic during a
voice turn, and reset() on every (re)start so a resume can't re-fire on stale
audio. Ships a bundled "hey hermes" openWakeWord model (tools/wakewords/,
trained with the openWakeWord pipeline, Apache-2.0) as the default; a built-in
name or a custom .onnx/.tflite path still works. download_models() is called
for any model so a fresh install fetches the shared feature models (else it
crashed on a missing melspectrogram.onnx).
- wake_word.surface ("auto"|"cli"|"tui"|"gui") + wake_surface_enabled() gate so
exactly one surface owns the listener and the session it opens.
- CLI: in-process detector; on wake → new session + single-utterance capture via
the existing voice pipeline, with an idle watchdog that re-arms the mic.
/wake [on|off|status] command.
- TUI + desktop GUI: share the Python tui_gateway, which runs the detector
server-side and exposes wake.start/stop/pause/resume/status + a wake.detected
event (routed back over the same transport that armed it). Desktop arms it on
connect, opens a fresh session + starts voice on wake, and hands the mic
between the detector and its browser voice loop.
- config.yaml wake_word section; [wake] extra + uv.lock; packaging ships the
bundled model in wheel and sdist.
- Also: collapse ElevenLabs voice-list 401 log spam; treat an empty STT
transcript (silence) as a quiet re-listen, not a "transcription failed" toast.
- Tests (mocked, no live audio/network) + packaging guard + feature docs.
|
||
|---|---|---|
| .. | ||
| docs | ||
| i18n/zh-Hans/docusaurus-plugin-content-docs/current | ||
| scripts | ||
| src | ||
| static | ||
| .gitignore | ||
| docusaurus.config.ts | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| sidebars.ts | ||
| tsconfig.json | ||
Website
This website is built using Docusaurus, a modern static website generator.
Installation
yarn
Local Development
yarn start
This command starts a local development server and opens up a browser window. Most changes are reflected live without having to restart the server.
Build
yarn build
This command generates static content into the build directory and can be served using any static contents hosting service.
Deployment
Using SSH:
USE_SSH=true yarn deploy
Not using SSH:
GIT_USER=<Your GitHub username> yarn deploy
If you are using GitHub pages for hosting, this command is a convenient way to build the website and push to the gh-pages branch.
Diagram Linting
CI runs ascii-guard to lint docs for ASCII box diagrams. Use Mermaid (````mermaid`) or plain lists/tables instead of ASCII boxes to avoid CI failures.