mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-23 16:36:23 +00:00
/api/audio/speak-stream WebSocket: one socket + one Web Audio clock per reply. The renderer feeds raw LLM deltas as they arrive; the server cuts sentences with the shared chunker and streams int16 PCM back while generation continues — speech overlaps generation with no per-sentence connection or synthesis gaps. Falls back to the POST data-URL path for old backends / non-chunked providers. Barge-in runs a MediaRecorder on the monitor's stream the whole time playback is live (rotated while quiet to bound pre-roll); talking over the agent cuts playback and the complete utterance goes straight to transcription and submit. /api/audio/transcribe returns 200/"" for no-speech results so quiet turns re-listen instead of toasting a 400. |
||
|---|---|---|
| .. | ||
| bootstrap-installer | ||
| desktop | ||
| shared | ||