Salvaged from PR #47588 and rebased onto the post-campaign streaming core: the StreamingTTSProvider ABC/registry and the ElevenLabs/OpenAI streamers already live on main (tools/tts_streaming.py), so this ports the pieces main lacked: - GeminiStreamer: streamGenerateContent?alt=sse -> base64 PCM chunks (24 kHz mono int16), reusing main's DEFAULT_GEMINI_TTS_* constants. - XAIStreamer: WebSocket wss://api.x.ai/v1/tts -> binary PCM frames, async->sync bridged via the _collect_async test seam. - tts.streaming.provider config knob: pin one streamer, or 'auto' to walk the priority list elevenlabs -> gemini -> openai -> xai. Unset keeps the never-swap-the-user's-voice default. - docs/streaming-tts.md: architecture, capability matrix, how to add a provider. - Unit tests for the knob, SSE parsing, and the WS bridge; key-gated E2E tests (skipped without credentials). Refs: #47588
4.3 KiB
Streaming TTS
Hermes can stream TTS audio as it arrives from the provider, instead of waiting
for the full audio before playing. This is used by voice mode (CLI/TUI live
conversation), the dashboard speak-stream WebSocket, and — via the gateway
StreamingTTSConsumer — any platform adapter that opts into streaming audio.
Voice replies start speaking after the first clause instead of after full
generation + synthesis.
Architecture
The streaming pipeline has four parts:
- Producer — the LLM emits text deltas as it generates a response
- Sentence chunker —
tools.tts_streaming.SentenceChunkeraccumulates deltas, strips<think>blocks (even split across deltas), and flushes complete sentences - TTS provider — a registered
StreamingTTSProviderturns each sentence into raw PCM chunks (int16 mono at the provider's declaredsample_rate) - Audio sink —
sounddevice.OutputStreamfor local playback (tools.tts_tool.stream_tts_to_speaker), or a gateway platform adapter'swrite_streaming_ttsseam (gateway/streaming_tts_consumer.py)
Providers with no chunked API still get per-sentence playback via the proven
sync text_to_speech_tool path, so edge (the default) is conversational too.
All spoken text is cleaned by tools.tts_text_normalize.prepare_spoken_text
(one cleaner, all paths).
How to pick a provider
By default the dispatcher streams with the provider you already configured
(tts.provider) when that provider has a chunked API — it never silently
swaps your voice for a different provider just to get streaming.
To override, set tts.streaming.provider in your config.yaml:
- a provider name (
elevenlabs,gemini,openai,xai) pins that streamer autowalks the priority listelevenlabs → gemini → openai → xaiand uses the first one whose credentials resolve — an explicit opt-in to "best chunked voice available"
tts:
provider: gemini
streaming:
provider: gemini # or "auto"
gemini:
model: gemini-2.5-flash-preview-tts
voice: Kore
Capability matrix
| Provider | Transport | Chunked PCM | Credentials |
|---|---|---|---|
| elevenlabs | chunked HTTP (pcm_24000) |
yes | ELEVENLABS_API_KEY / tts.elevenlabs |
| openai | chunked HTTP (with_streaming_response, pcm) |
yes | tts.openai.api_key → env → managed gateway |
| gemini | SSE (streamGenerateContent?alt=sse) |
yes | GEMINI_API_KEY / GOOGLE_API_KEY |
| xai | WebSocket (wss://api.x.ai/v1/tts) |
yes | xAI OAuth or XAI_API_KEY |
| edge, piper, kitten, neutts, mistral, minimax, deepinfra, … | — | no (per-sentence sync fallback) | as usual |
All credential lookups go through resolve_provider_secret()
(config > env/.env > credential pool) — never bare env reads. Streamed bodies
are capped at 16 MiB per sentence, mirroring the sync providers' bounded
upstream-body invariant.
Adding a new streaming provider
- Subclass
StreamingTTSProviderintools/tts_streaming.py - Set
sample_rate(andchannels/sample_widthif not int16 mono) - Implement
available()(a pure probe — never install anything) andstream(self, text) -> Iterator[bytes]yielding raw PCM chunks - Decorate with
@register("yourname") - Add tests in
tests/tools/test_tts_streaming.py
The ABC enforces the contract; the registry makes the provider discoverable;
the dispatcher (stream_tts_to_speaker) and the gateway consumer handle the
sentence buffer, stop events, and audio sink for free.
Gateway streaming (platform adapters)
gateway/streaming_tts_consumer.py bridges agent deltas to an adapter's
streaming-audio seam. Adapters opt in by overriding, on
BasePlatformAdapter:
supports_streaming_tts(chat_id, audio_format) -> boolbegin_streaming_tts / write_streaming_tts / finish_streaming_tts / abort_streaming_tts
All default to unsupported/no-op, so existing adapters are untouched. When a turn's streaming audio completes, the whole-file auto-TTS reply for that turn is suppressed (no double playback); when streaming fails before any audio was audible, the gateway falls back to the legacy whole-file voice reply.