mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-29 18:46:59 +00:00
The `text_to_speech` tool schema accepted only `text` and `output_path`, so style direction (tone, emotion, pacing, whispering) could never reach the OpenAI backend — even though `gpt-4o-mini-tts` (Hermes's OpenAI provider default) treats `instructions` as its primary voice-design control. This plumbs an optional `instructions` argument through the tool schema, the handler lambda, and `text_to_speech_tool()` into `_generate_openai_tts`, where it is forwarded to `client.audio.speech.create()` only when truthy. Empty/None values still omit the key entirely, preserving behavior on `tts-1`/`tts-1-hd` and strict OpenAI-compatible servers. The same passthrough unblocks self-hosted OpenAI-compatible voice-design servers (Qwen3-TTS-VoiceDesign on oMLX, etc.) that are already wired in via `tts.openai.base_url` — the established convention per #9004 and the TTS config docs — without inventing a new provider backend. Tests: `tests/tools/test_tts_instructions.py` covers backend passthrough, tool-level threading, schema declaration, and the empty-string/absent omission cases. `tests/tools/test_tts_max_text_length.py` fake_openai signature widened to accept the new kwarg. Refs NousResearch/hermes-agent#14196 |
||
|---|---|---|
| .. | ||
| developer-guide | ||
| getting-started | ||
| guides | ||
| integrations | ||
| reference | ||
| user-guide | ||
| index.mdx | ||
| user-stories.mdx | ||