mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-27 17:58:07 +00:00
Some OpenAI-compatible endpoints — notably Tencent Copilot (copilot.tencent.com) — only accept streaming chat requests; any non-streaming call returns HTTP 400 (code 11101, 'Non-stream chat request is currently not supported'). The main conversation loop already streams, so interactive chat works, but every auxiliary task (title generation, compression, web extraction) used the non-streaming path and failed on each call. _provider_requires_stream() detects stream-only endpoints (copilot.tencent.com built in, plus user-configurable auxiliary.stream_only_base_urls substring markers in config.yaml). Matching sync auxiliary calls route through _create_with_progress (force_stream=True) and async calls through the new _acreate_with_stream, aggregating the chunk stream — including tool-call deltas and reasoning deltas — into a complete response via the shared _ChatStreamAccumulator. Salvaged from PR #60686 by @kudi88 onto the progress-aware streaming machinery from #71508, addressing both sweeper-review gaps: the async path now consumes the stream with 'async for' (awaiting create() and iterating synchronously raised on AsyncOpenAI streams), and tool-call deltas are reassembled instead of dropped (MCP passes tools= through call_llm). Under force_stream there is no silent non-streaming retry — a stream-only provider rejects those by definition, so the original error surfaces to the normal recovery chains.
1 line
7 B
Text
1 line
7 B
Text
kudi88
|