mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-27 17:58:07 +00:00
* feat(analytics): record auxiliary model usage per task in session accounting Auxiliary LLM calls (vision, compression, title_generation, web_extract, session_search, ...) discarded their token usage, leaving dashboard analytics blind to aux model spend (issue #23270). - hermes_state.py: session_model_usage gains a task PK dimension (''=main loop) via v22 table-rebuild migration (SQLite can't alter a PK); record_auxiliary_usage() writes per-(model,provider,task) deltas WITHOUT touching sessions counters (gateway overwrites those with absolute main-loop totals — folding aux in would double-count or be clobbered). Aux rows never inherit the session's main-loop route. - agent/aux_accounting.py: ContextVar ambient accounting context (mirrors the portal_tags conversation context); record_aux_usage() normalizes usage via usage_pricing.normalize_usage, estimates cost, and is strictly best-effort. moa_reference/moa_aggregator excluded — conversation_loop already folds MoA usage+cost into the main delta. - agent/auxiliary_client.py: _validate_llm_response is the recording chokepoint — every successful non-streaming aux response passes through it exactly once, sync and async, including fallback paths (model read from the response itself stays accurate across fallbacks). - run_agent.py: run_conversation publishes/resets the accounting context; agent/title_generator.py republishes on its bare thread. - hermes_cli/web_server.py: /api/analytics/usage folds aux rows into by_model (aux-only models finally appear) and adds a by_task summary; /api/analytics/models surfaces aux rows on the Models page. Design per review of PR #62850 by @eeksock (thread-local + separate auxiliary_usage table): rebuilt on ContextVar (async-safe — thread-local cross-attributes concurrent coroutines on one event loop) and the existing session_model_usage table instead of a parallel accounting path, extended beyond vision to every aux task, and wired the analytics endpoints so the dashboard actually shows it. Credit to @eeksock for the approach and @tboatman for the detailed root-cause analysis. * test(moa): match _validate_llm_response mock to new accounting-hint signature * test(aux): accept accounting-hint kwargs in remaining _validate_llm_response mocks |
||
|---|---|---|
| .. | ||
| test_aux_usage_accounting.py | ||
| test_conversation_root.py | ||
| test_get_anchored_view.py | ||
| test_get_messages_around.py | ||
| test_resolve_resume_session_id.py | ||
| test_session_archiving.py | ||
| test_session_md_export.py | ||