mirror of https://github.com/NousResearch/hermes-agent.git synced 2026-05-18 04:41:56 +00:00

docs: deep audit — fix stale config keys, missing commands, and registry drift (#22784 )

* docs: deep audit — fix stale config keys, missing commands, and registry drift

Cross-checked ~80 high-impact docs pages (getting-started, reference, top-level
user-guide, user-guide/features) against the live registries:

  hermes_cli/commands.py    COMMAND_REGISTRY (slash commands)
  hermes_cli/auth.py        PROVIDER_REGISTRY (providers)
  hermes_cli/config.py      DEFAULT_CONFIG (config keys)
  toolsets.py               TOOLSETS (toolsets)
  tools/registry.py         get_all_tool_names() (tools)
  python -m hermes_cli.main <subcmd> --help (CLI args)

reference/
- cli-commands.md: drop duplicate hermes fallback row + duplicate section,
  add stepfun/lmstudio to --provider enum, expand auth/mcp/curator subcommand
  lists to match --help output (status/logout/spotify, login, archive/prune/
  list-archived).
- slash-commands.md: add missing /sessions and /reload-skills entries +
  correct the cross-platform Notes line.
- tools-reference.md: drop bogus '68 tools' headline, drop fictional
  'browser-cdp toolset' (these tools live in 'browser' and are runtime-gated),
  add missing 'kanban' and 'video' toolset sections, fix MCP example to use
  the real mcp_<server>_<tool> prefix.
- toolsets-reference.md: list browser_cdp/browser_dialog inside the 'browser'
  row, add missing 'kanban' and 'video' toolset rows, drop the stale
  '38 tools' count for hermes-cli.
- profile-commands.md: add missing install/update/info subcommands, document
  fish completion.
- environment-variables.md: dedupe GMI_API_KEY/GMI_BASE_URL rows (kept the
  one with the correct gmi-serving.com default).
- faq.md: Anthropic/Google/OpenAI examples — direct providers exist (not just
  via OpenRouter), refresh the OpenAI model list.

getting-started/
- installation.md: PortableGit (not MinGit) is what the Windows installer
  fetches; document the 32-bit MinGit fallback.
- installation.md / termux.md: installer prefers .[termux-all] then falls
  back to .[termux].
- nix-setup.md: Python 3.12 (not 3.11), Node.js 22 (not 20); fix invalid
  'nix flake update --flake' invocation.
- updating.md: 'hermes backup restore --state pre-update' doesn't exist —
  point at the snapshot/quick-snapshot flow; correct config key
  'updates.pre_update_backup' (was 'update.backup').

user-guide/
- configuration.md: api_max_retries default 3 (not 2); display.runtime_footer
  is the real key (not display.runtime_metadata_footer); checkpoints defaults
  enabled=false / max_snapshots=20 (not true / 50).
- configuring-models.md: 'hermes model list' / 'hermes model set ...' don't
  exist — hermes model is interactive only.
- tui.md: busy_indicator -> tui_status_indicator with values
  kaomoji|emoji|unicode|ascii (not kawaii|minimal|dots|wings|none).
- security.md: SSH backend keys (TERMINAL_SSH_HOST/USER/KEY) live in .env,
  not config.yaml.
- windows-wsl-quickstart.md: there is no 'hermes api' subcommand — the
  OpenAI-compatible API server runs inside hermes gateway.

user-guide/features/
- computer-use.md: approvals.mode (not security.approval_level); fix broken
  ./browser-use.md link to ./browser.md.
- fallback-providers.md: top-level fallback_providers (not
  model.fallback_providers); the picker is subcommand-based, not modal.
- api-server.md: API_SERVER_* are env vars — write to per-profile .env,
  not 'hermes config set' which targets YAML.
- web-search.md: drop web_crawl as a registered tool (it isn't); deep-crawl
  modes are exposed through web_extract.
- kanban.md: failure_limit default is 2, not '~5'.
- plugins.md: drop hard-coded '33 providers' count.
- honcho.md: fix unclosed quote in echo HONCHO_API_KEY snippet; document
  that 'hermes honcho' subcommand is gated on memory.provider=honcho;
  reconcile subcommand list with actual --help output.
- memory-providers.md: legacy 'hermes honcho setup' redirect documented.

Verified via 'npm run build' — site builds cleanly; broken-link count went
from 149 to 146 (no regressions, fixed a few in passing).

* docs: round 2 audit fixes + regenerate skill catalogs

Follow-up to the previous commit on this branch:

Round 2 manual fixes:
- quickstart.md: KIMI_CODING_API_KEY mentioned alongside KIMI_API_KEY;
  voice-mode and ACP install commands rewritten — bare 'pip install ...'
  doesn't work for curl-installed setups (no pip on PATH, not in repo
  dir); replaced with 'cd ~/.hermes/hermes-agent && uv pip install -e
  ".[voice]"'. ACP already ships in [all] so the curl install includes it.
- cli.md / configuration.md: 'auxiliary.compression.model' shown as
  'google/gemini-3-flash-preview' (the doc's own claimed default);
  actual default is empty (= use main model). Reworded as 'leave empty
  (default) or pin a cheap model'.
- built-in-plugins.md: added the bundled 'kanban/dashboard' plugin row
  that was missing from the table.

Regenerated skill catalogs:
- ran website/scripts/generate-skill-docs.py to refresh all 163 per-skill
  pages and both reference catalogs (skills-catalog.md,
  optional-skills-catalog.md). This adds the entries that were genuinely
  missing — productivity/teams-meeting-pipeline (bundled),
  optional/finance/* (entire category — 7 skills:
  3-statement-model, comps-analysis, dcf-model, excel-author, lbo-model,
  merger-model, pptx-author), creative/hyperframes,
  creative/kanban-video-orchestrator, devops/watchers,
  productivity/shop-app, research/searxng-search,
  apple/macos-computer-use — and rewrites every other per-skill page from
  the current SKILL.md. Most diffs are tiny (one line of refreshed
  metadata).

Validation:
- 'npm run build' succeeded.
- Broken-link count moved 146 -> 155 — the +9 are zh-Hans translation
  shells that lag every newly-added skill page (pre-existing pattern).
  No regressions on any en/ page.

2026-05-09 13:19:51 -07:00

14 KiB

Raw Blame History

sidebar_position	title	description
99	Honcho Memory	AI-native persistent memory via Honcho — dialectic reasoning, multi-agent user modeling, and deep personalization

Honcho Memory

Honcho is an AI-native memory backend that adds dialectic reasoning and deep user modeling on top of Hermes's built-in memory system. Instead of simple key-value storage, Honcho maintains a running model of who the user is — their preferences, communication style, goals, and patterns — by reasoning about conversations after they happen.

:::info Honcho is a Memory Provider Plugin Honcho is integrated into the Memory Providers system. All features below are available through the unified memory provider interface. :::

What Honcho Adds

Capability	Built-in Memory	Honcho
Cross-session persistence	✔ File-based MEMORY.md/USER.md	✔ Server-side with API
User profile	✔ Manual agent curation	✔ Automatic dialectic reasoning
Session summary	—	✔ Session-scoped context injection
Multi-agent isolation	—	✔ Per-peer profile separation
Observation modes	—	✔ Unified or directional observation
Conclusions (derived insights)	—	✔ Server-side reasoning about patterns
Search across history	✔ FTS5 session search	✔ Semantic search over conclusions

Dialectic reasoning: After each conversation turn (gated by dialecticCadence), Honcho analyzes the exchange and derives insights about the user's preferences, habits, and goals. These accumulate over time, giving the agent a deepening understanding that goes beyond what the user explicitly stated. The dialectic supports multi-pass depth (1–3 passes) with automatic cold/warm prompt selection — cold start queries focus on general user facts while warm queries prioritize session-scoped context.

Session-scoped context: Base context now includes the session summary alongside the user representation and peer card. This gives the agent awareness of what has already been discussed in the current session, reducing repetition and enabling continuity.

Multi-agent profiles: When multiple Hermes instances talk to the same user (e.g., a coding assistant and a personal assistant), Honcho maintains separate "peer" profiles. Each peer sees only its own observations and conclusions, preventing cross-contamination of context.

Setup

hermes memory setup    # select "honcho" from the provider list

Or configure manually:

# ~/.hermes/config.yaml
memory:
  provider: honcho

echo 'HONCHO_API_KEY=***' >> ~/.hermes/.env

Get an API key at honcho.dev.

Architecture

Two-Layer Context Injection

Every turn (in hybrid or context mode), Honcho assembles two layers of context injected into the system prompt:

Base context — session summary, user representation, user peer card, AI self-representation, and AI identity card. Refreshed on contextCadence. This is the "who is this user" layer.
Dialectic supplement — LLM-synthesized reasoning about the user's current state and needs. Refreshed on dialecticCadence. This is the "what matters right now" layer.

Both layers are concatenated and truncated to the contextTokens budget (if set).

Cold/Warm Prompt Selection

The dialectic automatically selects between two prompt strategies:

Cold start (no base context yet): General query — "Who is this person? What are their preferences, goals, and working style?"
Warm session (base context exists): Session-scoped query — "Given what's been discussed in this session so far, what context about this user is most relevant?"

This happens automatically based on whether base context has been populated.

Three Orthogonal Config Knobs

Cost and depth are controlled by three independent knobs:

Knob	Controls	Default
`contextCadence`	Turns between `context()` API calls (base layer refresh)	`1`
`dialecticCadence`	Turns between `peer.chat()` LLM calls (dialectic layer refresh)	`2` (recommended 1–5)
`dialecticDepth`	Number of `.chat()` passes per dialectic invocation (1–3)	`1`

These are orthogonal — you can have frequent context refreshes with infrequent dialectic, or deep multi-pass dialectic at low frequency. Example: contextCadence: 1, dialecticCadence: 5, dialecticDepth: 2 refreshes base context every turn, runs dialectic every 5 turns, and each dialectic run makes 2 passes.

Dialectic Depth (Multi-Pass)

When dialecticDepth > 1, each dialectic invocation runs multiple .chat() passes:

Pass 0: Cold or warm prompt (see above)
Pass 1: Self-audit — identifies gaps in the initial assessment and synthesizes evidence from recent sessions
Pass 2: Reconciliation — checks for contradictions between prior passes and produces a final synthesis

Each pass uses a proportional reasoning level (lighter early passes, base level for the main pass). Override per-pass levels with dialecticDepthLevels — e.g., ["minimal", "medium", "high"] for a depth-3 run.

Passes bail out early if the prior pass returned strong signal (long, structured output), so depth 3 doesn't always mean 3 LLM calls.

Session-Start Prewarm

On session init, Honcho fires a dialectic call in the background at the full configured dialecticDepth and hands the result directly to turn 1's context assembly. A single-pass prewarm on a cold peer often returns thin output — multi-pass depth runs the audit/reconcile cycle before the user ever speaks. If prewarm hasn't landed by turn 1, turn 1 falls back to a synchronous call with a bounded timeout.

Query-Adaptive Reasoning Level

The auto-injected dialectic scales dialecticReasoningLevel by query length: +1 level at ≥120 chars, +2 at ≥400, clamped at reasoningLevelCap (default "high"). Disable with reasoningHeuristic: false to pin every auto call to dialecticReasoningLevel. Available levels: minimal, low, medium, high, max.

Configuration Options

Honcho is configured in ~/.honcho/config.json (global) or $HERMES_HOME/honcho.json (profile-local). The setup wizard handles this for you.

Full Config Reference

Key	Default	Description
`contextTokens`	`null` (uncapped)	Token budget for auto-injected context per turn. Set to an integer (e.g. 1200) to cap. Truncates at word boundaries
`contextCadence`	`1`	Minimum turns between `context()` API calls (base layer refresh)
`dialecticCadence`	`2`	Minimum turns between `peer.chat()` LLM calls (dialectic layer). Recommended 1–5. In `tools` mode, irrelevant — model calls explicitly
`dialecticDepth`	`1`	Number of `.chat()` passes per dialectic invocation. Clamped to 1–3
`dialecticDepthLevels`	`null`	Optional array of reasoning levels per pass, e.g. `["minimal", "low", "medium"]`. Overrides proportional defaults
`dialecticReasoningLevel`	`'low'`	Base reasoning level: `minimal`, `low`, `medium`, `high`, `max`
`dialecticDynamic`	`true`	When `true`, model can override reasoning level per-call via tool param
`dialecticMaxChars`	`600`	Max chars of dialectic result injected into system prompt
`recallMode`	`'hybrid'`	`hybrid` (auto-inject + tools), `context` (inject only), `tools` (tools only)
`writeFrequency`	`'async'`	When to flush messages: `async` (background thread), `turn` (sync), `session` (batch on end), or integer N
`saveMessages`	`true`	Whether to persist messages to Honcho API
`observationMode`	`'directional'`	`directional` (all on) or `unified` (shared pool). Override with `observation` object for granular control
`messageMaxChars`	`25000`	Max chars per message sent via `add_messages()`. Chunked if exceeded
`dialecticMaxInputChars`	`10000`	Max chars for dialectic query input to `peer.chat()`
`sessionStrategy`	`'per-directory'`	`per-directory`, `per-repo`, `per-session`, or `global`

Session strategy controls how Honcho sessions map to your work:

per-session — each hermes run gets a fresh session. Clean starts, memory via tools. Recommended for new users.
per-directory — one Honcho session per working directory. Context accumulates across runs.
per-repo — one session per git repository.
global — single session across all directories.

Recall mode controls how memory flows into conversations:

hybrid — context auto-injected into system prompt AND tools available (model decides when to query).
context — auto-injection only, tools hidden.
tools — tools only, no auto-injection. Agent must explicitly call honcho_reasoning, honcho_search, etc.

Settings per recall mode:

Setting	`hybrid`	`context`	`tools`
`writeFrequency`	flushes messages	flushes messages	flushes messages
`contextCadence`	gates base context refresh	gates base context refresh	irrelevant — no injection
`dialecticCadence`	gates auto LLM calls	gates auto LLM calls	irrelevant — model calls explicitly
`dialecticDepth`	multi-pass per invocation	multi-pass per invocation	irrelevant — model calls explicitly
`contextTokens`	caps injection	caps injection	irrelevant — no injection
`dialecticDynamic`	gates model override	N/A (no tools)	gates model override

In tools mode, the model is fully in control — it calls honcho_reasoning when it wants, at whatever reasoning_level it picks. Cadence and budget settings only apply to modes with auto-injection (hybrid and context).

Observation (Directional vs. Unified)

Honcho models a conversation as peers exchanging messages. Each peer has two observation toggles that map 1:1 to Honcho's SessionPeerConfig:

Toggle	Effect
`observeMe`	Honcho builds a representation of this peer from its own messages
`observeOthers`	This peer observes the other peer's messages (feeds cross-peer reasoning)

Two peers × two toggles = four flags. observationMode is a shorthand preset:

Preset	User flags	AI flags	Semantics
`"directional"` (default)	me: on, others: on	me: on, others: on	Full mutual observation. Enables cross-peer dialectic — "what does the AI know about the user, based on what the user said and the AI replied."
`"unified"`	me: on, others: off	me: off, others: on	Shared-pool semantics — the AI observes the user's messages only, the user peer only self-models. Single-observer pool.

Override the preset with an explicit observation block for per-peer control:

"observation": {
  "user": { "observeMe": true,  "observeOthers": true },
  "ai":   { "observeMe": true,  "observeOthers": false }
}

Common patterns:

Intent	Config
Full observation (most users)	`"observationMode": "directional"`
AI shouldn't re-model the user from its own replies	`"ai": {"observeMe": true, "observeOthers": false}`
Strong persona the AI peer shouldn't update from self-observation	`"ai": {"observeMe": false, "observeOthers": true}`

Server-side toggles set via the Honcho dashboard win over local defaults — Hermes syncs them back at session init.

Tools

When Honcho is active as the memory provider, five tools become available:

Tool	Purpose
`honcho_profile`	Read or update peer card — pass `card` (list of facts) to update, omit to read
`honcho_search`	Semantic search over context — raw excerpts, no LLM synthesis
`honcho_context`	Full session context — summary, representation, card, recent messages
`honcho_reasoning`	Synthesized answer from Honcho's LLM — pass `reasoning_level` (minimal/low/medium/high/max) to control depth
`honcho_conclude`	Create or delete conclusions — pass `conclusion` to create, `delete_id` to remove (PII only)

CLI Commands

The hermes honcho subcommand is only registered when Honcho is the active memory provider (memory.provider: honcho in config.yaml). Run hermes memory setup and pick Honcho first; the subcommand appears on the next invocation.

hermes honcho status          # Connection status, config, and key settings
hermes honcho setup           # Redirects to `hermes memory setup`
hermes honcho strategy        # Show or set session strategy (per-session/per-directory/per-repo/global)
hermes honcho peer            # Show or update peer names + dialectic reasoning level
hermes honcho mode            # Show or set recall mode (hybrid/context/tools)
hermes honcho tokens          # Show or set token budget for context and dialectic
hermes honcho identity        # Seed or show the AI peer's Honcho identity
hermes honcho sync            # Sync Honcho config to all existing profiles
hermes honcho peers           # Show peer identities across all profiles
hermes honcho sessions        # List known Honcho session mappings
hermes honcho map             # Map current directory to a Honcho session name
hermes honcho enable          # Enable Honcho for the active profile
hermes honcho disable         # Disable Honcho for the active profile
hermes honcho migrate         # Step-by-step migration guide from openclaw-honcho

Migrating from `hermes honcho`

If you previously used the standalone hermes honcho setup:

Your existing configuration (honcho.json or ~/.honcho/config.json) is preserved
Your server-side data (memories, conclusions, user profiles) is intact
Set memory.provider: honcho in config.yaml to reactivate

No re-login or re-setup needed. Run hermes memory setup and select "honcho" — the wizard detects your existing config.

Full Documentation

See Memory Providers — Honcho for the complete reference.

14 KiB Raw Blame History Unescape Escape