hermes-agent/apps/desktop
ethernet 68abf0b33d
test(desktop): add E2E coverage for session lifecycle (#69580)
* test(desktop): e2e test for interim assistant message preservation (#65919)

Adds a Playwright E2E test that reproduces the fix from PR #65919 across
all three layers (agent core → tui_gateway → desktop renderer). The mock
inference server is upgraded with a multi-turn scripted response that
exercises several interleaved patterns:

  1. text + tool_call  → should produce an interim message
  2. text + tool_call  → another interim message
  3. no text + tool_call → NO interim (no visible text alongside tools)
  4. text + tool_call  → another interim message
  5. final answer (stop) → message.complete, different from all interims

Two describe blocks exercise display.interim_assistant_messages both on
(default) and off:
  - ON:  all interim texts + the final answer visible in the transcript
  - OFF: only the final answer visible, all interim texts wiped

Also fixes a footgun: test:e2e now runs `npm run build` as a pretest
hook so the renderer dist/ is always fresh. Previously, running
`npx playwright test` locally would silently load a stale dist/ that
predated renderer fixes — the python backend ran from source (had the
fix) but the renderer was frozen in an old bundle. CI already built
fresh, so the explicit build step there is removed to avoid duplication.

* test(desktop): e2e sidebar states — background dot, subagent, cross-session

Add sidebar-states.spec.ts with three E2E tests exercising the desktop
sidebar's session dot states driven by real gateway events:

1. Background process dot appears during a terminal(background=true)
   call and disappears after auto-dismiss; subagent (delegate_task)
   runs concurrently; final answer is visible in the transcript.

2. Background dot remains visible while a subagent runs concurrently
   (longer sleep 5 background process so the dot is catchable).

3. Cross-session dot transition: start a turn with a background process,
   wait for the turn to complete, open a new session, then verify the
   original session's dot transitions from 'background running' to
   'finished — unread' when the background process exits.

The mock server gains SIDEBAR_SCRIPT and SIDEBAR_CROSS_SCRIPT trigger
keywords that return tool_calls for terminal(background=true) and
delegate_task — the agent executes these for real (real background
process, real subagent), so the tests assert against genuine gateway
events rather than mocked UI state.

Verified: 3 passed (1.2m) under cage headless wlroots.

* test(desktop): e2e tests for tile-unread bug (tab passes, split fails)

Two scenarios for the tile-unread bug where a session that finishes
while visible on-screen gets the green 'finished unread' dot even
though the user is looking right at it.

The unread check in handleTransition (session-states.ts:174) only
compares against $selectedStoredSessionId and ignores $sessionTiles,
so a session visible in a tile gets marked unread even though it's
on screen.

1. TAB (hidden, PASSES): ⌃-click opens the session as a stacked tab
   that is NOT visible on screen. The unread dot IS correct here —
   the user isn't looking at it.

2. SPLIT (visible, FAILS): drag the session row to the workspace's
   right edge to create a side-by-side split tile. Both sessions are
   visible on screen. The unread dot is WRONG — the session is visible
   in the split tile, so it should not be marked 'unread'. This test
   is RED until the fix lands.

Also adds explicit page.screenshot() calls at key assertion points in
sidebar-states.spec.ts so the trace viewer has full-res captures of the
sidebar dot states during the test.

* test(desktop): cover compression and queued stop lifecycle

Add real desktop E2E coverage for session compression continuation and
queue parking after an explicit Stop. Extend the mock server with a
blocking scripted turn and submitted-prompt assertions.

* test(desktop): cover busy composer submit routing

Replace the invalid queued-stop E2E scenario: plain text redirects a busy
turn rather than entering the queue. Add focused submit-routing coverage for
plain text, slash commands, attachments, explicit Stop, and idle submission.
2026-07-22 22:56:13 +00:00
..
assets
e2e test(desktop): add E2E coverage for session lifecycle (#69580) 2026-07-22 22:56:13 +00:00
electron fmt(js): npm run fix on merge (#69503) 2026-07-22 16:30:14 +00:00
pr-assets chore(desktop): drop PR screenshot assets from tree 2026-07-07 13:04:32 -07:00
public fix(providers): align Fireworks integration with project policy 2026-07-11 05:43:35 -07:00
scripts feat(desktop/e2e): Playwright E2E suite with visual regression diffs 2026-07-20 11:44:40 -04:00
src test(desktop): add E2E coverage for session lifecycle (#69580) 2026-07-22 22:56:13 +00:00
AGENTS.md fix(desktop): avoid false remote gateway reauthentication (#68250) 2026-07-20 20:54:36 -04:00
components.json feat(desktop): add shared project UI primitives 2026-06-25 16:40:27 -05:00
DESIGN.md feat(desktop): button tooltip keybind hints + keybinds settings tab + unified worktree dialog (#65204) 2026-07-16 18:26:21 -04:00
eslint.config.mjs refactor(lint): hoist shared eslint + prettier config to root 2026-07-16 01:42:02 +05:30
index.html feat(desktop): composer status stack, live subagent windows, editable prompts (#44630) 2026-06-12 08:30:06 -05:00
package.json test(desktop): add E2E coverage for session lifecycle (#69580) 2026-07-22 22:56:13 +00:00
playwright.config.ts fix(desktop): address review — overlay a11y, e2e typecheck, nits 2026-07-20 14:37:39 -04:00
preview-demo.html
README.md Merge pull request #63103 from NousResearch/bb/desktop-docs-alignment 2026-07-12 04:34:21 -05:00
tsconfig.e2e.json fix(desktop): address review — overlay a11y, e2e typecheck, nits 2026-07-20 14:37:39 -04:00
tsconfig.electron.json feat(desktop): ts-ify everything 2026-07-08 16:24:16 -07:00
tsconfig.json fix(desktop): minor type fixes and devShell cage dep 2026-07-20 11:44:40 -04:00
vite.config.ts feat(desktop): billing settings tab (#61054) 2026-07-18 19:38:02 +05:30
vitest.config.ts fix(desktop): bump skills test timeout to fix cold-start flake (#68235) 2026-07-20 21:27:34 +00:00
vitest.setup.ts test(desktop): widen Testing Library async deadline to de-flake UI panels (#67849) 2026-07-20 04:09:40 +00:00

Hermes Desktop ☤

Download Documentation Discord License: MIT

The native desktop app for Hermes Agent — the self-improving AI agent from Nous Research. Same agent, same skills, same memory as the CLI and gateway, in a polished native window — chat with streaming tool output, side-by-side previews, a file browser, voice, and settings, no terminal required. Available for macOS, Windows, and Linux.

Chat with the full agentStreaming responses, live tool activity, structured tool summaries, and the same conversation history as every other Hermes surface.
Side-by-side previewsRender web pages, files, and tool outputs in a right-hand pane while you keep chatting.
File browserExplore and preview the working directory without leaving the app.
VoiceTalk to Hermes and hear it back.
Settings & onboardingManage providers, models, tools, and credentials from a real UI. First-run setup gets you to your first message in seconds.
Stays currentBuilt-in updates pull the latest agent and rebuild the app in place.

Install

Already have the Hermes CLI? Just run:

hermes desktop

It builds and launches the GUI against your existing install — same config, keys, sessions, and skills. On first launch Hermes walks you through picking a provider and model; nothing else to configure.

Prebuilt installers

Prebuilt installers are built and distributed via the Hermes Desktop website..


Updating

The app checks for updates in the background and offers a one-click update when one is ready. You can also update any time from the CLI:

hermes update

Requirements

The installer handles everything for you (Python 3.11+, a portable Git, ripgrep).


Development

Want to hack on the app itself? Install workspace deps from the repo root once, then run the dev server from this directory:

npm install          # from repo root — links apps/desktop, web, apps/shared
cd apps/desktop
npm run dev          # Vite renderer + Electron, which boots the Python backend

Point the app at a specific source checkout, or sandbox it away from your real config:

# throwaway HERMES_HOME, separate Electron userData, distinct app name to avoid the single-instance lock
../scripts/dev-sandbox.sh npm run dev
HERMES_DESKTOP_HERMES_ROOT=/path/to/clone npm run dev
HERMES_HOME=/tmp/throwaway npm run dev
npm run dev:fake-boot   # exercise the startup overlay with deterministic delays

Building installers

npm run dist:mac     # DMG + zip
npm run dist:win     # NSIS + MSI
npm run dist:linux   # AppImage + deb + rpm
npm run pack         # unpacked app under release/ (no installer)

Installers are built and uploaded to GitHub Releases manually. macOS/Windows signing & notarization happen automatically when the relevant credentials are present in the environment (CSC_LINK / CSC_KEY_PASSWORD / APPLE_* for macOS, WIN_CSC_* for Windows).

How it works

The packaged app ships the Electron shell and a native React chat surface. On first launch it can install the Hermes Agent runtime into HERMES_HOME (~/.hermes, or %LOCALAPPDATA%\hermes on Windows), using the same layout as a CLI install.

The app has three boundaries:

  • Electron resolves and validates a runnable backend, owns native filesystem/git/window capabilities, and exposes a narrow preload bridge.
  • React owns the Desktop routes, panes, interaction state, and @assistant-ui/react transcript.
  • Hermes Agent runs as a headless hermes serve process and exposes the tui_gateway JSON-RPC/WebSocket API. The renderer connects through apps/shared, which is also used by the browser dashboard.

Backend resolution is an ordered ladder:

  1. HERMES_DESKTOP_HERMES_ROOT
  2. the current source checkout during development
  3. a completed managed install
  4. HERMES_DESKTOP_HERMES, or hermes on PATH
  5. a system Python that can import the Hermes runtime
  6. the first-launch bootstrap installer

Candidates are probed before use; an existing shim or interpreter is not enough. A runtime that predates serve falls back to headless dashboard --no-open. This is compatibility for the backend command only and does not launch or embed the dashboard UI.

The Electron orchestration entry point is electron/main.ts; pure resolution, probe, hardening, and platform policies live in focused modules beside it. The renderer is under src/, with shared atoms in src/store and transport/native adapters in src/lib.

Before changing the app, read:

  • AGENTS.md: architecture, state ownership, resolver/fallback, transport, performance, and testing rules.
  • DESIGN.md: visual system, information architecture, motion, direct manipulation, and keyboard behavior.

Connections, projects, and switching

Desktop supports a managed local backend, explicit remote gateways, and Hermes Cloud connections. Remote and cloud modes use the same remote-capability path; authentication and discovery differ, not the renderer feature model.

Projects are the workspace abstraction. A project may own multiple folders, repositories, worktrees, and sessions; a bare new chat remains detached unless the user enters a project or configures a default project directory. Use the Projects UI rather than adding a second per-session folder-picker workflow.

Changing profiles or connection modes is a soft workspace switch, not another cold boot. The shell and current management overlay remain mounted while gateway-bound nanostores are wiped, query-backed data is invalidated, and the new connection repopulates skeletons. This prevents rows or transcripts from the previous gateway bleeding into the next one.

Verification

Run before opening a PR (lint may surface pre-existing warnings but must exit cleanly):

npm run fix
npm run typecheck
npm run lint
npm run test:ui
npm run test:desktop:platforms

Run npm run test:desktop:all for install, boot, update, packaging, or other release-path changes.

Troubleshooting

Boot logs land in HERMES_HOME/logs/desktop.log (includes backend output and recent Python tracebacks) — check it first if the app reports a boot failure.

macOS / Linux:

# Force a clean first-launch setup
rm "$HOME/.hermes/hermes-agent/.hermes-bootstrap-complete"
# Rebuild a broken Python venv
rm -rf "$HOME/.hermes/hermes-agent/venv"
# Reset a stuck macOS microphone prompt (macOS only)
tccutil reset Microphone com.nousresearch.hermes

Windows (PowerShell):

# Force a clean first-launch setup
Remove-Item "$env:LOCALAPPDATA\hermes\hermes-agent\.hermes-bootstrap-complete"
# Rebuild a broken Python venv
Remove-Item -Recurse -Force "$env:LOCALAPPDATA\hermes\hermes-agent\venv"

The default Hermes home on Windows is %LOCALAPPDATA%\hermes. Set the HERMES_HOME env var if you've relocated it.


Community


License

MIT — see LICENSE.

Built by Nous Research.