'Clicked voice chat but it never starts listening' — a mic-device
contention race. Starting a voice conversation fires wake.pause (to free
the mic from the wake-word listener) and opens the conversation's own mic
in two separate effects, both keyed on voiceConversationActive with no
ordering between them. wake.pause was fire-and-forget, so getUserMedia
often raced the wake listener's stream teardown (which joins a reader
thread + closes the device in a finally). On Windows the capture device is
effectively single-owner, so opening it while wake still held it failed and
the conversation never started listening.
Wake-word-initiated starts don't hit this: the backend's _on_detect calls
pause_listening synchronously before emitting wake.detected, so the mic is
already free by the time the frontend opens it. Only the button/hotkey path
raced — matching the report.
Fix: pauseWakeForVoice now returns an awaitable barrier for the in-flight
wake.pause round-trip, and useVoiceConversation awaits a new beforeMicOpen
hook (wired to that barrier) right before handle.start(), re-checking
enabled/muted/busy/idle after the wait. No behavior change when the wake
word isn't running (barrier is null → no wait).
tsc + eslint clean, 39 desktop voice/wake tests green.
Diagnosed from a Windows desktop's logs: the backend TTS was generating
and saving the reply mp3 every turn (provider: openai), so synthesis was
fine — the DESKTOP wasn't playing the first reply. Root cause is Chromium's
autoplay policy: audio (HTMLAudioElement.play() and AudioContext) is
suspended until the frame sees a user gesture. A voice conversation started
by the 'Hey Hermes' wake word has no preceding click, so the first reply's
play() was rejected with NotAllowedError (silently swallowed) and only
turn 2+ spoke. Manual voice-start worked because the button click WAS the
gesture — exactly the 'first message in a new voice session is silent,
manual start is fine' report.
Fix (defense in depth):
- chatWindowWebPreferences sets autoplayPolicy: 'no-user-gesture-required'
on every chat window (primary + secondary), so a deliberately-launched
native app never gates audio on a gesture. Primary fix.
- voice-playback.ts resumes a suspended AudioContext on stream 'start' and
retries HTMLAudioElement.play() once via an unlock context after a
NotAllowedError — covers the dashboard-embedded surface that doesn't get
the Electron policy.
Tests: 16 electron session-window tests (incl. new autoplay assertion),
tsc + eslint clean.
The /api/audio/speak-stream WebSocket URL was always minted for the primary
backend with no profile param, so streaming read-aloud used the default
profile's TTS provider even when another profile was active. Mint the
ticket for the active profile's connection and carry ?profile= (same seam
as /api/pty), via a read-only getApiRequestProfile() accessor.
Generated media often arrives as image markdown (). The chat
markdown renderer maps the img element to MarkdownImage, which renders a raw
<img>; for a video/audio source the browser cannot paint it and shows a
broken-image icon even though the file is valid and plays fine. Images worked
only because <img src=...png> is valid markup.
Route video/audio sources in MarkdownImage to the existing MediaAttachment,
which already picks the correct <video>/<audio> element (streaming protocol +
open-externally fallback) by media kind. The intended MEDIA:-tag -> #media:
link path is unaffected; this only fixes the bare image-markdown case.
Since #57944 the image path is built on hooks (async src resolution, failure
and loading states), so the routing check cannot be a plain early return
inside it -- it would have to sit after every hook call and would still fire
an image resolve for media never rendered as an image. MarkdownImage is
therefore a thin hookless router in front of MarkdownImageContent, which is
the previous body unchanged.
The regression tests join the existing markdown-text.media.test.tsx added by
#57944 rather than replacing it.
Fixes#40896
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018vU45SEvcxLkVeyyEE9gJ2
transcribeAudio/speakText/getElevenLabsVoices called the backend without
profileScoped(), so for a non-default profile they hit the default backend
config instead of the active profile - playback used the wrong TTS/voice
even though settings were saved to the active profile config.
Fixes#53441
The grant_spent notice fired for every subscription user with top-up
funds the moment their cap was reached and camped in the CLI/TUI status
bar and desktop toasts with no action to take — the account keeps
working off top-up. Remove it everywhere:
- agent/credits_tracker.py: drop grant_cond + the emit/clear block;
the dev fixture state now (correctly) produces no notice
- TUI: keep the turn-start clear of credits.grant_spent as back-compat
for older backends that still emit the key
- Desktop: drop the demo step and stale comment references
- Docs/config comments: remove grant-spent from credits_notices text
- Tests updated: policy/cold-start now assert the key never fires
Usage bands, depleted, and restored notices are unchanged; /usage still
reports the full balance breakdown.
- Extract the cmdk filter into exported rankSearchOption and cover it,
SearchableSelect selection/clear/placeholder, and ConfigField
searchable-schema routing with 12 vitest cases (sabotage-verified:
the backend schema test fails on unfixed main).
- Add searchPlaceholder/noResults/systemDefault strings to ja/ar/zh-hant
(zh + en came with the salvaged commits; defineLocale would have
fallen back to English otherwise).
- Add a backend invariant test: timezone ships as a searchable,
clearable select of sorted IANA ids with a UTC fallback.
- Add 'System default' clear option to SearchableSelect via clearLabel prop
- Add clearable flag to ConfigFieldSchema (schema-driven, not hardcoded)
- Add clearable: true to timezone schema override in web_server.py
- Fix CommandItem value for clear item: use clearLabel instead of '' so
cmdk can match it during search
- Fix backend: or ['UTC'] fallback for hosts without tzdata where
available_timezones() returns an empty set (not an exception)
- Add systemDefault i18n key (en, types, zh)
- Fix focus: use autoFocus on CommandInput instead of e.preventDefault()
- Fix filter: prioritize city segment match (return 2 for slash match)
- Fix handleSelect: always select, don't toggle-deselect
- Add aria-haspopup to trigger button
- Use defensive placeholder logic
The timezone field was a free-text input with no guidance on format.
Users had to know the exact IANA identifier (e.g. America/New_York)
to configure it. Replace it with a searchable combobox built on
Popover + cmdk Command — the same stack as Shadcn's Combobox.
Backend:
- Add `_timezone_options()` to web_server.py (cached at import time,
returns sorted zoneinfo.available_timezones() — ~598 identifiers)
- Add `"timezone"` to `_SCHEMA_OVERRIDES` with `type: "select"`,
`options`, and `searchable: true`
Frontend:
- New `SearchableSelect` component (Popover + cmdk Command)
— closed-world filterable dropdown for large option lists
- `ConfigField` routes to `SearchableSelect` when
`schema.searchable === true` (explicit opt-in, no threshold)
- Add `searchable?: boolean` to `ConfigFieldSchema` type
- Add i18n keys: `searchPlaceholder`, `noResults`
The `searchable` flag is deterministic — no existing field is
affected unless explicitly opted in. Future large-list fields
can adopt the same pattern by adding `searchable: true` to their
schema override.
Saying 'stop' in a voice chat did nothing — the transcript was just
submitted to the agent as a normal turn, so the conversation never ended.
The only way out was the mouse/hotkey. That's not how a hands-free voice
assistant should work.
useVoiceConversation now checks each finished utterance against a spoken
stop-command matcher BEFORE submitting: 'stop', 'stop listening', 'never
mind', 'goodbye', 'cancel', 'that's all', etc., optionally addressed
('hey hermes, stop'). A match ends the conversation (flips enabled=false,
which drives the existing end() teardown — mic close, playback stop, wake
re-arm) instead of sending a turn.
Deliberately conservative: only a WHOLE-utterance stop phrase matches, so
substantive requests that merely contain 'stop' ('stop the docker
container', 'how do I stop a process') still go through.
- apps/desktop/src/lib/voice-stop-word.ts: isVoiceStopCommand() matcher
- use-voice-conversation.ts: onStopWord option + intercept before submit
- use-composer-voice.ts: wire onStopWord -> end the conversation
- 6 matcher tests (bare/multi-word/addressed stop; substantive requests
with 'stop' pass through; bare address words don't match)
- docs: note the spoken-stop behavior
A short, bright, rising two-note ding (G5 -> C6) plays the moment 'Hey
Hermes' is detected, before voice capture starts, so it's obvious the
wake registered. Deliberately distinct from the turn-end completion cue
(that one settles; this one rises = 'listening'). Reuses the same
lightweight WebAudio synthesis as completion-sound.ts — no asset to ship
— and is gated by the shared sound-mute toggle ($hapticsMuted), so
muting turn-end sounds silences it too.
- apps/desktop/src/lib/wake-sound.ts: playWakeSound()
- wiring.tsx: fire it at the top of the wake.detected handler
- 3 tests (plays two-note chime, silent when muted, never throws with
no WebAudio)
Teknium: the ear must always be visible so a user can click to enable
passive listening. If it can't start (missing STT/TTS, deps still
installing, no mic permission), the click surfaces the reason in the
tooltip and the toggle stays off — but the control never disappears.
Removes the 'if (!wake.available && !wake.enabled) return null' hide
branch that made the button vanish on machines where a requirements
probe returned false (the Windows report). The only non-idle state is
paused-for-voice (disabled, in the voice-chat pill), since an active
voice conversation genuinely holds the mic.
Tests updated: 'stays visible when unavailable and not enabled' and
'surfaces the refusal reason in the tooltip, still visible' replace the
old hide assertion. 30 vitest green, tsc + eslint clean.
The ear vanished whenever a voice conversation ran (the ConversationPill
replaces the whole controls row) and whenever a transient start refusal
marked the feature unavailable — so a persistent, config-backed setting
silently disappeared mid-session. The wake word is passive by design:
it should be visibly listening no matter what the GUI is doing, with
exactly one pause state — an active voice chat holding the mic.
- ConversationPill now renders the ear in paused form (disabled, EarOff,
'paused during voice chat' tooltip) so voice chat shows the listener
yielding the mic instead of the toggle vanishing.
- WakeWordButton hides only when the feature can't run AND isn't enabled
in config; $wakeWord gains 'enabled' (config truth from wake.status /
start/stop responses) so transient 'unavailable' refusals no longer
unmount the button.
- Busy agent turns never touched the listener (it keeps listening
through agent loops; wake.detected already opens a fresh session),
and now they can't hide the toggle either.
- New i18n key wakeWordPausedVoice across en/ja/zh/zh-hant.
Tests: ear mounted during busy turn, mounted through refusal when
config-enabled, hidden when unavailable+disabled, paused ear disabled
inside the pill. 29 vitest green across controls + wake-word store.
Internal testing (macOS): clicking the ear froze the button for ~30s,
then it went blue but 'hey hermes' never fired even though STT worked.
Two distinct bugs:
1. Frozen-then-timeout button: first-use wake.start lazy-installs the
detection engine (onnxruntime is a large wheel), which blows the
desktop's default 30s WS request timeout — the RPC 'failed' client-
side while the backend kept installing and armed later on its own.
wake.start now gets a 180s budget and the pending state says
'arming — first use may take a minute while the engine installs'
instead of a silent disabled button.
2. Armed but deaf: macOS grants mic permission per PROCESS. The
renderer having mic access (STT working) does not grant the Python
backend anything — CoreAudio hands an unentitled process a 'working'
stream that delivers zeros forever, so the listener looks healthy
and can never hear the phrase. The detector now tracks frame peaks:
10s of consecutive near-zero frames sets audio_silent, logged with
the exact macOS Settings path, cleared automatically when audio
appears. Surfaced everywhere: wake.status (audio_silent + hint),
desktop ear tooltip (kept visible while listening), /wake status on
TUI (⚠ line) and classic CLI, plus a docs troubleshooting section.
Tests: detector silent-flag set/recover cycle (fake silent/loud
streams), desktop tooltip keeps the dead-mic hint while listening.
44 wake Python tests, 22 desktop + 14 TUI vitest green.
Ending a voice conversation left the wake word silently off even with
wake_word.enabled: true — the desktop fired one wake.resume and hoped;
if the mic was still held by the just-released WebRTC capture (or the
resume raced teardown), the listener stayed dead until the user
re-toggled it. The wake word is a persistent setting: on is on until
the user explicitly turns it off.
- Desktop: resumeWakeAfterVoice() replaces the fire-and-forget resume —
resume, then verify against wake.status (config 'enabled' is the
authority) and re-arm via wake.start, with spaced retries to ride out
mic-release latency. Passive path: never passes persist, never writes
config; respects an explicit off and another surface's mic lease.
- Backend: wake.status now reports 'enabled' (config truth) so clients
reconcile against the setting, not runtime listener state.
- Backend: _wake_resume_if_owner self-heals — a resume that throws
(mic still busy) retries in a background thread for up to 15s. A
False return (lease gone/moved) is final, never retried, so the
retry can't steal another surface's mic. Covers the TUI/gateway
voice.record path which had no recovery at all (CLI has its idle
watchdog; the gateway had nothing).
- 6 new vitest cases: re-arm on enabled+down, no persist on the passive
path, resume-alone success, disabled stays off, owned lease yields,
older-backend no-op.
The wake-pause effect assigned wakePausedRef.current inside useEffect,
tripping eslint's no-restricted-syntax guard against atom→ref mirroring.
The ref is actually a request token (did WE issue wake.pause?), not a
reactive mirror — moving the assignment into a pauseWakeForVoice
callback keeps the semantics and passes the rule without a disable.
Clicking the desktop ear button or running /wake on|off now writes
wake_word.enabled to config.yaml (live, saved for future sessions), so the
feature no longer requires hand-editing config before the UI toggle works.
- wake.start accepts persist:true (explicit gesture): flips
wake_word.enabled on in config before arming; response reports
enabled_persisted. Passive auto-arm paths (desktop gateway-ready,
TUI reconnect) never pass it, so a mic can't become persistently
enabled without a deliberate user action.
- wake.stop accepts persist:true: writes wake_word.enabled: false so
auto-arm stays off next session; reports disabled_persisted.
- Split the refusal reason: 'disabled' (feature off in config — a
persisted gesture turns it on) vs 'disabled_for_surface' (explicit
wake_word.surface scoping, which persist does NOT override).
- Classic CLI /wake on|off and bare-toggle persist the flag too
(skips the write when config already matches).
- Desktop tooltip now maps refusal codes to friendly text (mirrors
the TUI's START_REASON_TEXT) instead of showing raw codes like
disabled_for_surface.
- Docs: quick-start notes the toggle persists; ear-button mention.
One sherpa listener now enrolls every wake-enabled profile's phrase
(defaulting to "hey <profile name>") and reports WHICH phrase fired.
wake.detected gains a profile field; the desktop live-switches to the
matching profile (same path as the profile rail), opens a fresh session
there, and starts hands-free voice. The single-profile CLI/TUI print the
hermes -p switch command for foreign-profile phrases instead of
answering as the wrong profile. Opt out per listener with
wake_word.profile_routing: false.
Ear icon next to the voice controls: highlighted while listening, muted
when off, hidden when the backend reports the wake word unavailable.
Backed by a feature-owned $wakeWord nanostore synced by both the button
(wake.start/stop) and the gateway-ready auto-arm (status-then-arm), so
the UI always reflects the real listener state; start refusals surface
their reason as a tooltip notice. i18n en/zh/zh-hant/ja.
start_new_session was respected only by the CLI; the TUI and desktop GUI
always opened a fresh session on wake, ignoring the config. The gateway
now carries the flag in the wake.detected payload and both clients honor
it (open a fresh session vs. continue the current one), matching the CLI.
Wake opened a fresh session but voice didn't start: the start intent was
a fire-once window CustomEvent, and the fresh-session remount tore down /
recreated the composer's subscription, so the deferred dispatch landed in
the gap and was lost.
Replace it with a latched nanostore ($voiceConversationStartRequest +
takeVoiceConversationStart): the controller sets it on wake.detected, and
the composer claims it once on (re)mount when the gateway is open, waiting
out any transient `disabled`. Drop the now-unused composer voice-start
window event.
Ending a voice conversation manually left the wake detector paused for
good, so the wake word couldn't be used again. The composer paused the
detector on voice start but only resumed on the voiceConversationActive
-> false render; if ending voice tore the composer down first, that
render never landed and the resume was skipped.
Resume on unmount as well (latched on wakePausedRef so it fires exactly
once), and stop early-returning when the $gateway atom is momentarily
null. Add wake.pause/resume INFO logs for visibility.
The GUI armed the detector (wake.start) and the gateway fired
wake.detected, but the desktop never reacted: detection was wired through
a side-registered gatewayRef.current.on('wake.detected', …) listener that
was instance/timing-fragile (and silently dead across reconnects/HMR),
even though the raw events were arriving on the socket.
Route wake.detected through handleGatewayEventWithWake — the same onEvent
pipeline every gateway socket already feeds via useGatewayBoot — and open
a fresh session + start back-and-forth voice there. Drop the separate
.on() listener; the open-effect now only arms wake.start.
On wake, the desktop GUI now opens a fresh session AND starts the
browser voice conversation (continuous, with TTS), matching the CLI/TUI
hands-free flow instead of just opening a session.
- Add an explicit requestVoiceStart() intent to the composer bus
(idempotent start; toggle could stop an active loop).
- Composer owns mic hand-off: pause the server-side wake detector while
the browser voice loop is live, resume after (server no-ops when the
wake word isn't armed) — via the $gateway store accessor.
- Controller fires startFreshSessionDraft() + requestVoiceStart() on
wake.detected.
Makes the wake word a tri-surface feature with one configurable owner.
- wake_word.surface ("auto" | "cli" | "tui" | "gui") + shared
wake_surface_enabled() gate consulted by every surface, so exactly one
place owns the listener and the new session it opens.
- tui_gateway: wake.start/stop/pause/resume/status RPCs + a wake.detected
event, sharing one server-side detector for both TUI and desktop. The
detector yields the mic to voice.record (pause on capture start, resume
on terminal) and to the desktop's browser mic (wake.pause/resume).
- TUI (Ink): arm wake.start on gateway.ready; on wake.detected open a
fresh session and start voice capture.
- Desktop (Electron): arm wake.start on connect; on wake.detected open a
fresh session.
- CLI now gates on wake_surface_enabled("cli"); /wake status shows surface.
- Tests for the surface gate; docs cover the surface knob + cross-surface.
Hard 16 MB cap on readFileDataUrl blocked larger local attaches with no way to raise it. Settings -> Chat now has a free-form MB field. Main process owns the persisted value and clamps only absurd inputs.
A session is meant to be either the main thread or a tile, never both.
openSessionTile enforced this from the tile side (refusing to tile the
selected session), but the main side never dropped an existing tile — so
any path that routes main to a session that's already tiled (cold-start
remembered-session restore, a pasted/Cmd-K route, a notification jump)
painted the same transcript twice: the workspace pane from the route and
the tile pane in parallel, both fighting one runtime.
resumeSession is the single chokepoint every load-into-main path funnels
through, so close the redundant tile there once the session becomes the
selection. The warm cache/runtime binding survives for main to reuse.
Mirror the picker dropdown's provider collapse into the Edit Models dialog so
curating is one click per provider instead of scrolling through 30 models.
Each provider header is a full-width clickable label (same style as the
composer context-menu labels) with a DisclosureCaret next to the text and a
select-all Checkbox (indeterminate when partial). Model rows stay Switches.
The dropdown's collapse is fixed too: the current provider is now
collapsible (was forced open), and the label style matches the rest of the app.
Adds a search icon to the dialog input matching every other search field.
The subpath entry existed only so the TUI could import the client-side
projection fallback. The TUI spawns its gateway from this same checkout and
cannot version-skew with it, so that fallback was dead weight — it reads
`display` directly now. With its last caller gone the dispatch helper folds
back into the desktop, where an older backend is genuinely reachable.
apps/shared/package.json is byte-identical to main again.
Add data-[state=indeterminate] styling (same primary fill as checked) and a
dash glyph that shows when the root is in the indeterminate state, so a
partial select-all reads as a dash rather than a full check.
expandProviderDefaults prefers a provider's featured_models when present and
falls back to the existing top-N for providers that ship none (single-lab,
local, custom). Only aggregators get curated, so exactly the providers with
the everything-under-the-sun problem are trimmed; the rest are unchanged.
Every non-featured model stays one search or Edit Models toggle away.