Commit graph

17734 commits

Author SHA1 Message Date
ethernet
2d71e2f1e4 perf(tools): text prefilter before AST parse in tool discovery
`_module_registers_tools()` reads each `tools/*.py` file and fully
AST-parses it to check for a top-level `registry.register()` call.
90 files are scanned on every process start — but only 32 actually
register tools.

Add a cheap text prefilter: after reading the file (which we need to
do anyway for AST), check that both `"registry"` and `"register"`
appear in the source before calling `ast.parse`. A file with a
top-level `registry.register()` call must contain both strings, so
this is a perfect superset — zero false negatives. 50 of 90 files
skip the AST parse entirely.

The `source=` parameter is not threaded through `discover_builtin_tools`;
the prefilter lives entirely inside `_module_registers_tools`, keeping
the public API unchanged.

Benchmark (median of 10 runs, scanning 90 files):

  before (read + ast.parse all):  305.9ms
  after  (text prefilter + ast):   187.8ms
  speedup: 1.6x  (118ms saved)

Identical module set: 32 modules, same names, same order.
2026-07-13 18:14:15 -04:00
kshitijk4poor
658c011266 fix(deepinfra): restore provider-prefix aliases for model parsing
The _PROVIDER_PREFIXES frozenset in agent/model_metadata.py is static
and does not auto-extend from ProviderProfile. Removing deepinfra and
deep-infra from it broke provider:model prefix stripping for DeepInfra.
2026-07-14 03:34:25 +05:30
kshitijk4poor
ff52dce1fa test(cli): cover noted multimodal persistence handoff 2026-07-14 03:32:46 +05:30
kshitijk4poor
b708d10db0 test(session): type finalizer clean-history assertions 2026-07-14 03:32:46 +05:30
kshitijk4poor
8341d775a9 fix(session): restore clean API-local turn content 2026-07-14 03:32:46 +05:30
kshitijk4poor
32bdc67e10 fix(cli): snapshot close state under staging lock 2026-07-14 03:32:46 +05:30
kshitijk4poor
962189d9ea fix(cli): clear stale persistence override before staging 2026-07-14 03:32:46 +05:30
kshitijk4poor
50aebcbcff fix(session): preserve clean shortened close snapshots 2026-07-14 03:32:46 +05:30
kshitijk4poor
69fd846ef8 fix(session): serialize direct persistence flushes 2026-07-14 03:32:46 +05:30
kshitijk4poor
0b422559f3 fix(session): preserve clean multimodal persistence override 2026-07-14 03:32:45 +05:30
kshitijk4poor
a22a1079a3 fix(cli): preserve noted staged input on close 2026-07-14 03:32:45 +05:30
kshitijk4poor
475922f2ce fix(cli): serialize close persistence handoff
Preserve one durable staged input across terminal close and the worker's early turn flush, without duplicating resumed transcripts or creating a session with a null prompt. Fixes #63766.
2026-07-14 03:32:45 +05:30
kshitijk4poor
a27d51ef46 fix(cli): preserve resumed history during close flush
Retain a distinct CLI history baseline during the signal window before a turn's normal persistence flush. When CLI history aliases the live agent list, use marker-only persistence so a genuinely unflushed tail is written.
2026-07-14 03:32:45 +05:30
dsad
35ebf6ba67 fix(cli): persist close transcript without history alias 2026-07-14 03:32:45 +05:30
kshitijk4poor
ccb045ba7d fix(cron): resolve SessionDB timeout from config.yaml
Salvage of #63935. The original fix read HERMES_CRON_SESSION_DB_TIMEOUT
from a bare env var, but AGENTS.md requires non-secret behavioral
settings to live in config.yaml with an env var bridge only for
backward compatibility.

Changes:
- Add cron.session_db_timeout_seconds to DEFAULT_CONFIG (default 10s)
- Resolution order: HERMES_CRON_SESSION_DB_TIMEOUT env override →
  cron.session_db_timeout_seconds in config.yaml → 10s default
  (mirrors the existing script_timeout_seconds pattern)
- 0 = unlimited (opt-in for debugging, skips the bound)
- Strengthen test: assert the warning is logged on invalid env value
  (caplog was taken but never asserted)
- Add test: verify config.yaml resolution path works end-to-end

Co-authored-by: LoicHmh <26006141+LoicHmh@users.noreply.github.com>
2026-07-14 03:31:06 +05:30
Minhao HU
c675e7c793 fix(cron): bound SessionDB init so a hang can't wedge cron forever
run_job() constructs SessionDB() synchronously with no timeout of its
own, unlike the agent's run_conversation call further down, which is
already bounded by HERMES_CRON_TIMEOUT. A wedged sqlite3.connect (e.g.
a stale flock from a crashed sibling process) hangs this call
indefinitely.

That hang is invisible to every existing cron safeguard because it
happens before _submit_with_guard's future exists: the finally block
that discards the job ID from _running_job_ids never runs. The job
stays wedged "running" — every later tick logs "already running —
skipping" — until the whole gateway process is restarted.

Observed in production: a cron job's worker thread was confirmed via
a live py-spy thread dump to be parked inside SessionDB.__init__'s
sqlite3.connect for 3+ days, silently skipping every scheduled fire
in between across a gateway process that otherwise stayed healthy.

Bound the SessionDB() construction with its own timeout
(HERMES_CRON_SESSION_DB_TIMEOUT, default 10s), following the same
bounded-thread-pool pattern already used elsewhere in this file (the
delivery retry path, and the agent inactivity watchdog just below).
On timeout, log at ERROR and proceed with session_db=None instead of
degrading silently to debug level, since an actual hang here is a new
condition worth surfacing.

Adds tests/cron/test_sessiondb_init_hang.py, including an end-to-end
regression proving the dispatch guard is released and a subsequent
tick can fire the same job again after a simulated hang.
2026-07-14 03:31:06 +05:30
Brooklyn Nicholson
8bd4a419de fix(desktop): render reasoning text in the Thinking widget
The Thinking disclosure rendered blank for every reasoning-emitting model
(Fable, DeepSeek, GPT-5.5, ...). Two causes:

1. ReasoningTextPart read a `text` prop that assistant-ui never populates —
   reasoning parts arrive via context, same as text parts — so it always got
   an empty string. Read the text via useMessagePartReasoning() instead,
   mirroring how MarkdownText uses useMessagePartText().

2. The reasoning-only SmoothStreamingText / useSmoothReveal layer stalled at
   revealed="": the reasoning part stays isRunning for the whole message while
   the answer streams and thrashes re-renders, so the char-reveal never
   advanced past 0. Render reasoning through the same DeferStreamingText →
   surface path the assistant answer uses, and drop the dead smoothing code.
2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
d14bf23fa3 chore(desktop): build config — keep tsc emit out of src, gitignore artifacts 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
a57b33782e docs(desktop): hermes-desktop-plugins skill + starter template 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
aff205dcc2 chore(desktop): i18n strings for tabs, zones, and session menus 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
369d0eeeff refactor(desktop): retire desktop-controller for the contribution shell; views as contributions 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
10fbade64b feat(desktop): electron — openDir IPC + ⌘W menu bridge (tabs, not windows) 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
7f74b324c3 feat(desktop): store + lib — layout/preview/session atoms, escape-layers, keybind helpers 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
0f922002eb feat(desktop): contribution controller, surfaces, and wiring 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
0f398f8e9c feat(desktop): focused-session-aware titlebar + statusbar 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
2afbe77763 feat(desktop): session hooks — open-in-tile, per-session actions, resilient resume 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
860a3f67bb feat(desktop): chat view — drop overlays, composer scoping, tile integration 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
eae1d7d147 feat(desktop): ⌘W close-tab, ⌘⇧T reopen, ⌘T new tab, ⌘1-9 + ⌃Tab tab switching 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
ac4f596ca2 feat(desktop): pointer session drag/drop + row/tab menus with close others/right/all 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
e6fea77d14 feat(desktop): multi-session tiles — per-profile state, tile pane, pane mirror 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
f1379bd6c2 feat(desktop): routes, nav, and command palette as contributions 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
aae35c5ee4 feat(desktop): shared UI — per-session prompt overlays, gateway overlays, tab primitives 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
63a9bde77b feat(desktop): layout-tree renderer — splits, zones, pointer drag-session, tab strip 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
c388daa665 feat(desktop): layout-tree model + store + workspace geometry 2026-07-13 17:57:54 -04:00
Brooklyn Nicholson
7ed7170960 feat(desktop): plugin manager, runtime loader, and plugins settings 2026-07-13 17:57:53 -04:00
Brooklyn Nicholson
aefb36299a feat(desktop): contribution registry — namespaced areas, keybinds, palette 2026-07-13 17:57:53 -04:00
Brooklyn Nicholson
99ff67eb03 feat(desktop): plugin SDK surface — rest door, socket, react-query, UI kit 2026-07-13 17:57:53 -04:00
ethernet
29b8cacfab fix(ci): add missing Δ Wait column to skipped job rows 2026-07-13 17:46:11 -04:00
ethernet
e117478eb3 fix(ci): align baseline gantt bars to current job start
Baseline bars were positioned at their absolute timeline offset, which
included the baseline run's own wait time — making duration comparisons
hard since the bars were visually offset. Now each baseline bar starts
at the same left position as its corresponding current job, so the two
bars directly overlap for at-a-glance duration comparison.

Also removed the now-unused bl_t0 / bl_max / bl_jobs_timed variables
that tracked the baseline timeline position.
2026-07-13 17:46:11 -04:00
ethernet
f0b7cf3836 feat(ci): show per-job wait times in timing report
Add wait time computation: for each job, wait_s = started_at - max(completed_at
of all jobs that finished before it started). This is a timestamp heuristic
(no workflow YAML dependency parse needed) that's accurate for pipeline-shaped
CI where the critical path is linear at each stage.

Shows up in:
- Job table: new Wait + Δ Wait columns
- Gantt chart: hatched bar segment before the run bar
- Step details: '(wait Xs)' annotation in the summary line
- Markdown summary: Total wait row with delta vs baseline
- Stats: total_wait / bl_total_wait in compute_stats

Also completes the skipped-jobs UI:
- Skipped stat card in stats cards
- Skipped row in markdown summary table

Backward compatible: old cached baselines without wait_s annotate on load.
2026-07-13 17:46:11 -04:00
ethernet
7beca22bc0 fix(ci): exclude skipped jobs from timing deltas and stats
Skipped jobs (conclusion == 'skipped') have null/zero-duration timestamps
that polluted every downstream computation: they counted as 'unchanged'
(0 vs 0) in faster/slower tallies, showed meaningless '0.0s (0%)' deltas
in the job table, rendered phantom gantt bars, and inflated wall/compute
totals.

Add is_skipped() helper and apply it consistently:
- compute_stats: exclude from wall/compute + faster/slower/unchanged;
  add 'skipped' and 'bl_skipped' counts
- _gantt_bars: filter from current bars, baseline bars, and axis calc
- _job_table: show 'skipped' label instead of durations/deltas
- _step_details: skip entirely (no meaningful step data)
- _regressions: exclude from both current and baseline sides
2026-07-13 17:46:11 -04:00
brooklyn!
2a25d53ee5
Merge pull request #63995 from NousResearch/bb/salvage-63842-zoom-restore-sync
fix(desktop): sync UI Scale control after zoom restore (supersedes #63842)
2026-07-13 17:43:42 -04:00
Brooklyn Nicholson
39230d1738 refactor(desktop): funnel zoom apply+notify so restore can't desync
alelpoan's fix (emit hermes:zoom:changed after restore) is correct, but the
bug's root is duplication: setAndPersistZoomLevel and restorePersistedZoomLevel
each independently did setZoomLevel + send, and restore forgot the send.

Collapse both (and the lifecycle re-assert) into a single applyZoomLevel()
helper in zoom.ts that always applies-then-notifies — the regression can't
recur by omitting a send. Replace the source-grep test (which broke on main:
the sibling source-assertion pet test it copied was refactored to a behavioral
one, dropping the fs/path imports it relied on) with behavioral coverage of the
funnel, matching zoom.ts's "unit-testable without booting a BrowserWindow"
convention.

Co-authored-by: alelpoan <alelpoan@proton.me>
2026-07-13 17:39:03 -04:00
alelpoan
9f7a3cb1f6 test(desktop): cover restorePersistedZoomLevel renderer notification 2026-07-13 17:32:17 -04:00
alelpoan
a1a4c8ce1f fix(desktop): sync UI Scale setting after zoom restore on window load 2026-07-13 17:32:17 -04:00
kshitijk4poor
10dc1571bc fix(deepinfra): align refresh and TTS availability
Forward explicit catalog refreshes and make the TTS availability gate follow the configured provider instead of unrelated credentials.
2026-07-14 02:59:39 +05:30
kshitijk4poor
2fc3f9c1ff fix(deepinfra): harden multimodal provider routing
Prevent credential forwarding across catalog redirects, retain explicit opt-in semantics for paid media backends, fail closed on invalid provider configuration, avoid mixed-catalog and output-limit assumptions, and reserve native STT provider names.
2026-07-14 02:59:39 +05:30
Georgi Atsev
fe002eb124 feat(providers): Support DeepInfra as an LLM provider 2026-07-14 02:59:39 +05:30
brooklyn!
ed8ce1f96c
Merge pull request #63955 from NousResearch/bb/fix-windows-broken-login-bash
fix(windows): survive broken Git Bash login shells
2026-07-13 17:23:36 -04:00
ethernet
5d691374c3 fix(desktop): recognize little-endian Mach-O magic in native binary classifier
classifyNativeBinary only checked big-endian Mach-O/Fat magic bytes
(feedfacf, feedface, cafebabe). Real Darwin .node files from node-pty
prebuilds are stored little-endian on disk (cffaedfe = MH_CIGAM_64),
so every Darwin prebuild classified as null, and validateStagedBinaries
threw a platform mismatch on macOS — breaking npm run check for every
macOS contributor. CI didn't catch it because runners are Linux (ELF
path was correct) and tests only planted big-endian fake headers.

Add recognition for all six Mach-O/Fat byte orderings:
- MH_CIGAM (cefaedfe) — LE 32-bit
- MH_CIGAM_64 (cffaedfe) — LE 64-bit [the one real prebuilds use]
- FAT_CIGAM (bebafeca) — LE universal

Update makeFakeNode to write LE CIGAM_64 bytes for the darwin fixture
(matching real on-disk format) and add regression tests for all new
magic forms.
2026-07-13 17:22:17 -04:00