hermes-agent/apps/desktop/scripts/perf
Brooklyn Nicholson 45d4cf634d fix(desktop): time interaction frames on the clock that drives them
withFrames ran its own requestAnimationFrame ticker while the gesture body
independently awaited rAF per step. Two rAF consumers, so the observer's
deltas counted the driver's frames as well as the app's — it reported
~3fps for a drag that a single-clock probe measures at ~23fps, and it
never moved no matter what got fixed underneath.

Timing now comes from the same callbacks the body drives (__MARK__).

This also fixes a silent false-negative on the typing pass: it paced on
setTimeout, so the independent ticker was mostly sampling idle waits
between keystrokes and reported a flat 61fps. On the driving clock the
same interaction reports ~30fps with 27 of 40 frames over 33ms — which
matches the 'typing feels slow' symptom I previously could not reproduce.

TYPE now records __TYPE_TARGET__ and the runner throws when no composer is
found, so a pass that measures nothing fails loudly instead of scoring a
perfect 0 deficit — same guard DRAG already had.
2026-07-26 19:56:56 -05:00
..
lib bench(desktop): measure representative (warm-cache) cold start (#67733) 2026-07-19 23:21:45 +00:00
scenarios fix(desktop): time interaction frames on the clock that drives them 2026-07-26 19:56:56 -05:00
baseline.json bench(desktop): measure representative (warm-cache) cold start (#67733) 2026-07-19 23:21:45 +00:00
README.md feat(desktop): render-churn perf scenario 2026-07-26 06:11:00 -05:00
run.mjs bench(desktop): measure representative (warm-cache) cold start (#67733) 2026-07-19 23:21:45 +00:00
serve.mjs bench(desktop): systematized perf harness; sunset 12 one-off scripts (#67466) 2026-07-19 07:41:00 -04:00

Desktop perf harness

One systematized way to measure desktop rendering/interaction performance, diff it against a committed baseline, and fail on regressions. It replaces the dozen one-off measure-* / profile-* scripts that each reinvented the CDP client, arg parsing, stats, and output (and never had a baseline).

Quick start

# Isolated instance (recommended) — no running app or LLM credits needed.
# Its own --user-data-dir + HERMES_HOME means it never collides with `hgui`.
npm run perf -- --spawn

# Or: launch an isolated instance once, attach repeatedly (faster iteration).
npm run perf:serve            # leaves an instance on :9222
npm run perf                  # attaches, runs the CI suite, gates on baseline

# One scenario, with a CPU profile:
npm run perf -- stream --cpuprofile --tokens 800

# Representative PRODUCTION numbers (minified React, not the ~3x-slower dev build):
npm run perf -- cold-start stream keystroke transcript --spawn --prod

# Re-capture the baseline on your reference device, then commit baseline.json:
npm run perf -- cold-start stream keystroke transcript --spawn --prod --update-baseline

Dev vs prod

By default the harness measures the dev renderer (fast to spin up, good for relative regression checks). Pass --prod (with --spawn) to build a production renderer with the probe included (VITE_PERF_PROBE=1) and measure minified React — the representative shipped numbers. The committed baseline is captured with --prod.

Why isolation matters

The measurement this harness exists to run was historically blocked: a running hgui holds the Electron single-instance lock, so a second instance quit immediately. --spawn / perf:serve launch with their own --user-data-dir (separate lock scope), their own HERMES_HOME (separate backend + sessions), and their own --remote-debugging-port. Synthetic scenarios drive $messages directly via window.__PERF_DRIVE__, so no LLM credits are spent.

Scenarios

scenario tier measures replaces
stream ci streaming longtasks, frame p95/p99, mutation cadence measure-synthetic-stream, profile-synth-stream, profile-long-stream
stream --real backend same, from a real LLM stream measure-real-stream, profile-real-stream
keystroke ci composer keystroke → paint latency measure-latency, profile-typing, leak-typing
transcript ci large-transcript mount + paint cost (new)
render-churn ci per-component render attribution + store churn while N tabs stream (new)
cold-start cold launch → CDP → driver → first paint (fresh spawn/run) (new)
first-token backend Enter → first assistant token painted (TTFT) (new)
submit backend Enter → cleared → user msg painted, scroll jump measure-submit, measure-jump
session-switch backend route → first-paint → settle profile-session-switch
profile-switch backend rail click → sidebar settled measure-profile-switch

ci + cold scenarios need no backend/credits and are gated against baseline.json (cold-start requires --spawn since it measures a fresh launch, and must be run in its own invocation). backend scenarios need a live backend (and --spawn or a real session/credits) and are report-only.

CPU profiling is a cross-cutting --cpuprofile flag on any scenario (it wraps the run in Profiler.start/stop and prints a top-self-time table), replacing every standalone profile-* script.

Adding a scenario

Create scenarios/<name>.mjs exporting { name, tier, description, run(cdp, opts) } where run returns { metrics, detail } (metrics = flat numbers, lower is better), then register it in scenarios/index.mjs. If it's ci, add a baseline.json entry (or run --update-baseline).

Layout

  • lib/cdp.mjs — the one CDP client + target discovery + typing + CPU-profile wrapper + DOM selectors.
  • lib/stats.mjs — percentiles, histograms, CPU-profile self-time ranking.
  • lib/baseline.mjs — load/compare/update the baseline + regression gate.
  • lib/launch.mjs — attach, or spawn a fully isolated instance.
  • scenarios/ — one module per measurement.
  • run.mjs — entrypoint. serve.mjs — standalone isolated launcher.

Not migrated (kept as dev utilities)

eval.mjs, reload.mjs, reload-renderer.mjs, probe-renderer.mjs, probe-thread.mjs, click-session.mjs, diag-*.mjs are interactive dev helpers, not benchmarks. They can adopt lib/cdp.mjs in a follow-up.