hermes-agent/apps/desktop/scripts/perf/scenarios/transcript.mjs
brooklyn! d1c455acf7
bench(desktop): systematized perf harness; sunset 12 one-off scripts (#67466)
Replaces the dozen ad-hoc measure-*/profile-* scripts (each reinventing the
CDP client — 4 different copies — plus its own arg parsing, stats, output
path, and none with a baseline) with one framework under scripts/perf/:

- lib/cdp.mjs      one CDP client + target discovery + typing + CPU-profile wrapper + DOM selectors
- lib/stats.mjs    percentiles, histograms, CPU-profile self-time ranking
- lib/baseline.mjs load/compare/update baseline + regression gate (new capability)
- lib/launch.mjs   attach, OR spawn a fully ISOLATED instance
- scenarios/*      one module per measurement, registered in scenarios/index.mjs
- run.mjs / serve.mjs, baseline.json, README.md

Isolation solves the long-standing measurement blocker: a running `hgui` held
the Electron single-instance lock, so a second instance quit. `--spawn` /
`perf:serve` launch with their own --user-data-dir (separate lock scope), their
own HERMES_HOME (separate backend/sessions, config seeded from ~/.hermes so it
reaches a chat view without onboarding), and their own --remote-debugging-port.
Synthetic scenarios drive $messages via window.__PERF_DRIVE__, so no LLM credits.

Scenario -> sunset script mapping:
  stream            <- measure-synthetic-stream, profile-synth-stream, profile-long-stream
  stream --real     <- measure-real-stream, profile-real-stream
  keystroke         <- measure-latency, profile-typing, leak-typing
  transcript        <- (new: long-transcript mount cost)
  submit            <- measure-submit, measure-jump
  session-switch    <- profile-session-switch
  profile-switch    <- measure-profile-switch
CPU profiling is now a cross-cutting --cpuprofile flag, not 5 separate scripts.

CI-tier scenarios (stream, keystroke, transcript) need no backend/credits and
are gated against baseline.json (seed values; re-capture with --update-baseline
on a reference device). Backend-tier scenarios are report-only.

perf-probe.tsx gains loadTranscript() for the transcript scenario. No core
files touched; isolation is via CLI args, not env-gated app changes.

Verified: node --check all modules, tsc, eslint, and a unit smoke of the
stats + regression-gate logic. The end-to-end GUI run (which opens a window)
is left to run interactively via `npm run perf -- --spawn`.
2026-07-19 07:41:00 -04:00

50 lines
1.7 KiB
JavaScript

// Large-transcript mount cost. New scenario (no prior script measured this):
// loads N synthetic turns of mixed markdown into $messages and records the
// mount→paint time plus any longtasks the mount blocks the main thread with.
// This is the "open a long session" path — a first-impression latency.
import { sleep } from '../lib/cdp.mjs'
const OBSERVE = `
(() => {
window.__TM__ = { longtasks: [] }
try {
const po = new PerformanceObserver((l) => {
for (const e of l.getEntries()) window.__TM__.longtasks.push(e.duration)
})
po.observe({ entryTypes: ['longtask'] })
window.__TM__.po = po
} catch {}
return 'observing'
})()
`
export default {
name: 'transcript',
tier: 'ci',
description: 'Mount + paint cost of loading a long transcript.',
async run(cdp, opts = {}) {
const turns = Number(opts.turns ?? 200)
await cdp.send('Runtime.enable')
await cdp.eval(OBSERVE)
const mountMs = await cdp.eval(`window.__PERF_DRIVE__.loadTranscript(${turns})`)
// Let post-mount longtasks (content-visibility passes, virtualizer) settle.
await sleep(1500)
const longtasks = await cdp.eval('window.__TM__.longtasks')
await cdp.eval('try { window.__TM__.po && window.__TM__.po.disconnect() } catch {}')
await cdp.eval('window.__PERF_DRIVE__.reset()')
return {
metrics: {
transcript_mount_ms: Math.round(mountMs * 10) / 10,
transcript_longtask_ms: Math.round(longtasks.reduce((a, b) => a + b, 0) * 10) / 10,
transcript_longtask_max_ms: Math.round((longtasks.length ? Math.max(...longtasks) : 0) * 10) / 10
},
detail: { turns, messages: turns * 2, longtasks: longtasks.length }
}
}
}