mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-21 16:18:55 +00:00
Replaces the dozen ad-hoc measure-*/profile-* scripts (each reinventing the CDP client — 4 different copies — plus its own arg parsing, stats, output path, and none with a baseline) with one framework under scripts/perf/: - lib/cdp.mjs one CDP client + target discovery + typing + CPU-profile wrapper + DOM selectors - lib/stats.mjs percentiles, histograms, CPU-profile self-time ranking - lib/baseline.mjs load/compare/update baseline + regression gate (new capability) - lib/launch.mjs attach, OR spawn a fully ISOLATED instance - scenarios/* one module per measurement, registered in scenarios/index.mjs - run.mjs / serve.mjs, baseline.json, README.md Isolation solves the long-standing measurement blocker: a running `hgui` held the Electron single-instance lock, so a second instance quit. `--spawn` / `perf:serve` launch with their own --user-data-dir (separate lock scope), their own HERMES_HOME (separate backend/sessions, config seeded from ~/.hermes so it reaches a chat view without onboarding), and their own --remote-debugging-port. Synthetic scenarios drive $messages via window.__PERF_DRIVE__, so no LLM credits. Scenario -> sunset script mapping: stream <- measure-synthetic-stream, profile-synth-stream, profile-long-stream stream --real <- measure-real-stream, profile-real-stream keystroke <- measure-latency, profile-typing, leak-typing transcript <- (new: long-transcript mount cost) submit <- measure-submit, measure-jump session-switch <- profile-session-switch profile-switch <- measure-profile-switch CPU profiling is now a cross-cutting --cpuprofile flag, not 5 separate scripts. CI-tier scenarios (stream, keystroke, transcript) need no backend/credits and are gated against baseline.json (seed values; re-capture with --update-baseline on a reference device). Backend-tier scenarios are report-only. perf-probe.tsx gains loadTranscript() for the transcript scenario. No core files touched; isolation is via CLI args, not env-gated app changes. Verified: node --check all modules, tsc, eslint, and a unit smoke of the stats + regression-gate logic. The end-to-end GUI run (which opens a window) is left to run interactively via `npm run perf -- --spawn`.
38 lines
1.1 KiB
JSON
38 lines
1.1 KiB
JSON
{
|
|
"_meta": {
|
|
"note": "SEED baseline only. Values are approximations from apps/desktop/scripts/profile-typing-lag.md (May 2026, 34MB session, PRE incremental-lex). Run `npm run perf -- --update-baseline` on YOUR reference device to capture real numbers, then commit. Tolerances are intentionally loose until then.",
|
|
"platform": null,
|
|
"node": null,
|
|
"updated": null
|
|
},
|
|
"scenarios": {
|
|
"stream": {
|
|
"tolerance": { "tolFrac": 0.5, "tolAbs": 2 },
|
|
"metrics": {
|
|
"longtasks_n": 2,
|
|
"longtask_max_ms": 127,
|
|
"frame_p95_ms": 25.6,
|
|
"frame_p99_ms": 31.4,
|
|
"slow_frames_33": 6,
|
|
"intermut_p95_ms": 45
|
|
}
|
|
},
|
|
"keystroke": {
|
|
"tolerance": { "tolFrac": 0.5, "tolAbs": 3 },
|
|
"metrics": {
|
|
"keystroke_p50_ms": 8,
|
|
"keystroke_p95_ms": 17,
|
|
"keystroke_p99_ms": 28,
|
|
"keystroke_slow_16": 6
|
|
}
|
|
},
|
|
"transcript": {
|
|
"tolerance": { "tolFrac": 0.5, "tolAbs": 20 },
|
|
"metrics": {
|
|
"transcript_mount_ms": 450,
|
|
"transcript_longtask_ms": 350,
|
|
"transcript_longtask_max_ms": 160
|
|
}
|
|
}
|
|
}
|
|
}
|