mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-27 17:58:07 +00:00
Stop drip-feeding scenarios: extend the harness to cover the latencies that actually dominate perceived speed, and measure them on a REAL production build. - --prod: build a production renderer with the probe included (VITE_PERF_PROBE=1, off in normal builds) and launch it from dist/. Measures minified React, so numbers are representative shipped figures instead of ~3x-inflated dev ones. - cold-start scenario (tier "cold"): launch → CDP → driver → first paint, via a fresh isolated spawn per run. Captures spawn_to_cdp_ms, spawn_to_driver_ms, fcp_ms. - first-token scenario (backend tier): Enter → first assistant token painted — the TTFT latency an agent app is uniquely judged on. - run.mjs gained --prod (build once), cold-start fresh-spawn loop, and gates ci+cold tiers against the baseline. Baseline re-captured on a PRODUCTION build (median of 5), darwin-arm64 — all green. Representative numbers: cold-start spawn→interactive ~1.6s, FCP ~0.5s stream frame p95 22ms, 1 longtask keystroke p50 2ms, p95 8.7ms transcript mount 145ms, 82ms longtask (400-msg open) The prod build also settled the open question from the dev numbers: the transcript-mount "lead" (221ms longtask in dev) is only ~72-82ms in prod — not actionable. Measurement did its job.
58 lines
1.5 KiB
JSON
58 lines
1.5 KiB
JSON
{
|
|
"_meta": {
|
|
"note": "Median of 5 runs, darwin-arm64, `npm run perf -- cold-start stream keystroke transcript --spawn --prod` — a PRODUCTION renderer (minified React), so these are representative shipped numbers, not dev-inflated. Re-baseline per device with the same command + `--update-baseline`. Tolerances are loose to absorb cross-machine variance; cold-start especially varies with disk/OS state.",
|
|
"platform": "darwin-arm64",
|
|
"node": "v24.11.0",
|
|
"updated": "2026-07-19T21:38:19.701Z"
|
|
},
|
|
"scenarios": {
|
|
"stream": {
|
|
"tolerance": {
|
|
"tolFrac": 0.6,
|
|
"tolAbs": 5
|
|
},
|
|
"metrics": {
|
|
"longtasks_n": 1,
|
|
"longtask_max_ms": 67,
|
|
"frame_p95_ms": 22,
|
|
"frame_p99_ms": 23.7,
|
|
"slow_frames_33": 1,
|
|
"intermut_p95_ms": 36.1
|
|
}
|
|
},
|
|
"keystroke": {
|
|
"tolerance": {
|
|
"tolFrac": 0.6,
|
|
"tolAbs": 4
|
|
},
|
|
"metrics": {
|
|
"keystroke_p50_ms": 2.1,
|
|
"keystroke_p95_ms": 8.7,
|
|
"keystroke_p99_ms": 16.9,
|
|
"keystroke_slow_16": 2
|
|
}
|
|
},
|
|
"transcript": {
|
|
"tolerance": {
|
|
"tolFrac": 0.75,
|
|
"tolAbs": 40
|
|
},
|
|
"metrics": {
|
|
"transcript_mount_ms": 145,
|
|
"transcript_longtask_ms": 82,
|
|
"transcript_longtask_max_ms": 82
|
|
}
|
|
},
|
|
"cold-start": {
|
|
"tolerance": {
|
|
"tolFrac": 0.6,
|
|
"tolAbs": 150
|
|
},
|
|
"metrics": {
|
|
"spawn_to_cdp_ms": 326,
|
|
"spawn_to_driver_ms": 1647,
|
|
"fcp_ms": 500
|
|
}
|
|
}
|
|
}
|
|
}
|