hermes-agent/mcp-research-data
teknium1 d799184ede
bench: discovery-bound suite — paraphrase/absence/survey tasks isolate the listing's structural advantage
Bridge vs listing only (Opus 4.8, 830 real UE schemas, 3 reps/cell).
Excluding one both-modes mock artifact: listing 24/24 vs bridge 20/24,
searches/task 0.2 vs 4.0. Bridge failures: core-tool substitution at
frontier tier (ran the host test suite via terminal instead of
discovering RunTests, 2/3 reps), up to 8 searches to prove a negative,
and search-vocabulary misses on paraphrase. Listing asserts absence in
zero searches and answers a 5-way capability survey in 1 API call.
2026-07-22 05:16:30 -07:00
..
ue_bench_rows.json bench: Unreal-scale live benchmark — Epic's real 830 UE 5.8 schemas replayed (Opus 4.8) 2026-07-22 05:16:30 -07:00
ue_bench_summary.json bench: Unreal-scale live benchmark — Epic's real 830 UE 5.8 schemas replayed (Opus 4.8) 2026-07-22 05:16:30 -07:00
ue_discovery_rows.json bench: discovery-bound suite — paraphrase/absence/survey tasks isolate the listing's structural advantage 2026-07-22 05:16:30 -07:00
ue_hard_haiku_rows.json bench: adversarial 830-tool gauntlet — confusion clusters, type-aware error mocks, strict scoring 2026-07-22 05:16:30 -07:00
ue_hard_rows.json bench: adversarial 830-tool gauntlet — confusion clusters, type-aware error mocks, strict scoring 2026-07-22 05:16:30 -07:00