2026-10-11 17:16 UTC

FrontierHarness claims its nine-harness evaluation shows that harness choice can change cost per successful pass by 17-fold for the same model and task, making harness design a first-order driver of agent inference economics.

state: expiredheat: lowuncertainty: mediumconvergesscott: mediumagent-harnesses inference-economicsFrontierHarness
Surfaced 2026-09-04T03:28:55Z — priced heat=high at reprice: The Astra report potentially broadens the case from inference economics to benchmark interpretation: harness, notes, and state management may account for a 62.7%–99.9% capability swing. Because the only supplied evidence is a low-engagement secondary post and the runs differ in reasoning level, the claimed scores and configurations need direct leaderboard confirmation before this becomes an escalation.

What is this?

FrontierHarness presents an evaluation comparing nine agent harnesses while holding the model and task constant, claiming a 17-fold spread in cost per successful pass. The cited earliest artifact appears to be a repository initially published as “Runta Eval,” but the supplied snippet is truncated and does not establish who operates FrontierHarness, the full methodology, or the exact experimental controls. Broader supplied results support the general proposition that harness and execution-environment design can materially affect agent cost and measured capability, but they do not independently verify this specific 17× result.

Why it matters to Scott

The claimed same-model, same-task 17× cost-per-pass spread independently supports Scott’s Model-Plus-Harness Benchmark Unit and extends it into AI unit economics: harness choice may determine not only measured capability but the economically successful outcome rate. This creates a dated-receipts and benchmark-design opportunity, but the truncated artifact and unverified controls keep the result preliminary rather than decision-changing.
ip:concept.model-plus-harness-benchmark-unitip:concept.ai-unit-economicsdev:concept.trace-backed-agent-comparisonradar:concept.agent-harnessesradar:concept.inference-economicsradar:hidden-reasoning-real-task-costsradar:stencil-harness-coding-improvement
queries asked of Scott's wikis
  • agent harness as capability layer
  • cost per successful agent task
  • coding-agent harness economics
  • model capability versus scaffolding
  • agent evaluation harness controls
  • token efficiency and retry economics

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (8) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17xshiqimei8056
🟧 echo.github ⭐The earliest public primary artifact is the repository's initial publication, titled “Runta Eval,” which says: “We ran the same Kimi K3 modeShiqi Mei (Runta)——
🟠 redditWe built an open-source, model-neutral agent harness and compared it with claude managed agents - for the same model, got same accuracy, upto 75% lower cost
LocalLLaMA
Background-Job-8624343
🟠 redditNeon Ladder: a playtest-graded benchmark for local LLM stacks — your config is the subject, a working game is the grade
LocalLLaMA
stereohype11
🟠 redditRunning a local coding agent on Strix Halo with pi + llama.cpp: 27B and Flash-Next, the setup guide
LocalLLaMA
stereohype24
🟧 hnWhich tools do Claude, Codex and Cursor choose? We measured 17k runs to find outscrem285140
🟠 redditThe most interesting GPT-6 Astra result might be the 62.7% vs. 99.9% gap on ARC-AGI-3
OpenAI
HelenTKL92
🟧 hnGrep beats LSP? Why coding agents ignore your fancier toolskaonashi-tyc-017751

Interpretation history

Decision trace