FrontierHarness presents an evaluation comparing nine agent harnesses while holding the model and task constant, claiming a 17-fold spread in cost per successful pass. The cited earliest artifact appears to be a repository initially published as “Runta Eval,” but the supplied snippet is truncated and does not establish who operates FrontierHarness, the full methodology, or the exact experimental controls. Broader supplied results support the general proposition that harness and execution-environment design can materially affect agent cost and measured capability, but they do not independently verify this specific 17× result.
2026-09-04T13:37:46Z
Repeated discussion refreshes have produced no controlled cost, outcome, trace, or replication evidence, and no direct confirmation of the Astra gap. The broader harness-economics effect is established enough to retain, but this episode’s specific 17× magnitude has faded without validation and no confirming event is expected.
2026-09-04T12:32:09Z
The refreshed Armature discussion adds no controlled cost, outcome, trace, or replication evidence, so it is repetitive amplification rather than a change in meaning. Harness economics remains broadly corroborated, while FrontierHarness’s specific 17× magnitude and the reported Astra gap remain unverified.
2026-09-04T11:28:56Z
The refreshed Armature discussion adds no controlled cost, outcome, trace, or replication evidence; it is repetitive amplification of known tool-selection behavior. Harness economics remains broadly corroborated, while FrontierHarness’s specific 17× magnitude and the reported Astra gap remain unverified.
2026-09-04T10:30:00Z
The refreshed Armature and grep-versus-LSP comments add only anecdotal implementation observations, not controlled cost, outcome, or replication evidence. Harness economics remains broadly corroborated, while FrontierHarness’s specific 17× magnitude and the reported Astra gap remain unverified and stale.
2026-09-04T09:32:08Z
The refreshed Armature comments remain anecdotal discussion of agent tool preferences and add no measured outcome, cost, or replication evidence. Harness economics is still broadly corroborated, but neither FrontierHarness’s 17× magnitude nor the reported Astra gap has gained support.
2026-09-04T08:27:17Z
The refreshed comments add anecdotes about tool choice and token reduction but no controlled measurements, replication, or direct confirmation of the claimed 17× economics or Astra leaderboard gap. The broad harness-economics thesis remains corroborated, while the headline magnitudes remain unverified and stale.
2026-09-04T07:40:30Z
The refreshed discussions add no measurements, replication, or direct leaderboard confirmation; they only repeat known implementation observations. Harness choice remains a corroborated economic variable, but FrontierHarness’s 17× magnitude and the Astra gap remain unverified and no longer merit near-term attention.
2026-09-04T06:24:26Z
The brief confirmation window closed without direct ARC Prize leaderboard evidence, and refreshed discussion adds no measurements or replication. The broad harness-economics effect remains corroborated, but neither the Astra gap nor FrontierHarness’s 17× magnitude warrants continued high heat.
2026-09-04T05:23:22Z
Refreshed comments add no measurements, traces, or direct confirmation of either the 17× economics result or the reported Astra harness gap. The broad harness-economics thesis remains corroborated, with near-term heat sustained only by the pending ARC Prize leaderboard check.
2026-09-04T04:26:07Z
The grep-versus-LSP item suggests a plausible tool-selection mechanism but supplies no measurements, traces, or cost-per-success evidence, so it does not strengthen the specific 17× claim. The case remains hot only while direct ARC leaderboard confirmation of the Astra harness gap is pending.
2026-09-04T04:22:10Z
evidence attached: hn.story.49560260 — The report that coding agents prefer simple grep over richer language-server tools bears directly on how harness and tool design affect agent effectiveness and cost.
2026-09-04T03:28:55Z
The Astra report potentially broadens the case from inference economics to benchmark interpretation: harness, notes, and state management may account for a 62.7%–99.9% capability swing. Because the only supplied evidence is a low-engagement secondary post and the runs differ in reasoning level, the claimed scores and configurations need direct leaderboard confirmation before this becomes an escalation.
2026-09-04T03:22:33Z
evidence attached: reddit.post.1w6rwth — The observation provides relevant evidence that Astra’s headline benchmark results vary dramatically with harness, notes, and state-management configuration.
2026-09-04T02:28:27Z
The refreshed discussion adds no controlled outcome, cost, trace, or replication evidence; it is further amplification of the already-known tool-selection study. Harness economics remains broadly corroborated, but FrontierHarness’s specific 17× magnitude is still unverified.
2026-09-04T01:28:44Z
The refreshed Armature discussion adds no controlled outcome, cost, or replication evidence. The broad harness-economics effect remains corroborated, while FrontierHarness’s specific 17× magnitude remains unverified and increasingly stale.
2026-09-04T00:28:28Z
The refreshed Armature comments add no outcome, cost, or methodological evidence, so they do not strengthen the link between tool-selection behavior and inference economics. The broad harness-economics effect remains corroborated, while FrontierHarness’s specific 17× magnitude remains unverified.
2026-09-03T23:34:47Z
Refreshed comments add only implementation questions and speculation, with no controlled outcome, cost, or replication evidence. The broad harness-economics effect remains corroborated, but the specific 17× magnitude is still unverified and the discussion is repetitive.
2026-09-03T22:44:31Z
The 17,000-run Armature study adds scale evidence that coding agents’ tool-selection behavior varies across systems, but it does not establish effects on task outcomes or cost per successful pass. The broader harness-economics thesis remains corroborated while FrontierHarness’s specific 17× magnitude remains unverified.
2026-09-03T22:23:11Z
evidence attached: hn.story.49557206 — A 17,000-run comparison of tool selection provides useful independent evidence that coding-agent harness behavior materially affects outcomes and costs.
2026-09-03T20:36:15Z
The refreshed discussion adds no controlled replication, traces, or methodological evidence and does not change the distinction between a corroborated broad harness-economics effect and the unverified 17× headline magnitude. Repetitive commentary now lowers the case’s near-term temperature.
2026-09-03T18:55:44Z
The added whole-stack benchmark and working local-agent configuration broaden implementation-level support for treating harness and runtime as part of the evaluated system. They add no controlled cost-per-success comparison or replication of FrontierHarness’s 17× magnitude, so the case remains corroborated rather than accelerating.
2026-09-03T17:24:35Z
evidence attached: reddit.post.1w6c3ad — A verified local coding-agent setup and working multi-file build provide independent real-world evidence that the complete stack, not just the model, determines practical performance.
2026-09-03T17:24:35Z
evidence attached: reddit.post.1w6cjm5 — Its whole-stack, playtest-graded benchmark independently reinforces that runtime and harness configuration can determine real coding-agent outcomes beyond model benchmarks.
2026-09-03T16:46:35Z
Refreshed comments continue to identify useful measurement questions—failure modes, human rescue, caching, and realistic harness configuration—but add no replication, traces, or methodological disclosure. The broad harness-economics effect remains corroborated, while FrontierHarness’s specific 17× magnitude remains unverified.
2026-09-03T13:32:20Z
The independent TrueForge comparison corroborates the broader claim that harness design can materially change same-model task economics, moving the case beyond a single evaluator’s result. It does not validate FrontierHarness’s specific 17× magnitude, whose portability and controls remain unresolved.
2026-09-03T12:22:30Z
evidence attached: reddit.post.1w65ise — A concrete independent harness comparison reports similar solve rates with 63% fewer tokens and 30% lower cost, materially supporting harness choice as a first-order inference-economics variable.
2026-09-02T23:39:46Z
New discussion sharpens methodological concerns—especially harness customization and model-specific compatibility—which make the reported 17× spread less portable, but provides no replication or validation. The case remains a preliminary primary-source result rather than evidence of a general economic law.
2026-09-02T16:52:15Z
The refreshed discussion adds no methodological verification or independent replication; it is lightweight reaction rather than evidence for the specific 17× cost-per-pass claim. The case remains a relevant but preliminary primary-source result, with the nine-versus-twelve harness discrepancy and economic controls unresolved.
2026-09-02T16:29:52Z
grounded: converges/medium — The claimed same-model, same-task 17× cost-per-pass spread independently supports Scott’s Model-Plus-Harness Benchmark Unit and extends it into AI unit economic
2026-09-02T16:26:04Z
origin walked (codex/luna, conf 0.94): anchor hn.story.49538490 -> echo.github.da57499bd2 by Shiqi Mei (Runta)
2026-09-02T16:24:34Z
case created — The published evaluation makes a specific and consequential cross-harness cost claim distinct from existing harness-optimization cases.