HarnessOpt-Bench is a benchmark for testing whether frontier LLMs can iteratively improve another agent’s harness—including prompts, tools, control flow, memory, and orchestration code—under a fixed evaluation budget. The optimizer receives a seed harness and graded evaluation feedback, while its final candidate is scored by normalized improvement over the seed on an inaccessible held-out test split intended to distinguish genuine harness gains from overfitting or exploitation. The reported results say current frontier models can improve harnesses, but unevenly across tasks and often without enough separation for fine-grained model rankings. The supplied snippets associate the paper page with ScaleAI and host a PDF on Remotasks, but do not establish the individual authors or fully substantiate the stronger claim that the protocol prevents all grading-signal exploitation.
HarnessOpt-Bench independently operationalizes Scott’s positions that capability resides in the model-plus-harness system, harnesses can improve through evaluated recursive loops, and hidden acceptance tests are needed to limit specification gaming. Its uneven results and incompletely substantiated leakage protections also create a direct validation and critique opportunity for his evaluation-driven, trace-backed agent work rather than merely illustrating a familiar pattern.
ip:concept.model-plus-harness-benchmark-unitip:concept.self-improving-loopsip:framework.hidden-gates-frameworkip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisonradar:concept.recursive-self-improvementradar:concept.agent-evaluationradar:concept.benchmark-integrityradar:autodesign-meta-harness-optimization
queries asked of Scott's wikis
- agent harnesses as the real capability layer
- automated harness optimization and recursive improvement
- held-out evaluations for coding agents
- benchmark leakage and grading-signal exploitation
- stochastic agent evaluation under fixed budgets
- prompts tools memory orchestration as optimizable code
2026-10-01T01:26:37Z
Expired on neglect rather than outcome: five-plus weeks after launch, HarnessOpt-Bench itself has drawn no independent engagement (announcement posts at score 1, no replication, no critique of its anti-gaming protocol), while every new signal absorbed into this case has been adjacent harness tooling that never touches the benchmark. The surrounding harness-optimization wave is real but lives at the concept level — which stays independently tracked and hot — so any future replication or leakage critique of HarnessOpt-Bench specifically should arrive as a fresh signal, not a revival of this episode.
2026-09-23T17:56:08Z
The FrontierHarness Eval submission adds an adjacent performance claim, but its title alone establishes neither a released implementation nor a reproducible result, and it does not test HarnessOpt-Bench’s safeguards. Harness engineering remains active around this case without new evidence supporting this benchmark’s validity or urgency.
2026-09-22T21:23:03Z
evidence attached: hn.story.49808026 — A direct released-harness benchmark claim bears on whether harnesses can materially improve agent performance, though independent validation is still needed.
2026-09-13T02:30:08Z
RoastMyHarness adds a concrete adjacent comparison-tool lead: its submitter describes testing Pi customizations against bare Pi on DeepSWE tasks. This does not independently validate HarnessOpt-Bench’s optimization results or anti-gaming protocol, so the benchmark’s assessment remains unchanged.
2026-09-13T02:21:40Z
evidence attached: reddit.post.1weufjc — This released harness-comparison tool supplies practical evidence for measuring how skills, extensions, and agent configuration change coding results and cost.
2026-09-11T21:33:27Z
Beagle’s submitter now describes an initial release implementing DarwinX with verifier-backed tasks and regression safeguards, making it a more concrete adjacent tooling lead. This supplies neither independent validation of HarnessOpt-Bench nor evidence that verifier-backed optimization prevents test leakage or grading-signal exploitation, so the benchmark’s assessment remains unchanged.
2026-09-11T00:28:38Z
Beagle adds a relevant tooling lead for evaluating and evolving agent harnesses, but the supplied title provides no methods or results establishing its practical advantage. It broadens the surrounding engineering direction without validating HarnessOpt-Bench’s held-out measurements or anti-gaming protections.
2026-09-10T23:22:51Z
evidence attached: hn.story.49650988 — Beagle is independent corroborating tooling for evaluating and evolving agent harnesses at scale, directly strengthening the case’s benchmark-and-optimization hypothesis.
2026-09-10T05:24:51Z
AutoResearchExam adds another distinct effort to evaluate agent improvement and generalization, but its title-only evidence supplies neither a usable technical advance nor validation of HarnessOpt-Bench. The surrounding research direction is broadening; this benchmark’s held-out protocol and anti-gaming claims remain uncorroborated.
2026-09-10T05:22:04Z
evidence attached: hn.story.49638674 — The newly surfaced AutoResearchExam benchmark bears directly on measuring agents’ ability to improve and generalize, materially contextualizing evaluation of recursive agent optimization.
2026-09-09T20:38:35Z
Hyper–bench adds a distinct benchmark effort for agents that build agents, strengthening the broader direction but not independently validating HarnessOpt-Bench. The supplied title establishes no methods, results, or leakage protections, so the core hypothesis remains uncorroborated rather than advancing on adjacent activity.
2026-09-09T20:23:20Z
evidence attached: hn.story.49633138 — This is independent corroborating coverage of benchmark efforts for agents that build or optimize other agents.
2026-09-07T22:34:33Z
The new self-recursion post adds topical overlap but, with only a title supplied, establishes neither a technical advance nor independent validation of HarnessOpt-Bench. The benchmark’s held-out evaluation and anti-gaming claims remain open; adjacent harness activity should not be mistaken for corroboration.
2026-09-07T22:22:56Z
evidence attached: hn.story.49603308 — The post directly concerns optimizing agents that recursively improve their own harnesses, bearing on the open recursive-harness evaluation case.
2026-09-06T17:26:27Z
No new evidence changes the benchmark’s status: its claimed leakage protections and optimization measurements remain unvalidated, with the paper represented here only by reconstructed testimony. Adjacent harness projects support the broader engineering direction, not this benchmark’s reliability; further review should focus on artifact inspection or independent use.
2026-09-04T16:31:17Z
Kullback adds a second concrete implementation signal that trace-backed harness evaluation and optimization is becoming an engineering direction, rather than remaining only a benchmark proposal. Its sparse launch claim provides no results or leakage controls, so it still does not corroborate HarnessOpt-Bench’s held-out protocol or anti-gaming claims.
2026-09-04T16:23:06Z
evidence attached: hn.story.49565975 — The released harness turns agent traces into executable evaluation and RL environments, materially supporting the open question of practical infrastructure for recursive agent improvement.
2026-09-03T12:29:48Z
No artifact inspection, replication, or substantive critique has emerged; adjacent harness evidence supports the broader direction but still does not validate HarnessOpt-Bench’s protocol or results. The case remains worth monitoring at a slower cadence for technical uptake rather than repeated mentions.
2026-09-01T12:27:05Z
The reported 70-fold token-use spread adds independent support for the broader claim that harness choice materially affects agent economics, but the headline lacks enough methodological detail to interpret the difference. It does not validate HarnessOpt-Bench’s held-out protocol, anti-gaming protections, or measurements, so the core case remains uncorroborated.
2026-09-01T12:23:47Z
evidence attached: hn.story.49520778 — Reports a striking 70-fold token-use difference for the same model across agent harnesses, independently reinforcing that harness design materially affects inference economics.
2026-08-30T15:32:56Z
The case has produced no artifact inspection, replication, or independent validation of HarnessOpt-Bench’s anti-gaming protocol; AutoSaddler remains only adjacent evidence that harness optimization is a broader direction. The benchmark is still relevant but no longer warrants frequent review absent technical uptake or validation.
2026-08-28T14:40:03Z
AutoSaddler provides a thin independent signal that automated harness optimization is emerging as a broader engineering direction, moving the case beyond a lone proposal. It does not validate HarnessOpt-Bench’s held-out protocol, anti-gaming claims, or reported capability results.
2026-08-28T14:25:17Z
evidence attached: hn.story.49478099 — AutoSaddler is independent evidence that automatic optimization of agent harnesses is becoming a concrete research and engineering direction.
2026-08-27T21:42:02Z
The attached Reddit post is another author-led presentation of the same benchmark, not independent corroboration or artifact validation. The case remains a highly relevant primary-source proposal whose anti-leakage design and capability findings still need inspection or replication.
2026-08-27T21:24:34Z
evidence attached: reddit.post.1w052xg — The post independently surfaces the benchmark’s held-out evaluation and sandbox-isolation design, materially contextualizing the recursive harness-optimization case.
2026-08-27T20:43:46Z
No new evidence or independent validation has arrived; the benchmark remains a highly relevant primary-source proposal whose leakage protections and capability findings still need artifact inspection or replication. The unchanged re-observation does not sustain medium heat.
2026-08-27T20:32:34Z
grounded: converges/high — HarnessOpt-Bench independently operationalizes Scott’s positions that capability resides in the model-plus-harness system, harnesses can improve through evaluat
2026-08-27T20:29:13Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1w05763 -> echo.paper.8390da05aa by Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, and Yuan Xue
2026-08-27T20:27:21Z
case created — The observation describes a concrete benchmark and code release addressing a consequential agent-harness evaluation problem.