2026-10-11 17:11 UTC

HarnessOpt-Bench’s authors claim their held-out benchmark can measure whether frontier models improve other agents’ harnesses without exploiting test data or grading signals, enabling safer evaluation of recursive agent optimization.

state: expiredheat: lowuncertainty: highconvergesscott: highrecursive-self-improvement agent-harnesses evaluation-securityHarnessOpt-Bench

What is this?

HarnessOpt-Bench is a benchmark for testing whether frontier LLMs can iteratively improve another agent’s harness—including prompts, tools, control flow, memory, and orchestration code—under a fixed evaluation budget. The optimizer receives a seed harness and graded evaluation feedback, while its final candidate is scored by normalized improvement over the seed on an inaccessible held-out test split intended to distinguish genuine harness gains from overfitting or exploitation. The reported results say current frontier models can improve harnesses, but unevenly across tasks and often without enough separation for fine-grained model rankings. The supplied snippets associate the paper page with ScaleAI and host a PDF on Remotasks, but do not establish the individual authors or fully substantiate the stronger claim that the protocol prevents all grading-signal exploitation.

Why it matters to Scott

HarnessOpt-Bench independently operationalizes Scott’s positions that capability resides in the model-plus-harness system, harnesses can improve through evaluated recursive loops, and hidden acceptance tests are needed to limit specification gaming. Its uneven results and incompletely substantiated leakage protections also create a direct validation and critique opportunity for his evaluation-driven, trace-backed agent work rather than merely illustrating a familiar pattern.
ip:concept.model-plus-harness-benchmark-unitip:concept.self-improving-loopsip:framework.hidden-gates-frameworkip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisonradar:concept.recursive-self-improvementradar:concept.agent-evaluationradar:concept.benchmark-integrityradar:autodesign-meta-harness-optimization
queries asked of Scott's wikis
  • agent harnesses as the real capability layer
  • automated harness optimization and recursive improvement
  • held-out evaluations for coding agents
  • benchmark leakage and grading-signal exploitation
  • stochastic agent evaluation under fixed budgets
  • prompts tools memory orchestration as optimizable code

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

08-05 14:00⭐ origin echo-reconstructedThe primary paper introduces HarnessOpt-Bench, evaluating five frontier LLMs as harness optimizers across four tasks and 111 runs, using hel
Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, and Yuan Xue on paper (echo) · attributed from reddit.post.1w05763
—
08-27 20:13first on r/MachineLearning · published · +534.2hCan AI Improve Itself? RSI Might Be the Answer [R]
shehio
—
08-27 20:17first on r/artificial · published · +534.3hCan an AI make other AIs better? We benchmarked 5 frontier LLMs at rewriting other agents' harnesses, scored on a test set they never see (HarnessOpt-Bench, arXiv + MIT code)
shehio
—
08-28 13:21first on hacker news · published · +551.4hAutoSaddler: Automatic Harness Optimization
drseu55
—
09-13 01:40first on r/LocalLLaMA · published · +923.7hBenchmark your custom Pi tools
AnotherObsceneBean
—
08-27 20:13amplified on r/MachineLearningreddit.post.1w052xg
shehio
peak 1 · 2 comments · 3% of case engagement
08-27 20:17amplified on r/artificialreddit.post.1w05763
shehio
peak 1 · 0 comments · 1% of case engagement
08-28 13:21amplified on hacker news 👑hn.story.49478099
drseu55
peak 22 · 1 comments · 43% of case engagement
09-01 12:00amplified on hacker newshn.story.49520778
sunilkumardash9
peak 3 · 0 comments · 6% of case engagement
09-04 15:22amplified on hacker newshn.story.49565975
kkkamur
peak 1 · 0 comments · 2% of case engagement
09-07 21:47amplified on hacker newshn.story.49603308
alansaber
peak 1 · 0 comments · 2% of case engagement
5 more amplifiers in ainews.case_chain
08-27 20:20our radar first saw it · +534.4hdiscovery anchor: reddit.post.1w05763—

Evidence (12) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditCan an AI make other AIs better? We benchmarked 5 frontier LLMs at rewriting other agents' harnesses, scored on a test set they never see (HarnessOpt-Bench, arXiv + MIT code)
artificial
shehio10
🟧 echo.paper ⭐The primary paper introduces HarnessOpt-Bench, evaluating five frontier LLMs as harness optimizers across four tasks and 111 runs, using helVarun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, and Yuan Xue——
🟠 redditCan AI Improve Itself? RSI Might Be the Answer [R]
MachineLearning
shehio12
🟧 hnAutoSaddler: Automatic Harness Optimizationdrseu55221
🟧 hnHermes, Claude Code, and Codex ran an identical model. Token use varied 70-foldsunilkumardash930
🟧 hnShow HN: Kullback – Synthetic RL Environments from Traceskkkamur10
🟧 hnOptimising Harness Self-Recursionalansaber10
🟧 hnHyper–bench: Evaluating agents that build agentstosh20
🟧 hnAutoResearchExam: Measuring agents' ability to improve and generalizematt_d30
🟧 hnBeagle – a framework to evaluate and evolve agent harnesses at scaleqainsights11
🟠 redditBenchmark your custom Pi tools
LocalLLaMA
AnotherObsceneBean185
🟧 hnI beat all the popular harnesses on FrontierHarness Evaltontinton41

Interpretation history

Decision trace