2026-10-11 18:02 UTC

Independent replication will determine whether AutoDesign’s meta-harness optimization reliably improves long-horizon agent design over manually engineered harnesses.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harnesses long-horizon-agents harness-optimization

What is this?

AutoDesign is presented as a system that uses a meta-harness to learn a DesignHarness for long-horizon academic-poster generation. Its report and derivative summaries claim that the learned harness raised the average PosterBench score from 54.99 to 67.39 across seven code-agent/model configurations, while an autonomous run used 253 tool calls and 11 editing turns in under 40 minutes for under $3. The supplied snippets do not identify the authors and do not establish independent replication: they report the authors’ results or discuss a separate Meta-Harness system, so AutoDesign’s reliability versus manually engineered harnesses remains unconfirmed here.

Why it matters to Scott

AutoDesign independently operationalizes Scott’s position that agent capability resides in the model-plus-harness system and that evaluated runs can breed improved scaffolding, closely converging with Replay-Driven Design Evolution and Reflexive Agent Design. The claimed cross-configuration gain could provide dated-receipts value, but absent independent replication or enough methodological detail, it does not yet validate reliability, portability, or superiority over manually engineered harnesses.
ip:concept.model-plus-harness-benchmark-unitip:framework.reflexive-agent-designip:framework.replay-driven-design-evolutionip:concept.self-improving-loopsdev:concept.trace-backed-agent-comparisonradar:evoharnessrl-self-evolving-agent-harnessradar:trained-harness-cross-model-transferradar:stencil-harness-coding-improvementradar:concept.agent-harnessesradar:concept.long-horizon-agents
queries asked of Scott's wikis
  • automated harness optimization vs hand-engineered agent harnesses
  • long-horizon agent harness design and evaluation
  • outer-loop optimization of coding-agent systems
  • agent execution traces as harness-learning memory
  • cross-model portability of learned agent harnesses
  • benchmark leakage and independent replication for agent systems

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Designmatt_d10
🟧 echo.paper ⭐The primary artifact is the authors’ arXiv technical report, submitted 13 August 2026. It introduces AutoDesign, where “a meta-harness optimYaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li——

Interpretation history

Decision trace