2026-10-11 17:10 UTC

Independent reproduction will determine whether MirrorCode’s detailed, checkable specifications enable frontier coding agents to autonomously reimplement substantial real-world software over multi-week-equivalent horizons.

state: expiredheat: lowuncertainty: highknownscott: highcoding-agents long-horizon-agents agent-evalsEpoch AIMETRAnthropic

What is this?

MirrorCode is a benchmark co-developed by Epoch AI and METR for testing whether coding agents can reimplement substantial existing software without seeing its source code, using detailed specifications and canonical outputs as checks. Preliminary results report that Anthropic’s Claude Opus 4.6 autonomously rebuilt gotree, a Go bioinformatics toolkit of roughly 16,000 lines and more than 40 commands, suggesting that some tightly specified coding projects can be sustained over task horizons associated with weeks of human work. The supplied material emphasizes that this is a particular, specification-rich setup and that broader implications remain uncertain; it does not establish that the result has yet been independently reproduced.

Why it matters to Scott

The radar already tracks this same development in `radar:mirrorcode-autonomous-project-scope`. Its eventual reproduction or failure would directly test Scott’s claims that substantial software can be regenerated from specification and behavioural oracles, while clarifying whether apparent weeks-long capability reflects the model alone or the complete agent system and evaluation setup.
ip:framework.ai-legacy-takeoverip:concept.specification-as-assetip:framework.long-running-agentsip:concept.benchmarking-the-wrong-unitradar:mirrorcode-autonomous-project-scoperadar:concept.software-reconstructionradar:concept.agent-evaluationradar:concept.long-horizon-agents
queries asked of Scott's wikis
  • long-horizon coding-agent harnesses and autonomy
  • executable specifications as agent verification oracles
  • coding-agent evaluation beyond issue-resolution benchmarks
  • agent task decomposition, recovery, and multi-session memory
  • autonomous software reimplementation and clean-room workflows
  • human-time estimates versus agent task-horizon measurement

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditMirrorCode: Evidence AI can already do some weeks-long coding tasks
singularity
141_13377117
🟧 echo.blog ⭐Preliminary MirrorCode results report that Claude Opus 4.6 autonomously reimplemented gotree, a bioinformatics toolkit with roughly 16,000 lEpoch AI and METR——
🟧 hn200B Tokens Later: A Month of Letting AI Agents Decompile MW2Gander573920

Interpretation history

Decision trace