MirrorCode is a benchmark co-developed by Epoch AI and METR for testing whether coding agents can reimplement substantial existing software without seeing its source code, using detailed specifications and canonical outputs as checks. Preliminary results report that Anthropic’s Claude Opus 4.6 autonomously rebuilt gotree, a Go bioinformatics toolkit of roughly 16,000 lines and more than 40 commands, suggesting that some tightly specified coding projects can be sustained over task horizons associated with weeks of human work. The supplied material emphasizes that this is a particular, specification-rich setup and that broader implications remain uncertain; it does not establish that the result has yet been independently reproduced.
The radar already tracks this same development in `radar:mirrorcode-autonomous-project-scope`. Its eventual reproduction or failure would directly test Scott’s claims that substantial software can be regenerated from specification and behavioural oracles, while clarifying whether apparent weeks-long capability reflects the model alone or the complete agent system and evaluation setup.
ip:framework.ai-legacy-takeoverip:concept.specification-as-assetip:framework.long-running-agentsip:concept.benchmarking-the-wrong-unitradar:mirrorcode-autonomous-project-scoperadar:concept.software-reconstructionradar:concept.agent-evaluationradar:concept.long-horizon-agents
queries asked of Scott's wikis
- long-horizon coding-agent harnesses and autonomy
- executable specifications as agent verification oracles
- coding-agent evaluation beyond issue-resolution benchmarks
- agent task decomposition, recovery, and multi-session memory
- autonomous software reimplementation and clean-room workflows
- human-time estimates versus agent task-horizon measurement
2026-08-27T12:26:59Z
No independent reproduction, failure report, artifact, or methodological disclosure arrived across the full monitoring ladder. This does not disprove MirrorCode, but the active episode has faded and should reopen only on substantive validation evidence.
2026-08-25T11:30:37Z
The MW2 item establishes only that an adjacent month-long, high-token decompilation experiment exists; without results, artifacts, verification details, or autonomy disclosures, it is not yet an independent reproduction or capability finding. MirrorCode therefore remains a consequential but unvalidated claim awaiting substantive external evidence.
2026-08-25T11:24:05Z
evidence attached: hn.story.49431593 — Independent month-long evidence of agents spending 200B tokens on substantial software decompilation materially informs the open question of long-horizon autonomous coding performance.
2026-08-24T11:22:34Z
Another staleness-only review adds no reproduction, failure report, or methodological disclosure, so the validation question remains unchanged. Keep the case open but suppress routine revisits until substantive external evidence appears.
2026-08-22T10:28:51Z
The staleness threshold adds no substantive evidence and does not resolve a validation question whose natural horizon is longer than a few days. Keep the case open but defer further review until an independent reproduction, failure report, or meaningful methodology disclosure appears.
2026-08-20T09:40:08Z
The Reddit velocity spike is repetitive amplification of the original preliminary result, not independent reproduction or new methodological evidence. MirrorCode remains a consequential but unvalidated test of specification-driven long-horizon coding, with no basis for promotion or renewed attention.
2026-08-19T14:35:23Z
No independent reproduction, failure report, or methodological disclosure has appeared; the elapsed time and repeated engagement-only refreshes do not change the case’s meaning. Keep it open as a long-horizon validation watch, but cool and revisit only when substantive external evidence arrives.
2026-08-17T14:08:00Z
Refreshed discussion adds skepticism about specification cost, public reference implementations, and verifier dependence, but no independent reproduction or methodological finding. The case remains a preliminary, testable claim whose meaning still hinges on external validation.
2026-08-17T06:27:57Z
No independent reproduction, implementation, or methodological evidence has arrived; the slight engagement increase does not change the preliminary claim. The case remains a high-relevance but unresolved test of specification-driven, long-horizon coding agents.
2026-08-17T06:26:04Z
grounded: known/high — The radar already tracks this same development in `radar:mirrorcode-autonomous-project-scope`. Its eventual reproduction or failure would directly test Scott’s
2026-08-17T06:23:01Z
case created — A first-party benchmark artifact makes a concrete, independently testable claim about coding-agent performance on unusually long-horizon software reconstruction.