2026-10-11 18:00 UTC

Independent runs of Epoch AI’s MirrorCode evaluation will determine the maximum repository-scale software project current coding agents can complete with limited human intervention.

state: expiredheat: lowuncertainty: highconvergesscott: highcoding-agents agent-benchmarksEpoch AI

What is this?

MirrorCode is a long-horizon software-engineering benchmark built by Epoch AI and co-developed and funded with METR; one secondary snippet instead attributes a collaboration to Meta, but the supplied primary-adjacent evidence supports METR. It asks coding agents to autonomously reimplement target programs from observable behavior and tests, across 25 programs in areas such as Unix utilities, interpreters, bioinformatics, and cryptography. Preliminary results report that Claude Opus 4.6 rebuilt the roughly 16,000-line Go toolkit gotree, a task contributors estimated would take a skilled unaided engineer 2–17 weeks, though the snippets do not establish that independent replications have yet determined a general maximum project scale.

Why it matters to Scott

MirrorCode independently operationalizes several of Scott’s load-bearing positions: repository-scale capability should be measured through long-horizon execution, executable behavioral specifications, and observable test-based convergence rather than patch-level answers. Epoch AI’s benchmark creates a strong dated-receipts and validation opportunity for his long-running-agent architecture, although the supplied evidence does not yet show independent runs establishing a general maximum project scale.
ip:framework.long-running-agentsip:source.breaking-the-1hr-barrierip:concept.benchmarking-the-wrong-unitip:concept.characterisation-testingip:concept.measurable-convergencedev:concept.resumable-agent-job-control-planeradar:concept.coding-agentsradar:concept.long-horizon-agentsradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.benchmark-integrityradar:cursor-sqlite-doc-reconstruction
queries asked of Scott's wikis
  • repository-scale coding-agent evaluations
  • autonomous coding harnesses and human intervention
  • behavioral specifications versus source-code context
  • long-horizon agent reliability and inference scaling
  • benchmark validity for agent-completed software projects
  • measuring coding agents by task horizon

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (8) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnWhat's the largest software project AI can complete on its own?yusufozkan106105
🟧 echo.blog ⭐Introduces MirrorCode as an evaluation of the largest software projects AI agents can complete autonomously.Epoch AI——
🟧 hnAsk HN: I had Codex and GPT 5.6 Sol running for 12 days, 870k+ LOC. Now what?iagooar12
🟧 hnI Spent 2B Tokens Writing a C++ Compiler So You Don't Have Tolemming20
🟠 redditUpdate on the dating app I built with Claude Code: 30 agents run the day to day now, I still make every call, and I printed 350 posters myself
ClaudeAI
kyle77745016
🟧 hnLLM DeepSWE Pareto Frontierijidak10
🟧 hnNeal: Codex and Claude in a loop shipped a 549-commit migrationnavels10
🟠 redditI made Claude Souls
ClaudeAI
vrdrift13

Interpretation history

Decision trace