MirrorCode is a long-horizon software-engineering benchmark built by Epoch AI and co-developed and funded with METR; one secondary snippet instead attributes a collaboration to Meta, but the supplied primary-adjacent evidence supports METR. It asks coding agents to autonomously reimplement target programs from observable behavior and tests, across 25 programs in areas such as Unix utilities, interpreters, bioinformatics, and cryptography. Preliminary results report that Claude Opus 4.6 rebuilt the roughly 16,000-line Go toolkit gotree, a task contributors estimated would take a skilled unaided engineer 2–17 weeks, though the snippets do not establish that independent replications have yet determined a general maximum project scale.
MirrorCode independently operationalizes several of Scott’s load-bearing positions: repository-scale capability should be measured through long-horizon execution, executable behavioral specifications, and observable test-based convergence rather than patch-level answers. Epoch AI’s benchmark creates a strong dated-receipts and validation opportunity for his long-running-agent architecture, although the supplied evidence does not yet show independent runs establishing a general maximum project scale.
ip:framework.long-running-agentsip:source.breaking-the-1hr-barrierip:concept.benchmarking-the-wrong-unitip:concept.characterisation-testingip:concept.measurable-convergencedev:concept.resumable-agent-job-control-planeradar:concept.coding-agentsradar:concept.long-horizon-agentsradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.benchmark-integrityradar:cursor-sqlite-doc-reconstruction
queries asked of Scott's wikis
- repository-scale coding-agent evaluations
- autonomous coding harnesses and human intervention
- behavioral specifications versus source-code context
- long-horizon agent reliability and inference scaling
- benchmark validity for agent-completed software projects
- measuring coding agents by task horizon
2026-08-15T17:29:59Z
Repeated checks have produced only unvalidated project anecdotes, with no independent MirrorCode run or published result narrowing the autonomous project-scale boundary. Expire the active episode and reopen if substantive benchmark results or replication appear.
2026-08-13T16:35:50Z
The agent-built native game adds another concrete but unvalidated project anecdote; without inspectable artifacts, intervention logs, or independent evaluation, it does not establish limited-intervention completion or narrow MirrorCode’s project-scale boundary. The case remains cold pending an independent MirrorCode run or substantive published benchmark results.
2026-08-13T16:23:51Z
evidence attached: reddit.post.1vndp67 — A concrete agent-built native game project offers weak but relevant evidence about the repository-scale software projects coding agents can complete.
2026-08-11T16:44:46Z
Repeated staleness adds no independent MirrorCode run or validated limited-intervention completion; practical anecdotes still do not establish an autonomous project-scale boundary, so the case remains cold pending published results.
2026-08-09T16:34:28Z
The staleness review and minor engagement increase add no independent MirrorCode run, validated limited-intervention completion, or evidence of an autonomous project-scale boundary. Keep the case cold while awaiting benchmark replication or substantive published results.
2026-08-07T16:26:18Z
The 549-commit migration further supports the practical plausibility of sustained repository-scale agent work, but the sparse report does not establish project complexity, quality, completeness, or limited human intervention. It is neither an independent MirrorCode run nor evidence of a maximum autonomous project scale, so the case remains cold pending validated benchmark results.
2026-08-07T16:21:39Z
evidence attached: hn.story.49212155 — A concrete 549-commit migration adds independent evidence about coding agents completing substantial repository-scale work.
2026-08-06T20:31:27Z
The added comments offer more practitioner anecdotes about large AI-assisted projects, but they still involve substantial human architecture or exhibit uncontrolled agent drift. No independent MirrorCode run or validated limited-intervention completion has appeared, so the case’s meaning remains unchanged and cold pending benchmark replication.
2026-08-06T19:22:08Z
The sparse DeepSWE frontier listing concerns capability and efficiency but does not test long-horizon repository completion, limited human intervention, or independently replicate MirrorCode. The case remains a promising benchmark hypothesis without evidence establishing an autonomous project-scale boundary.
2026-08-06T19:21:32Z
evidence attached: hn.story.49200656 — A DeepSWE Pareto-frontier result may provide relevant evidence about current coding-agent capability and efficiency, though the sparse listing limits confidence.
2026-08-05T09:25:20Z
No substantive evidence has appeared since the last review: the minor engagement increase adds neither an independent MirrorCode run nor a validated limited-intervention completion. Keep the case cold until replication or benchmark results establish an autonomous project-scale boundary.
2026-08-04T23:28:18Z
The latest activity adds no identifiable independent MirrorCode run or validated autonomous project completion; discussion continues to amplify repository-scale anecdotes without establishing limited-intervention capability or an upper bound. Repeated non-substantive updates now warrant cooling the case until replication or benchmark results appear.
2026-08-04T22:26:18Z
The 30-agent deployment adds another independent indication that repository-scale agent work is operationally feasible, but its creator still makes every consequential decision, underscoring that project size and agent count do not establish limited-intervention autonomy. No independent MirrorCode run or validated upper bound has appeared, so rising discussion increases attention without advancing maturity.
2026-08-04T22:21:31Z
evidence attached: reddit.post.1vfogeb — A real deployment reportedly has Claude Code coordinating 30 agents and maintaining a 70k-line project, materially contextualizing the scale and human-oversight limits of autonomous coding workflows.
2026-08-04T12:25:25Z
The compiler report adds a second independent anecdote that very large token budgets can sustain repository-scale agent work, but it provides no validation of completeness, quality, or limited human intervention. It therefore broadens practical plausibility without establishing MirrorCode’s maximum-project claim or resolving its generalization limits.
2026-08-04T12:21:35Z
evidence attached: hn.story.49167134 — Independent real-world report of a coding agent spending 2B tokens building a substantial compiler materially informs the frontier of autonomous project scope.
2026-08-04T01:22:00Z
No substantive new evidence is identifiable: the attached material still provides neither an independent MirrorCode run nor a validated completed-project result. The case remains promising but unchanged, with amplification outpacing evidence that could establish a maximum autonomous project scale.
2026-08-04T00:24:27Z
The latest change is only marginal engagement on existing discussion; no independent MirrorCode run or validated project completion has appeared. The case remains relevant but its meaning has not advanced beyond the benchmark’s initial promise and unresolved validity limits.
2026-08-03T21:23:18Z
No independent MirrorCode replication or validated completed-project result has appeared; the added activity remains anecdotal amplification and does not narrow the benchmark’s scope or validity uncertainty.
2026-08-03T20:28:04Z
The new long-running-agent report makes repository-scale autonomy more plausible in practice, but unvalidated LOC volume is not an independent MirrorCode run and does not establish completed-project scale. Growing discussion also sharpens the benchmark’s central validity question: reproducing behavior may not generalize to developing novel software.
2026-08-03T20:21:47Z
evidence attached: hn.story.49160339 — An independent real-world report of agents producing 870k lines over 12 days bears directly on repository-scale autonomous software development, though validation remains limited.
2026-08-03T17:24:36Z
grounded: converges/high — MirrorCode independently operationalizes several of Scott’s load-bearing positions: repository-scale capability should be measured through long-horizon executio
2026-08-03T17:22:16Z
case created — This is a bounded new coding-agent evaluation with a first-party artifact and enough discussion to merit follow-up.