2026-10-11 16:37 UTC

The study’s authors report that Claude-authored pull requests reached an 84% merge rate versus 85% for humans, 74% for Codex, and 43% for Devin, suggesting near-human acceptance in the sampled work and substantial product-level differences.

state: seedheat: lowuncertainty: highconvergesscott: mediumcoding-agent-evaluation software-engineering-agentsAnthropicOpenAI

What is this?

A paper titled “Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance,” listed for MSR ’26, compares pull-request acceptance across coding agents including Claude, OpenAI Codex, and Devin. The supplied summary reports merge rates of 84% for Claude, 85% for humans, 74% for Codex, and 43% for Devin, implying near-human acceptance for Claude in the sampled work. However, the paper snippet itself shows a 61.6% acceptance rate for Devin, so the exact figures or populations may differ; the snippets do not identify the study’s authors or establish that merge rate directly measures coding quality.

Why it matters to Scott

The reported near-human Claude merge rate independently supports Scott’s existing claim that coding agents can perform production-grade work, while the large cross-product spread reinforces his model-plus-harness evaluation unit and trace-backed comparison practice. It is a potential dated-receipts and tooling-selection signal, but the conflicting Devin figures and uncertainty over populations and whether merge acceptance measures quality limit its weight pending inspection or replication.
ip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonip:concept.evidence-class-ladderradar:post-merge-agentic-code-benchmarkradar:concept.coding-agent-benchmarksradar:concept.agent-evaluation
queries asked of Scott's wikis
  • pull-request merge rate as coding-agent evaluation
  • task-stratified coding-agent benchmarks
  • agent harness effects versus model capability
  • human review bottlenecks for coding agents
  • production acceptance metrics for autonomous code
  • coding-agent product performance differences

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 889h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-04 14:38⭐ origin directly observedAI coding agent PR merge rates: Claude 84%, Codex 74%, Devin 43%, humans 85%
d-yoda on hacker news
—
09-05 16:03first on r/singularity · published · +25.4hMicrosoft's distinguished engineer says "typing code is absolutely over," and Windows 11 is already being built that way. Nadella previously said 20-30% code at MSFT is AI-coded, and Windows security updates now include AI-assisted fixes to fight AI-enabled threats.
WPHero
—
09-09 17:14first on r/OpenAI · published · +122.6hFable 5.1 v Astra 6 on good ol' Gradle
Party_Till_I_Die
—
09-11 17:31first on hacker news · published · +170.9hLet AI Agents Rewrite a 92M-Message-a-Day Node Service to Go
tnolet
—
09-14 15:37first on r/artificial · published · +241.0hGPT-5.6 Luna vs GPT-6 Astra: is a $1.20 model good enough for code review?
entelligenceai17
—
09-23 07:43first on r/ClaudeAI · published · +449.1hI gave 5 AI models the same 3 coding assignments and graded them like a teacher. Claude Opus 5.5 came first, Claude Fable 5.1 second
Short_Regular_7191
—
09-04 14:38amplified on hacker newshn.story.49565335
d-yoda
peak 1 · 0 comments · 0% of case engagement
09-05 16:03amplified on r/singularity 👑reddit.post.1w84ngu
WPHero
peak 371 · 112 comments · 80% of case engagement
09-09 17:14amplified on r/OpenAIreddit.post.1wbrrhp
Party_Till_I_Die
peak 6 · 2 comments · 1% of case engagement
09-11 17:31amplified on hacker newshn.story.49662139
tnolet
peak 2 · 0 comments · 1% of case engagement
09-13 03:03amplified on r/singularityreddit.post.1wew453
jmclondon97
peak 13 · 80 comments · 15% of case engagement
09-14 15:37amplified on r/artificialreddit.post.1wg6tk1
entelligenceai17
peak 5 · 4 comments · 1% of case engagement
1 more amplifiers in ainews.case_chain
09-04 15:21our radar first saw it · +0.7hdiscovery anchor: hn.story.49565335—
pace: p85 vs 519 stories at the 720h mark (now 889h old) — ahead of qwen-drive-1-release (1.0x), behind deepseek-v4-flash-vision-release (1.0x)

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐AI coding agent PR merge rates: Claude 84%, Codex 74%, Devin 43%, humans 85%d-yoda10
🟠 redditMicrosoft's distinguished engineer says "typing code is absolutely over," and Windows 11 is already being built that way. Nadella previously said 20-30% code at MSFT is AI-coded, and Windows security updates now include AI-assisted fixes to fight AI-enabled threats.
singularity
WPHero371111
🟠 redditFable 5.1 v Astra 6 on good ol' Gradle
OpenAI
Party_Till_I_Die62
🟧 hnLet AI Agents Rewrite a 92M-Message-a-Day Node Service to Gotnolet20
🟠 redditSoftware Engineering is Over! Not so fast…
singularity
jmclondon971180
🟠 redditGPT-5.6 Luna vs GPT-6 Astra: is a $1.20 model good enough for code review?
artificial
entelligenceai1754
🟠 redditI gave 5 AI models the same 3 coding assignments and graded them like a teacher. Claude Opus 5.5 came first, Claude Fable 5.1 second
ClaudeAI
Short_Regular_719143

Interpretation history

Decision trace