A paper titled “Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance,” listed for MSR ’26, compares pull-request acceptance across coding agents including Claude, OpenAI Codex, and Devin. The supplied summary reports merge rates of 84% for Claude, 85% for humans, 74% for Codex, and 43% for Devin, implying near-human acceptance for Claude in the sampled work. However, the paper snippet itself shows a 61.6% acceptance rate for Devin, so the exact figures or populations may differ; the snippets do not identify the study’s authors or establish that merge rate directly measures coding quality.
The reported near-human Claude merge rate independently supports Scott’s existing claim that coding agents can perform production-grade work, while the large cross-product spread reinforces his model-plus-harness evaluation unit and trace-backed comparison practice. It is a potential dated-receipts and tooling-selection signal, but the conflicting Devin figures and uncertainty over populations and whether merge acceptance measures quality limit its weight pending inspection or replication.
ip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonip:concept.evidence-class-ladderradar:post-merge-agentic-code-benchmarkradar:concept.coding-agent-benchmarksradar:concept.agent-evaluation
queries asked of Scott's wikis
- pull-request merge rate as coding-agent evaluation
- task-stratified coding-agent benchmarks
- agent harness effects versus model capability
- human review bottlenecks for coding agents
- production acceptance metrics for autonomous code
- coding-agent product performance differences
2026-09-23T17:57:43Z
The latest informal assignment comparison measures manually graded coding performance, not authored-PR acceptance, and does not resolve the study’s conflicting figures. Despite the spread flag, the attachments span separate coding-agent stories rather than an expanding discussion or replication of this study, so this case remains low-attention.
2026-09-23T08:21:46Z
evidence attached: reddit.post.1wnzf3w — This small, informal coding comparison adds weak contextual evidence about model differences on manually graded work, but does not independently validate PR acceptance rates.
2026-09-14T16:33:20Z
The new benchmark reports bug-finding and cost differences in PR review, not acceptance of agent-authored PRs, so it does not independently corroborate the study’s comparative merge rates. It is a separate evaluation lead rather than a reason to promote this case or broaden its hypothesis.
2026-09-14T16:22:49Z
evidence attached: reddit.post.1wg6tk1 — A real-PR benchmark adds independent evidence on coding-agent review quality, bug detection, and cost tradeoffs.
2026-09-13T03:21:47Z
The new workplace anecdote does not substantiate the attachment’s claimed deployment failure: the supplied excerpt stops before any outcome or review details. It neither corroborates nor contradicts the comparative merge rates, leaving this an unresolved evaluation lead rather than evidence for tooling selection.
2026-09-13T03:21:38Z
evidence attached: reddit.post.1wew453 — Real-world deployment anecdote provides cautionary context on large agent-generated changes, weak review, and subsequent failure.
2026-09-11T18:46:55Z
The Node-to-Go rewrite attachment is a headline-only implementation lead, not a verified production outcome or corroboration of comparative PR acceptance. It does not change the study’s evidentiary standing or justify broadening this case into a general coding-agent capability story.
2026-09-11T18:22:46Z
evidence attached: hn.story.49662139 — A production-scale Node-to-Go rewrite is independent case-study evidence bearing on how far coding agents can execute consequential repository changes.
2026-09-09T17:25:37Z
The Gradle practitioner report adds a concrete single-issue evaluation lead with linked diffs and cost/token receipts, but it tests different models and outcomes and does not corroborate the PR merge-rate comparison. The original study remains unresolved pending comparable populations and reconciled Devin figures; the new report is useful evaluation-method context, not a general tooling ranking.
2026-09-09T17:23:52Z
evidence attached: reddit.post.1wbrrhp — An independent coding-agent evaluation with cost, token, and diff receipts adds useful evidence about model-level task performance.
2026-09-08T11:26:25Z
This is a stale evaluation lead, not a developing performance result: the latest review adds no substantive evidence, and Microsoft adoption commentary does not corroborate the comparative merge rates. Keep the hypothesis unresolved and space out reviews pending clarified study populations, reconciled Devin figures, or independent measured outcomes.
2026-09-06T10:30:08Z
The refreshed Microsoft discussion adds no evidence bearing on comparative PR acceptance; adoption claims and complaints about software quality neither validate nor refute the study. The comparison remains an unresolved evaluation lead, with further review best gated on methodology, reconciled Devin figures, or independent measured outcomes.
2026-09-06T05:23:50Z
The refreshed discussion is repetitive adoption commentary, not a new enterprise-adoption event or independent evidence for the PR rankings. The study remains an unresolved evaluation lead rather than a tooling-selection receipt; review should wait for methodology, reconciled Devin figures, or comparable measured outcomes.
2026-09-06T01:26:06Z
The refreshed Microsoft comments add no measured outcomes or independent support for the PR comparison; enterprise-adoption rhetoric cannot validate the agent rankings. The study remains an unresolved evaluation lead, with substantive review warranted by clarified populations and reconciled Devin figures rather than comment churn.
2026-09-06T00:24:11Z
The refreshed comments remain repetitive adoption commentary, with no new evidence about comparative PR acceptance or downstream quality. Keep the study as an unresolved evaluation lead, and space out reviews unless methodology, reconciled figures, or independent outcomes arrive.
2026-09-05T23:24:01Z
The refreshed Microsoft discussion adds neither independent validation of the PR comparison nor evidence connecting acceptance rates to downstream software quality. Keep this as a study lead, with further review gated on substantive methodology or outcomes rather than adoption commentary.
2026-09-05T20:25:51Z
The refreshed Microsoft discussion remains adoption commentary, not evidence that distinguishes coding-agent acceptance or quality. The study is still an evaluation lead rather than a tooling-selection basis; further comment-driven reviews add little until the conflicting Devin figures and sample comparability are addressed.
2026-09-05T19:30:39Z
The refreshed comments repeat adoption rhetoric and speculation without adding measured outcomes or independent support for the PR comparison. The study remains a potentially useful evaluation lead, not a tooling-selection receipt, pending reconciliation of the Devin figures and inspection of task and population comparability.
2026-09-05T17:30:21Z
The refreshed discussion adds speculation and reactions to Microsoft’s reported adoption, not measured evidence about agent performance. The merge-rate comparison remains a study lead rather than a tooling-selection receipt until its populations, task controls, and conflicting Devin figures are reconciled.
2026-09-05T16:28:26Z
The Microsoft item adds reported enterprise-adoption context, not independent validation of the study’s agent rankings or near-human acceptance comparison. The refreshed discussion supplies no measured outcomes; conflicting Devin figures and unresolved sample comparability still prevent using these rates as a tooling-selection receipt.
2026-09-05T16:22:43Z
evidence attached: reddit.post.1w84ngu — Reported Microsoft adoption and AI-authored Windows work provide useful enterprise corroboration that coding agents are moving into mainstream software production.
2026-09-04T15:53:05Z
No new evidence or discussion corroborates the reported comparison, so the case remains an unvalidated study claim rather than a production-grade capability receipt. The conflicting Devin figures and missing methodology continue to limit interpretation.
2026-09-04T15:38:06Z
grounded: converges/medium — The reported near-human Claude merge rate independently supports Scott’s existing claim that coding agents can perform production-grade work, while the large cr
2026-09-04T15:34:43Z
case created — The paper offers concrete comparative field evidence on coding-agent acceptance rather than another synthetic task benchmark.