2026-10-11 16:36 UTC

coding-agent-evaluation

band: coolmomentum: stable score: 0.125
temperature history

Episodes (3)

Independent use will determine whether 514 provides practical managed environments and behavioral data for simulating coding-agent users and evaluating agent workflows.
expiredconvergesscott: medium
The study’s authors report that Claude-authored pull requests reached an 84% merge rate versus 85% for humans, 74% for Codex, and 43% for Devin, suggesting near-human acceptance in the sampled work and substantial product-level differences.
seedconvergesscott: medium
vyang472 claims the released five-bugs experiment records 26 coding attempts passing visible tests while failing the same unseen text-preservation case, with one stronger-test rerun fixing it, suggesting specification coverage rather than model scale constrained correctness on this task.
seedknownscott: low