Independent analysis and replication will determine whether the Post-Merge Fate of Agentic Code benchmark validly measures downstream maintenance outcomes and reveals systematic differences between agent-generated and conventional code.
state: expiredheat: lowuncertainty: highnovelscott: nonecoding-agents software-maintenance agent-evaluation
What is this?
“Measuring the Post-Merge Fate of Agentic Code” is a longitudinal empirical study comparing agentic and human contributions across 182 software repositories by tracking later modifications, defects, and vulnerabilities. The paper reports similar overall maintenance rates but significantly more corrective maintenance, security weaknesses, and dependency vulnerabilities for agent-generated contributions. The supplied snippets do not identify the authors, establish that this is a validated benchmark, or provide an independent replication; the linked replication package concerns a different coding-agent impact study.
Why it matters to Scott
No intersection found: there are no Scott wiki hits connecting this maintenance benchmark or its claims to his existing work or positions, and no radar hits showing the study is already tracked.
queries asked of Scott's wikis
- coding-agent evaluation beyond task completion
- post-merge maintenance as an agent quality metric
- longitudinal evaluation of agent-generated code
- provenance and attribution of AI-written code
- coding-agent security and dependency risk
- agent harnesses for corrective maintenance
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-07-27T19:21:45Z
No independent analysis, replication, or substantive discussion emerged after launch, so the benchmark’s validity and reported maintenance differences remain uncorroborated with no active signal to track.
2026-07-25T11:22:51Z
grounded: novel/none — No intersection found: there are no Scott wiki hits connecting this maintenance benchmark or its claims to his existing work or positions, and no radar hits sho
2026-07-25T11:22:12Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49046461 -> echo.paper.b347f14d94 by Chunqiu Steven Xia and Courtney Miller
2026-07-25T11:21:24Z
case created — The project introduces a bounded evaluation claim about agent-generated code after merge, but currently has only its launch evidence.
Decision trace
- 07-28 05:21expireNo independent analysis, replication, or substantive discussion emerged after launch, so the benchmark’s validity and reported maintenance differences remain uncorroborated with no active signal to tr
- 07-25 21:22groundNo intersection found: there are no Scott wiki hits connecting this maintenance benchmark or its claims to his existing work or positions, and no radar hits showing the study is already tracked.
- 07-25 21:22promote_anchororigin walk conf 0.99
- 07-25 21:21createThe project introduces a bounded evaluation claim about agent-generated code after merge, but currently has only its launch evidence.