Independent evaluations will determine whether user code edits during agent execution materially change coding-agent performance or rankings, validating interactive intervention as a necessary benchmark dimension.
state: expiredheat: lowuncertainty: highconvergesscott: mediumcoding-agent-benchmarks interactive-coding-agents swe-touchSWE-Touch
What is this?
SWE-Touch is presented as a coding-agent evaluation framework that introduces validated “Counter-Edits” while an agent is working, testing whether agents can cope with users changing the code during execution. Its stated purpose is to determine whether this interactive condition materially affects agent performance or leaderboard rankings, extending benchmarks beyond static tasks. The supplied results establish that agent performance varies substantially by benchmark and task, but they do not identify SWE-Touch’s authors, report its findings, or provide independent validation of its methodology or claimed ranking effects.
Why it matters to Scott
SWE-Touch independently operationalizes Scott’s claim that coding-agent benchmarks measure the wrong unit when they omit human interaction and changing external state. It creates a dated-receipts and evaluation-design opportunity around Human–AI Collaboration and Path Testing, but no findings or independent validation are supplied yet, so it does not currently establish that interventions change performance or rankings.
ip:concept.benchmarking-the-wrong-unitip:concept.human-ai-collaborationip:concept.path-testingip:framework.reflexive-agent-designradar:concept.agent-benchmarksradar:concept.agent-evaluationradar:concept.coding-agents
queries asked of Scott's wikis
- interactive human intervention in coding-agent workflows
- coding-agent benchmarks versus real-world pair programming
- agent recovery from concurrent or external code edits
- harness evaluation of dynamic environment changes
- shared state and conflict handling in coding agents
- benchmark dimensions for human-agent collaboration
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-09T19:40:35Z
No independent reproduction, implementation, or ranking analysis followed the paper within the observation window. The reported performance drop remains an isolated first-party result, so the case has not developed beyond its initial benchmark proposal.
2026-08-07T19:31:36Z
The paper’s reported 7.7-point resolve-rate drop is first-party evidence that concurrent user edits matter, but no independent evaluation or implementation has appeared to validate the effect or ranking implications.
2026-08-04T04:29:27Z
grounded: converges/medium — SWE-Touch independently operationalizes Scott’s claim that coding-agent benchmarks measure the wrong unit when they omit human interaction and changing external
2026-08-04T04:27:36Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49164247 -> echo.paper.6c376ff4b5 by Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, and Kang Liu
2026-08-04T04:26:28Z
case created — SWE-Touch introduces a bounded and novel evaluation claim, but currently has only one low-engagement observation.
Decision trace
- 08-10 05:40expireNo independent reproduction, implementation, or ranking analysis followed the paper within the observation window. The reported performance drop remains an isolated first-party result, so the case has
- 08-10 05:40alert_silentThe only trigger is staleness, with no new consequential evidence; Scott need not be interrupted unless an independent evaluation or adoption appears.
- 08-10 05:40alert_routeThe only trigger is staleness, with no new consequential evidence; Scott need not be interrupted unless an independent evaluation or adoption appears.
- 08-08 05:31repriceThe paper’s reported 7.7-point resolve-rate drop is first-party evidence that concurrent user edits matter, but no independent evaluation or implementation has appeared to validate the effect or ranki
- 08-08 05:31alert_silentNo new consequential delta arrived; the case remains an uncorroborated benchmark proposal and can wait for an independent reproduction or adoption.
- 08-08 05:31alert_routeNo new consequential delta arrived; the case remains an uncorroborated benchmark proposal and can wait for an independent reproduction or adoption.
- 08-04 14:29groundSWE-Touch independently operationalizes Scott’s claim that coding-agent benchmarks measure the wrong unit when they omit human interaction and changing external state. It creates a dated-receipts and
- 08-04 14:27promote_anchororigin walk conf 0.99
- 08-04 14:26createSWE-Touch introduces a bounded and novel evaluation claim, but currently has only one low-engagement observation.