2026-10-11 17:10 UTC

Independent evaluations will determine whether user code edits during agent execution materially change coding-agent performance or rankings, validating interactive intervention as a necessary benchmark dimension.

state: expiredheat: lowuncertainty: highconvergesscott: mediumcoding-agent-benchmarks interactive-coding-agents swe-touchSWE-Touch

What is this?

SWE-Touch is presented as a coding-agent evaluation framework that introduces validated “Counter-Edits” while an agent is working, testing whether agents can cope with users changing the code during execution. Its stated purpose is to determine whether this interactive condition materially affects agent performance or leaderboard rankings, extending benchmarks beyond static tasks. The supplied results establish that agent performance varies substantially by benchmark and task, but they do not identify SWE-Touch’s authors, report its findings, or provide independent validation of its methodology or claimed ranking effects.

Why it matters to Scott

SWE-Touch independently operationalizes Scott’s claim that coding-agent benchmarks measure the wrong unit when they omit human interaction and changing external state. It creates a dated-receipts and evaluation-design opportunity around Human–AI Collaboration and Path Testing, but no findings or independent validation are supplied yet, so it does not currently establish that interventions change performance or rankings.
ip:concept.benchmarking-the-wrong-unitip:concept.human-ai-collaborationip:concept.path-testingip:framework.reflexive-agent-designradar:concept.agent-benchmarksradar:concept.agent-evaluationradar:concept.coding-agents
queries asked of Scott's wikis
  • interactive human intervention in coding-agent workflows
  • coding-agent benchmarks versus real-world pair programming
  • agent recovery from concurrent or external code edits
  • harness evaluation of dynamic environment changes
  • shared state and conflict handling in coding agents
  • benchmark dimensions for human-agent collaboration

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnSWE-Touch: Benchmarking Coding Agents When Users Touch the Codein-silico10
🟧 echo.paper ⭐The original paper introduces SWE-Touch, a framework using validated “Counter-Edits” to test whether coding agents can handle users changingYuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, and Kang Liu——

Interpretation history

Decision trace