coding-agent-benchmarks
band: coolmomentum: stable
score: 0.0
Episodes (2)
Trajectory notes
- 2026-08-09T19:40:35Z: swe-touch-interactive-agent-benchmark closed (faded) โ SWE-Touch independently operationalizes Scottโs claim that coding-agent benchmarks measure the wrong unit when they omit human interaction and changing external state. It creates a dated-receipts and evaluation-design oppor
- 2026-08-07T18:31:48Z: swe-rebench-multilingual-agent-validation closed (faded) โ No intersection found: Scottโs wikis contain no supplied position or project tied to multilingual coding-agent benchmark reproducibility, and the radar has no prior page tracking SWE-rebench or this leaderboard update.