2026-10-11 18:03 UTC

Independent reruns will determine whether SWE-rebench reproducibly reveals stable coding-agent capability differences across Go, Java, Python, Rust, and TypeScript software-engineering tasks.

state: expiredheat: lowuncertainty: highnovelscott: nonecoding-agents coding-agent-benchmarks agent-evaluationSWE-rebench

What is this?

SWE-rebench is a standardized, continuously refreshed benchmark for evaluating coding agents on real GitHub software-engineering tasks, with controlled execution environments intended to isolate model capability from harness differences. Its V2 task collection spans 20 languages, including Go, Java, Python, Rust, and TypeScript, while the announced leaderboard update compares 13 models and four agents. The supplied snippets associate the research with Ibragim Badertdinov and Nebius, but do not clearly establish full project ownership. They describe stable infrastructure and repeated standardized runs, but do not provide clear evidence of genuinely independent reruns confirming that capability differences are reproducible.

Why it matters to Scott

No intersection found: Scott’s wikis contain no supplied position or project tied to multilingual coding-agent benchmark reproducibility, and the radar has no prior page tracking SWE-rebench or this leaderboard update.
queries asked of Scott's wikis
  • coding-agent benchmark reproducibility and variance
  • model capability versus agent harness effects
  • multilingual coding-agent performance differences
  • fresh-task benchmarks and contamination resistance
  • real-world software-engineering agent evaluation
  • coding-agent evaluation across repeated runs

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TSibragim_bad2811
🟧 echo.x ⭐Primary announcement for the multilingual SWE-rebench leaderboard update. The linked announcement says: “We’ve just released a major update Ibragim Badertdinov (@ibragim_bad)——
🟧 hnThe Kotlin Benchmark for AI Coding Agentsyruzin20

Interpretation history

Decision trace