Independent reruns will determine whether SWE-rebench reproducibly reveals stable coding-agent capability differences across Go, Java, Python, Rust, and TypeScript software-engineering tasks.
state: expiredheat: lowuncertainty: highnovelscott: nonecoding-agents coding-agent-benchmarks agent-evaluationSWE-rebench
What is this?
SWE-rebench is a standardized, continuously refreshed benchmark for evaluating coding agents on real GitHub software-engineering tasks, with controlled execution environments intended to isolate model capability from harness differences. Its V2 task collection spans 20 languages, including Go, Java, Python, Rust, and TypeScript, while the announced leaderboard update compares 13 models and four agents. The supplied snippets associate the research with Ibragim Badertdinov and Nebius, but do not clearly establish full project ownership. They describe stable infrastructure and repeated standardized runs, but do not provide clear evidence of genuinely independent reruns confirming that capability differences are reproducible.
Why it matters to Scott
No intersection found: Scott’s wikis contain no supplied position or project tied to multilingual coding-agent benchmark reproducibility, and the radar has no prior page tracking SWE-rebench or this leaderboard update.
queries asked of Scott's wikis
- coding-agent benchmark reproducibility and variance
- model capability versus agent harness effects
- multilingual coding-agent performance differences
- fresh-task benchmarks and contamination resistance
- real-world software-engineering agent evaluation
- coding-agent evaluation across repeated runs
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-08-07T18:31:48Z
No independent SWE-rebench rerun or direct reproducibility evidence has appeared within the case horizon; the Kotlin benchmark remains adjacent rather than validating. With no expected confirming event and no relevance to Scott’s current work, this episode has faded without resolving the hypothesis.
2026-08-01T17:26:25Z
The Kotlin benchmark shows independent interest in language-specific coding-agent evaluation, but it neither reruns SWE-rebench nor validates stable differences across its five languages. The reproducibility claim therefore remains untested rather than corroborated.
2026-08-01T17:21:38Z
evidence attached: hn.story.49136137 — A real-world Kotlin benchmark provides independent evidence about whether coding-agent performance varies materially by programming language and task distribution.
2026-07-31T19:28:14Z
The added activity remains discussion around the original announcement, not an independent rerun or implementation. The reproducibility hypothesis is still untested, so the case stays a low-priority seed despite the hot coding-agent neighborhood.
2026-07-31T18:21:53Z
No engagement movement since last look and no independent reruns or third-party corroboration have surfaced — still a single-source announcement (author's own X post) with no evidence of reproducibility being tested. Relevance remains none per grounding; not yet worth tracking closely.
2026-07-31T17:25:09Z
grounded: novel/none — No intersection found: Scott’s wikis contain no supplied position or project tied to multilingual coding-agent benchmark reproducibility, and the radar has no p
2026-07-31T17:24:24Z
origin walked (codex/luna, conf 0.92): anchor hn.story.49124336 -> echo.x.8656f3b3e2 by Ibragim Badertdinov (@ibragim_bad)
2026-07-31T17:22:45Z
case created — The published multi-language comparison is a bounded benchmark-validation episode directly relevant to tracking coding-agent capability.
Decision trace
- 08-08 04:31expireNo independent SWE-rebench rerun or direct reproducibility evidence has appeared within the case horizon; the Kotlin benchmark remains adjacent rather than validating. With no expected confirming even
- 08-08 04:31alert_silentThe only new delta is elapsed staleness, not substantive evidence; there is nothing consequential to surface or hold for.
- 08-08 04:31alert_routeThe only new delta is elapsed staleness, not substantive evidence; there is nothing consequential to surface or hold for.
- 08-02 03:26repriceThe Kotlin benchmark shows independent interest in language-specific coding-agent evaluation, but it neither reruns SWE-rebench nor validates stable differences across its five languages. The reproduc
- 08-02 03:21attachA real-world Kotlin benchmark provides independent evidence about whether coding-agent performance varies materially by programming language and task distribution.
- 08-02 03:21propose_attachA real-world Kotlin benchmark provides independent evidence about whether coding-agent performance varies materially by programming language and task distribution.
- 08-01 05:28repriceThe added activity remains discussion around the original announcement, not an independent rerun or implementation. The reproducibility hypothesis is still untested, so the case stays a low-priority s
- 08-01 05:20mark_dirtyengagement_update
- 08-01 04:21repriceNo engagement movement since last look and no independent reruns or third-party corroboration have surfaced — still a single-source announcement (author's own X post) with no evidence of reproduc
- 08-01 04:20mark_dirtyengagement_update
- 08-01 03:25groundNo intersection found: Scott’s wikis contain no supplied position or project tied to multilingual coding-agent benchmark reproducibility, and the radar has no prior page tracking SWE-rebench or this l
- 08-01 03:24promote_anchororigin walk conf 0.92
- 08-01 03:22createThe published multi-language comparison is a bounded benchmark-validation episode directly relevant to tracking coding-agent capability.