2026-10-11 17:14 UTC

The SWE-Race builders claim their benchmark of 188 real concurrency bugs harvested from merged PRs across ~100 Python projects โ€” each graded by the project's own tests in isolated, history-stripped containers โ€” becomes an adopted reference for coding-agent concurrency repair, with their reported GLM-5.3-Flash-matches-GPT-5.6-Luna result holding under outside use.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation coding-agents concurrency

What is this?

SWE-Race is, per its builders' own announcement (the sole evidence on file), a coding-agent benchmark of 188 real concurrency bugs harvested from merged pull requests across roughly 100 Python projects, where an agent's patch is graded by running the project's own tests in isolated, history-stripped containers โ€” a design intended to block shortcut answers and contamination โ€” and its headline finding is that GLM-5.3-Flash performs on par with GPT-5.6-Luna. The supplied web search corroborates none of this independently: results returned only Society of Women Engineers and other 'SWE' homonyms, with the one relevant item a Wikipedia disambiguation entry that mentions SWE-Bench (placing the name in the SWE-Bench lineage of LLM coding benchmarks) but says nothing about SWE-Race itself. Adoption of the benchmark and external replication of the model-comparison result are therefore unverified; the case currently rests entirely on the builder announcement and the case file's own description.

Why it matters to Scott

Converges: the builders' anti-cheat construction โ€” grading by the project's own tests inside history-stripped containers so the agent can't reach the answer it's asked to reproduce โ€” is a public, dated instance of Scott's Future-Leakage Rule and Hidden Gates information-asymmetry positions, and the GLM-5.3-Flash โ‰ˆ GPT-5.6-Luna headline feeds his capability-symmetry argument, though as a weights-only comparison it is exactly the kind of claim his model-plus-harness benchmark unit says to discount pending harness disclosure. Medium rather than high: it rests on a single builder announcement with adoption and replication unverified, and the radar already tracks sibling PR-derived hidden-test evals โ€” the news is the liftable history-strip technique and one more parity datapoint, not a changed position.
ip:concept.future-leakage-ruleip:framework.hidden-gates-frameworkip:concept.specification-gamingip:concept.model-plus-harness-benchmark-unitip:concept.capability-symmetryradar:selfbench-pr-derived-evalsradar:concept.coding-agent-evaluationradar:concept.benchmark-integrityradar:concept.coding-modelsradar:frontierharness-17x-cost-variationradar:goldset-python-repair-corpus
queries asked of Scott's wikis
  • coding agent benchmark evaluation methodology
  • concurrency race condition repair debugging
  • benchmark contamination anti-cheat grading design
  • SWE-Bench critique coding eval limits
  • open-weight model parity frontier coding results
  • containerized test harness reproducible agent evals

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 194h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-03 14:00โญ origin echo-reconstructedThe origin is the benchmark's own v0.1 release (single commit "SWE-Race v0.1 public split (95 tasks)", 2026-10-04). README: "95 real concurr
Evaligo Labs (repo authored by "Danny Lev", the same person as Reddit poster u/heyitsdannyle) on github (echo) ยท attributed from reddit.post.1wyw0my
โ€”
10-06 07:03first on r/MachineLearning ยท published ยท +65.1hSWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
heyitsdannyle
โ€”
10-06 07:03amplified on r/MachineLearning ๐Ÿ‘‘reddit.post.1wyw0my
heyitsdannyle
peak 4 ยท 8 comments ยท 100% of case engagement
10-06 07:20our radar first saw it ยท +65.3hdiscovery anchor: reddit.post.1wyw0myโ€”
pace: p47 vs 1188 stories at the 168h mark (now 194h old) โ€” ahead of all-your-agents-session-monitor (1.1x), behind armature-coding-agent-vendor-selection (0.9x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditSWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
MachineLearning
heyitsdannyle48
๐ŸŸง echo.github โญThe origin is the benchmark's own v0.1 release (single commit "SWE-Race v0.1 public split (95 tasks)", 2026-10-04). README: "95 real concurrEvaligo Labs (repo authored by "Danny Lev", the same person as Reddit poster u/heyitsdannyle)โ€”โ€”

Interpretation history

Decision trace