The SWE-Race builders claim their benchmark of 188 real concurrency bugs harvested from merged PRs across ~100 Python projects โ each graded by the project's own tests in isolated, history-stripped containers โ becomes an adopted reference for coding-agent concurrency repair, with their reported GLM-5.3-Flash-matches-GPT-5.6-Luna result holding under outside use.
state: seedheat: lowuncertainty: mediumconvergesscott: mediumagent-evaluation coding-agents concurrency
What is this?
SWE-Race is, per its builders' own announcement (the sole evidence on file), a coding-agent benchmark of 188 real concurrency bugs harvested from merged pull requests across roughly 100 Python projects, where an agent's patch is graded by running the project's own tests in isolated, history-stripped containers โ a design intended to block shortcut answers and contamination โ and its headline finding is that GLM-5.3-Flash performs on par with GPT-5.6-Luna. The supplied web search corroborates none of this independently: results returned only Society of Women Engineers and other 'SWE' homonyms, with the one relevant item a Wikipedia disambiguation entry that mentions SWE-Bench (placing the name in the SWE-Bench lineage of LLM coding benchmarks) but says nothing about SWE-Race itself. Adoption of the benchmark and external replication of the model-comparison result are therefore unverified; the case currently rests entirely on the builder announcement and the case file's own description.
Why it matters to Scott
Converges: the builders' anti-cheat construction โ grading by the project's own tests inside history-stripped containers so the agent can't reach the answer it's asked to reproduce โ is a public, dated instance of Scott's Future-Leakage Rule and Hidden Gates information-asymmetry positions, and the GLM-5.3-Flash โ GPT-5.6-Luna headline feeds his capability-symmetry argument, though as a weights-only comparison it is exactly the kind of claim his model-plus-harness benchmark unit says to discount pending harness disclosure. Medium rather than high: it rests on a single builder announcement with adoption and replication unverified, and the radar already tracks sibling PR-derived hidden-test evals โ the news is the liftable history-strip technique and one more parity datapoint, not a changed position.
ip:concept.future-leakage-ruleip:framework.hidden-gates-frameworkip:concept.specification-gamingip:concept.model-plus-harness-benchmark-unitip:concept.capability-symmetryradar:selfbench-pr-derived-evalsradar:concept.coding-agent-evaluationradar:concept.benchmark-integrityradar:concept.coding-modelsradar:frontierharness-17x-cost-variationradar:goldset-python-repair-corpus
queries asked of Scott's wikis
- coding agent benchmark evaluation methodology
- concurrency race condition repair debugging
- benchmark contamination anti-cheat grading design
- SWE-Bench critique coding eval limits
- open-weight model parity frontier coding results
- containerized test harness reproducible agent evals
Measured heat
now 0 pts/hpeak 4 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 194h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p47 vs 1188 stories at the 168h mark (now 194h old) โ ahead of all-your-agents-session-monitor (1.1x), behind armature-coding-agent-vendor-selection (0.9x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-06T08:05:33Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wyw0my -> echo.github.f94d295414 by Evaligo Labs (repo authored by "Danny Lev", the same person as Reddit poster u/heyitsdannyle)
2026-10-06T07:42:50Z
grounded: converges/medium โ Converges: the builders' anti-cheat construction โ grading by the project's own tests inside history-stripped containers so the agent can't reach the answer it'
2026-10-06T07:34:57Z
case created โ A builder-announced evaluation track with a concrete construct (concurrency repair, anti-cheat container design) and a substantive model-comparison finding, distinct from the open proactive-bug-discovery and enterprise-SWE benchmark cases.
Decision trace
- 10-11 12:07review_screenjev screen: no material development (noul=0.04)
- 10-07 00:21sensor_dirtycomment_update
- 10-06 19:05promote_anchororigin walk conf 0.85
- 10-06 18:42groundConverges: the builders' anti-cheat construction โ grading by the project's own tests inside history-stripped containers so the agent can't reach the answer it's asked to reproduce
- 10-06 18:34createA builder-announced evaluation track with a concrete construct (concurrency repair, anti-cheat container design) and a substantive model-comparison finding, distinct from the open proactive-bug-discov