GoBench presenter Roland31415 claims its 9×9 Go evaluation correlates with ARC-AGI 2 at r=0.83 while retaining substantial headroom and exposing gains from coding-tool preparation, potentially providing an unsaturated benchmark for reasoning and tool-assisted agent capability.
state: watchingheat: lowuncertainty: highconvergesscott: lowllm-evaluation reasoning-benchmarks coding-agentsRoland31415
What is this?
The case describes GoBench as an evaluation of language models playing 9×9 Go, presented by Roland31415; an evidence title attributes the original announcement to Roland Gao, but the supplied web snippets do not verify that attribution or link the two names. The presenter reportedly claims a correlation of r=0.83 with ARC-AGI-2, substantial remaining headroom, and gains from coding-tool preparation. None of the returned web snippets directly covers GoBench, so its methodology, correlation, and reported gains remain uncorroborated here; the pivotools snippet provides related context about evaluating ARC-AGI-2 with a stateful coding environment, not confirmation of GoBench’s results.
Why it matters to Scott
GoBench’s reported gains from coding-tool preparation tentatively converge with Scott’s Model-Plus-Harness Benchmark Unit: measured capability depends on the execution setup, not weights alone. However, the supplied material does not corroborate the gains, correlation or remaining headroom, or establish applicability to his trace-backed agent comparisons, so this is currently another claimed example rather than a reason to change his evaluations; no supplied radar hit tracks GoBench itself.
ip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:ai-benchmark-saturation-distortionradar:ship-harness-benchradar:concept.llm-evaluation
queries asked of Scott's wikis
- agent harness evaluation versus base model capability
- coding tools executable reasoning preparation gains
- benchmark validity correlation generalization real-world performance
- unsaturated benchmarks model selection regression testing
- stateful REPL agent planning evaluation
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 650h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p69 vs 1032 stories at the 336h mark (now 650h old) — ahead of claude-cowork-windows-update-command-failure (1.0x), behind gemini-25-october-migration-gap (1.0x)
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-09-16T20:44:47Z
The author's new comparison turns the previously truncated coding-tool claim into a concrete, testable harness effect, making GoBench more worth tracking. It remains a single-origin result, not independent evidence that Go strength measures general reasoning or transfers to useful coding-agent performance.
2026-09-16T20:22:38Z
evidence attached: reddit.post.1wi75t2 — The benchmark author supplies concrete comparative GoBench scores and shows a large harness effect, materially informing the open benchmark-validation case.
2026-09-16T19:30:00Z
grounded: converges/low — GoBench’s reported gains from coding-tool preparation tentatively converge with Scott’s Model-Plus-Harness Benchmark Unit: measured capability depends on the ex
2026-09-16T19:23:53Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1wi68jg -> echo.blog.d5cbcf6c66 by Roland Gao
2026-09-16T19:22:49Z
case created — The presenter supplies a distinct benchmark design and quantitative results, but the truncated observation does not establish the scout's claimed code and paper releases or independently validate the findings.
Decision trace
- 10-11 15:40drop_targetsquiet through full ladder or over cap 8
- 09-25 00:42review_screenjev screen: no material development (noul=0.09)
- 09-17 17:30review_screenThe changes add questions, praise, and commentary about already known benchmark details without new results, methods, or independently verified evidence.
- 09-17 17:21sensor_dirtycomment_update
- 09-17 06:44repriceThe author's new comparison turns the previously truncated coding-tool claim into a concrete, testable harness effect, making GoBench more worth tracking. It remains a single-origin result, not i
- 09-17 06:22attachThe benchmark author supplies concrete comparative GoBench scores and shows a large harness effect, materially informing the open benchmark-validation case.
- 09-17 06:21propose_attachThe benchmark author supplies concrete comparative GoBench scores and shows a large harness effect, materially informing the open benchmark-validation case.
- 09-17 05:30groundGoBench’s reported gains from coding-tool preparation tentatively converge with Scott’s Model-Plus-Harness Benchmark Unit: measured capability depends on the execution setup, not weights alone. Howeve
- 09-17 05:23promote_anchororigin walk conf 0.99
- 09-17 05:22createThe presenter supplies a distinct benchmark design and quantitative results, but the truncated observation does not establish the scout's claimed code and paper releases or independently validate