2026-10-11 16:37 UTC

GoBench presenter Roland31415 claims its 9×9 Go evaluation correlates with ARC-AGI 2 at r=0.83 while retaining substantial headroom and exposing gains from coding-tool preparation, potentially providing an unsaturated benchmark for reasoning and tool-assisted agent capability.

state: watchingheat: lowuncertainty: highconvergesscott: lowllm-evaluation reasoning-benchmarks coding-agentsRoland31415

What is this?

The case describes GoBench as an evaluation of language models playing 9×9 Go, presented by Roland31415; an evidence title attributes the original announcement to Roland Gao, but the supplied web snippets do not verify that attribution or link the two names. The presenter reportedly claims a correlation of r=0.83 with ARC-AGI-2, substantial remaining headroom, and gains from coding-tool preparation. None of the returned web snippets directly covers GoBench, so its methodology, correlation, and reported gains remain uncorroborated here; the pivotools snippet provides related context about evaluating ARC-AGI-2 with a stateful coding environment, not confirmation of GoBench’s results.

Why it matters to Scott

GoBench’s reported gains from coding-tool preparation tentatively converge with Scott’s Model-Plus-Harness Benchmark Unit: measured capability depends on the execution setup, not weights alone. However, the supplied material does not corroborate the gains, correlation or remaining headroom, or establish applicability to his trace-backed agent comparisons, so this is currently another claimed example rather than a reason to change his evaluations; no supplied radar hit tracks GoBench itself.
ip:concept.model-plus-harness-benchmark-unitdev:concept.trace-backed-agent-comparisonradar:ai-benchmark-saturation-distortionradar:ship-harness-benchradar:concept.llm-evaluation
queries asked of Scott's wikis
  • agent harness evaluation versus base model capability
  • coding tools executable reasoning preparation gains
  • benchmark validity correlation generalization real-world performance
  • unsaturated benchmarks model selection regression testing
  • stateful REPL agent planning evaluation

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 650h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-14 14:00⭐ origin echo-reconstructedThe original GoBench announcement by Roland Gao, dated “Original: September 15, 2026,” says GoBench measures how well frontier language mode
Roland Gao on blog (echo) · attributed from reddit.post.1wi68jg
—
09-16 18:54first on r/MachineLearning · published · +52.9hGoBench: Evaluating LLMs on the game of Go [R]
Roland31415
—
09-16 19:26first on r/OpenAI · published · +53.5hGPT-6 Astra scores highest on GoBench
Roland31415
—
09-16 18:54amplified on r/MachineLearningreddit.post.1wi68jg
Roland31415
peak 25 · 8 comments · 31% of case engagement
09-16 19:26amplified on r/OpenAI 👑reddit.post.1wi75t2
Roland31415
peak 54 · 21 comments · 70% of case engagement
09-16 19:20our radar first saw it · +53.3hdiscovery anchor: reddit.post.1wi68jg—
pace: p69 vs 1032 stories at the 336h mark (now 650h old) — ahead of claude-cowork-windows-update-command-failure (1.0x), behind gemini-25-october-migration-gap (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditGoBench: Evaluating LLMs on the game of Go [R]
MachineLearning
Retrieved article excerpt

Open article · Retrieved 2026-09-16T19:22:39.591783+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Roland31415238
🟧 echo.blog ⭐The original GoBench announcement by Roland Gao, dated “Original: September 15, 2026,” says GoBench measures how well frontier language modeRoland Gao——
🟠 redditGPT-6 Astra scores highest on GoBench
OpenAI
Roland314155421

Interpretation history

Decision trace