ZeroBench leaderboard results show GPT-6 Astra surpassing the human baseline across all three metrics of the difficult visual-reasoning benchmark โ the first reported model to do so if the scores hold.
state: resolvedheat: lowuncertainty: mediumknownscott: lowgpt-6-astra visual-reasoning benchmark-evaluationOpenAI
What is this?
OpenAI released GPT-6 Astra on September 3, 2026 โ its largest training run to date (~100k GPUs, Stargate Texas), trained with earlier OpenAI models supervising the run. The announcement claims state-of-the-art results across many benchmarks, including reported scores on the difficult multi-step visual-reasoning benchmark ZeroBench that surpass the human baseline across its metrics. However, the actual ZeroBench leaderboard material supplied (zerobench.github.io and the BenchLM mirror) shows GPT-5.4 leading the current public snapshot at 41.0% with no Astra entry or human-baseline rows visible, so the specific ZeroBench claim is not corroborated by the leaderboard evidence provided โ only by secondary coverage. The surrounding reporting also carries caveats: Astra's headline scores elsewhere (e.g., 99.9% on ARC-AGI-3) depended on undisclosed test settings and a provider adapter harness, and The New Stack notes the model's own Preparedness Framework classifies it at the Critical cybersecurity threshold, gating rollout accordingly.
Why it matters to Scott
The radar already tracks this story in two open cases: radar:openai-gpt-astra-release (the Astra announcement itself, gated on documentation/API substantiation) and radar:gpt-6-astra-arc-agi-3-score (the same release's 99.95% ARC-AGI-3 claim, made through a provider adapter harness with undisclosed settings). The ZeroBench claim is another instance of the identical pattern โ an announcement-only benchmark score, not yet on the leaderboard snapshot the grounding cites, that must be defended at announcement evidence class per ip:concept.evidence-class-ladder and read through ip:concept.model-plus-harness-benchmark-unit. Nothing new on either side: no independent replication has arrived, and OpenAI is not converging on any Scott position. The one latent hook worth watching inside the existing case: if Astra's visual reasoning were ever independently confirmed at superhuman-baseline level, it would pressure Scott's text-vision routing approach (ip:concept.text-is-the-models-home-turf, ip:concept.denoised-semantic-dom) โ but an unreplicated claim can't carry that contradiction yet.
ip:concept.evidence-class-ladderip:concept.model-plus-harness-benchmark-unitip:concept.text-is-the-models-home-turfradar:openai-gpt-astra-releaseradar:gpt-6-astra-arc-agi-3-scoreradar:concept.model-evaluationradar:concept.multimodal-models
queries asked of Scott's wikis
- benchmark score inflation and harness/adapter effects on eval results
- contamination-controlled evaluation and benchmark skepticism
- visual reasoning and multimodal model evaluation limits
- human-baseline-relative benchmark scoring methodology
- frontier release announcements versus independent replication
- ARC-style benchmark saturation and moving-target eval design
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-09-25T12:43:04Z
The episode's attention window closed without the claim ever leaving announcement class: the single Reddit thread decayed to zero velocity (~335 pts, 54 comments, one platform, no derivative coverage anywhere), and no leaderboard update, independent source, or replication arrived in the ~2 days since grounding showed zerobench.github.io with no Astra entry. Nothing about the claim's meaning changed โ it simply stopped developing, and any leaderboard substantiation routes through radar:openai-gpt-astra-release, which is already gated on documentation/API substantiation. The recent velocity_spike sensors were the decay tail of the original wave, not renewed spread.
2026-09-23T17:24:19Z
grounded: known/low โ The radar already tracks this story in two open cases: radar:openai-gpt-astra-release (the Astra announcement itself, gated on documentation/API substantiation)
2026-09-23T17:19:24Z
case created โ A widely discussed (184-point) distinct capability claim on a separate benchmark from the RLI result, meriting its own case.
Decision trace
- 09-25 22:43resolveThe episode's attention window closed without the claim ever leaving announcement class: the single Reddit thread decayed to zero velocity (~335 pts, 54 comments, one platform, no derivative cove
- 09-25 01:21sensor_dirtyvelocity_spike
- 09-24 16:21sensor_dirtyvelocity_spike
- 09-24 10:21sensor_dirtycomment_update
- 09-24 07:24sensor_dirtyvelocity_spike
- 09-24 03:24groundThe radar already tracks this story in two open cases: radar:openai-gpt-astra-release (the Astra announcement itself, gated on documentation/API substantiation) and radar:gpt-6-astra-arc-agi-3-score (
- 09-24 03:19createA widely discussed (184-point) distinct capability claim on a separate benchmark from the RLI result, meriting its own case.