2026-10-11 17:14 UTC

ZeroBench leaderboard results show GPT-6 Astra surpassing the human baseline across all three metrics of the difficult visual-reasoning benchmark โ€” the first reported model to do so if the scores hold.

state: resolvedheat: lowuncertainty: mediumknownscott: lowgpt-6-astra visual-reasoning benchmark-evaluationOpenAI

What is this?

OpenAI released GPT-6 Astra on September 3, 2026 โ€” its largest training run to date (~100k GPUs, Stargate Texas), trained with earlier OpenAI models supervising the run. The announcement claims state-of-the-art results across many benchmarks, including reported scores on the difficult multi-step visual-reasoning benchmark ZeroBench that surpass the human baseline across its metrics. However, the actual ZeroBench leaderboard material supplied (zerobench.github.io and the BenchLM mirror) shows GPT-5.4 leading the current public snapshot at 41.0% with no Astra entry or human-baseline rows visible, so the specific ZeroBench claim is not corroborated by the leaderboard evidence provided โ€” only by secondary coverage. The surrounding reporting also carries caveats: Astra's headline scores elsewhere (e.g., 99.9% on ARC-AGI-3) depended on undisclosed test settings and a provider adapter harness, and The New Stack notes the model's own Preparedness Framework classifies it at the Critical cybersecurity threshold, gating rollout accordingly.

Why it matters to Scott

The radar already tracks this story in two open cases: radar:openai-gpt-astra-release (the Astra announcement itself, gated on documentation/API substantiation) and radar:gpt-6-astra-arc-agi-3-score (the same release's 99.95% ARC-AGI-3 claim, made through a provider adapter harness with undisclosed settings). The ZeroBench claim is another instance of the identical pattern โ€” an announcement-only benchmark score, not yet on the leaderboard snapshot the grounding cites, that must be defended at announcement evidence class per ip:concept.evidence-class-ladder and read through ip:concept.model-plus-harness-benchmark-unit. Nothing new on either side: no independent replication has arrived, and OpenAI is not converging on any Scott position. The one latent hook worth watching inside the existing case: if Astra's visual reasoning were ever independently confirmed at superhuman-baseline level, it would pressure Scott's text-vision routing approach (ip:concept.text-is-the-models-home-turf, ip:concept.denoised-semantic-dom) โ€” but an unreplicated claim can't carry that contradiction yet.
ip:concept.evidence-class-ladderip:concept.model-plus-harness-benchmark-unitip:concept.text-is-the-models-home-turfradar:openai-gpt-astra-releaseradar:gpt-6-astra-arc-agi-3-scoreradar:concept.model-evaluationradar:concept.multimodal-models
queries asked of Scott's wikis
  • benchmark score inflation and harness/adapter effects on eval results
  • contamination-controlled evaluation and benchmark skepticism
  • visual reasoning and multimodal model evaluation limits
  • human-baseline-relative benchmark scoring methodology
  • frontier release announcements versus independent replication
  • ARC-style benchmark saturation and moving-target eval design

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

09-23 13:00โญ origin directly observedGPT-6 Astra makes a massive leap on ZeroBench (an extremely difficult vision benchmark), surpassing the human baseline across all three metrics
Waiting4AniHaremFDVR on r/singularity
โ€”
09-23 13:00amplified on r/singularity ๐Ÿ‘‘reddit.post.1wo5c7t
Waiting4AniHaremFDVR
peak 329 ยท 54 comments ยท 100% of case engagement
09-23 13:20our radar first saw it ยท +0.3hdiscovery anchor: reddit.post.1wo5c7tโ€”

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญGPT-6 Astra makes a massive leap on ZeroBench (an extremely difficult vision benchmark), surpassing the human baseline across all three metrics
singularity
Waiting4AniHaremFDVR33554

Interpretation history

Decision trace