2026-10-11 16:37 UTC

visual-reasoning

band: warmmomentum: stable score: 0.315
temperature history

Episodes (2)

ZeroBench leaderboard results show GPT-6 Astra surpassing the human baseline across all three metrics of the difficult visual-reasoning benchmark โ€” the first reported model to do so if the scores hold.
resolvedknownscott: low
The Humanity's Sixth Sense benchmark claims a large human-model gap on intuitive visual, spatial, causal, and social reasoning (humans 93.1% vs GPT-6-Astra 53.6%), proposing a new capability reference for multimodal reasoning.
seedconvergesscott: high