2026-10-11 16:38 UTC

The Humanity's Sixth Sense benchmark claims a large human-model gap on intuitive visual, spatial, causal, and social reasoning (humans 93.1% vs GPT-6-Astra 53.6%), proposing a new capability reference for multimodal reasoning.

state: seedheat: mediumuncertainty: mediumconvergesscott: highvisual-reasoning agent-evaluation multimodal-benchmarksCharuru

What is this?

Humanity's Sixth Sense (HSS) is a new benchmark from Scale AI (announced ~1 day before search) targeting intuitive visual reasoning — the implicit spatial, causal, temporal, and social inferences people make at a glance from images and video. It comprises 522 open-ended tasks organized under a structured taxonomy, with human baselines at 93.1% vs. the strongest tested model (GPT-6-Astra) at 53.6% and median model at 30.9%. The benchmark is positioned for ICLR 2027 and released on Hugging Face as an open dataset. The snippets confirm the human-model gap numbers and Scale AI authorship; they do not detail the evaluation protocol, task construction, or whether the GPT-6-Astra result is from a released or internal model.

Why it matters to Scott

Converges with Scott's human-baseline measurement framework (ip:concept.human-baseline-measurement) and evaluation-driven development position (ip:concept.evaluation-driven-development): a new Scale AI benchmark independently adopts a measured human baseline (93.1%) as the reference for multimodal reasoning gaps, extending the pattern tracked in radar episodes like ActiveVision and ZeroBench. The benchmark's focus on spatial, causal, and social visual reasoning also touches his Agents Retina and Skeleton of a Visual work on multimodal sensorium engineering.
ip:concept.human-baselineip:concept.human-baseline-measurementip:concept.evaluation-driven-developmentip:concept.capability-auditip:framework.three-tier-error-budgetsip:source.the-agents-retina-ebookip:source.skeleton-of-a-visual-ebookip:concept.cognitive-irip:framework.cognition-dimension-ladderip:concept.verification-loopsradar:activevision-repeated-perception-gapradar:astra-zerobench-human-baselineradar:frontier-agents-interactive-maze-failuresradar:hle-diamond-benchmarkradar:epoch-ai-innovation-benchmarkradar:concept.ai-benchmarksradar:concept.model-evaluationradar:concept.benchmark-saturationradar:concept.benchmark-integrityradar:concept.evaluationradar:concept.multimodal-modelsradar:concept.vision-language-modelsradar:ai-benchmark-saturation-distortionradar:cleaned-benchmarks-frontier-rankings
queries asked of Scott's wikis
  • benchmark saturation and the need for new evaluation paradigms beyond human-authored static sets
  • multimodal reasoning evaluation: spatial, causal, and social reasoning as distinct capabilities
  • open dataset release practices and community adoption signals for capability references
  • Scale AI's evaluation infrastructure and benchmarking strategy
  • human-baseline methodology: how human performance is measured and whether it sets a meaningful ceiling
  • ICLR 2027 submission landscape: which capability gaps the community is targeting next

Measured heat

now 0 pts/hpeak 120 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 147h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-05 13:00⭐ origin echo-reconstructedarXiv:2610.08966, "Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models", submitted 6 Oct 2026 by Xingang Gu
Xingang Guo et al. (Scale AI, in partnership with Elorian) on paper (echo) · attributed from reddit.post.1x0twes
—
10-08 15:27first on r/singularity · published · +74.5hIntroducing Humanity’s Sixth Sense, a new benchmark testing intuitive visual reasoning from spatial and causal reasoning to social understanding. The gap is significant. Humans score 93.1%, while the strongest model, GPT-6-astra, reaches 53.6%. The median model scores just 30.9%.
Charuru
—
10-08 15:27amplified on r/singularity 👑reddit.post.1x0twes
Charuru
peak 426 · 80 comments · 100% of case engagement
10-08 15:38our radar first saw it · +74.7hdiscovery anchor: reddit.post.1x0twes—
pace: p84 vs 1247 stories at the 96h mark (now 147h old) — ahead of gpt6-prompt-cache-controls (1.0x), behind openai-researcher-separation-safety-sharing (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditIntroducing Humanity’s Sixth Sense, a new benchmark testing intuitive visual reasoning from spatial and causal reasoning to social understanding. The gap is significant. Humans score 93.1%, while the strongest model, GPT-6-astra, reaches 53.6%. The median model scores just 30.9%.
singularity
Retrieved article excerpt

Open article · Retrieved 2026-10-08T17:49:29.181587+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Charuru41980
🟧 echo.paper ⭐arXiv:2610.08966, "Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models", submitted 6 Oct 2026 by Xingang GuXingang Guo et al. (Scale AI, in partnership with Elorian)——

Interpretation history

Decision trace