The Humanity's Sixth Sense benchmark claims a large human-model gap on intuitive visual, spatial, causal, and social reasoning (humans 93.1% vs GPT-6-Astra 53.6%), proposing a new capability reference for multimodal reasoning.
state: seedheat: mediumuncertainty: mediumconvergesscott: highvisual-reasoning agent-evaluation multimodal-benchmarksCharuru
What is this?
Humanity's Sixth Sense (HSS) is a new benchmark from Scale AI (announced ~1 day before search) targeting intuitive visual reasoning — the implicit spatial, causal, temporal, and social inferences people make at a glance from images and video. It comprises 522 open-ended tasks organized under a structured taxonomy, with human baselines at 93.1% vs. the strongest tested model (GPT-6-Astra) at 53.6% and median model at 30.9%. The benchmark is positioned for ICLR 2027 and released on Hugging Face as an open dataset. The snippets confirm the human-model gap numbers and Scale AI authorship; they do not detail the evaluation protocol, task construction, or whether the GPT-6-Astra result is from a released or internal model.
Why it matters to Scott
Converges with Scott's human-baseline measurement framework (ip:concept.human-baseline-measurement) and evaluation-driven development position (ip:concept.evaluation-driven-development): a new Scale AI benchmark independently adopts a measured human baseline (93.1%) as the reference for multimodal reasoning gaps, extending the pattern tracked in radar episodes like ActiveVision and ZeroBench. The benchmark's focus on spatial, causal, and social visual reasoning also touches his Agents Retina and Skeleton of a Visual work on multimodal sensorium engineering.
ip:concept.human-baselineip:concept.human-baseline-measurementip:concept.evaluation-driven-developmentip:concept.capability-auditip:framework.three-tier-error-budgetsip:source.the-agents-retina-ebookip:source.skeleton-of-a-visual-ebookip:concept.cognitive-irip:framework.cognition-dimension-ladderip:concept.verification-loopsradar:activevision-repeated-perception-gapradar:astra-zerobench-human-baselineradar:frontier-agents-interactive-maze-failuresradar:hle-diamond-benchmarkradar:epoch-ai-innovation-benchmarkradar:concept.ai-benchmarksradar:concept.model-evaluationradar:concept.benchmark-saturationradar:concept.benchmark-integrityradar:concept.evaluationradar:concept.multimodal-modelsradar:concept.vision-language-modelsradar:ai-benchmark-saturation-distortionradar:cleaned-benchmarks-frontier-rankings
queries asked of Scott's wikis
- benchmark saturation and the need for new evaluation paradigms beyond human-authored static sets
- multimodal reasoning evaluation: spatial, causal, and social reasoning as distinct capabilities
- open dataset release practices and community adoption signals for capability references
- Scale AI's evaluation infrastructure and benchmarking strategy
- human-baseline methodology: how human performance is measured and whether it sets a meaningful ceiling
- ICLR 2027 submission landscape: which capability gaps the community is targeting next
Measured heat
now 0 pts/hpeak 120 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 147h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p84 vs 1247 stories at the 96h mark (now 147h old) — ahead of gpt6-prompt-cache-controls (1.0x), behind openai-researcher-separation-safety-sharing (1.0x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-10-08T20:54:27Z
origin walked (opencode/cheap-glm, conf 0.95): anchor reddit.post.1x0twes -> echo.paper.410af4e8e1 by Xingang Guo et al. (Scale AI, in partnership with Elorian)
2026-10-08T18:06:37Z
grounded: converges/high — Converges with Scott's human-baseline measurement framework (ip:concept.human-baseline-measurement) and evaluation-driven development position (ip:concept.evalu
2026-10-08T17:53:45Z
case created — New benchmark with concrete human-vs-frontier-model gap numbers posted to Reddit; adoption as a reference depends on independent uptake.
Decision trace
- 10-10 20:31sensor_dirtyvelocity_spike
- 10-10 12:33sensor_dirtyvelocity_spike
- 10-10 06:37sensor_dirtycomment_update
- 10-10 00:36sensor_dirtyvelocity_spike
- 10-09 15:39sensor_dirtyvelocity_spike
- 10-09 09:34sensor_dirtycomment_update
- 10-09 09:34sensor_dirtyvelocity_spike
- 10-09 07:59attention_routeThe editor compared this story and chose to keep watching.
- 10-09 07:54attention_candidatecreate
- 10-09 07:54promote_anchororigin walk conf 0.95
- 10-09 05:06groundConverges with Scott's human-baseline measurement framework (ip:concept.human-baseline-measurement) and evaluation-driven development position (ip:concept.evaluation-driven-development): a new Sc
- 10-09 04:53createNew benchmark with concrete human-vs-frontier-model gap numbers posted to Reddit; adoption as a reference depends on independent uptake.