2026-10-11 17:13 UTC

ARC Prize's published results claim OpenAI's GPT-6 Astra scored 99.95% on ARC-AGI-3 using a provider adapter harness, which would mark frontier-level generalization on abstract reasoning tasks if the methodology holds up.

state: expiredheat: lowuncertainty: highknownscott: lowgpt-6-astra reasoning-benchmark coding-agentsOpenAI

What is this?

The case claims OpenAI's 'GPT-6 Astra' scored 99.95% on ARC-AGI-3, but no web result actually mentions this model or score โ€” the supplied snippets instead show real ARC-AGI-3 results where frontier models (Gemini 3.1 Pro, Claude Opus 4.6, GPT-5.4) scored near zero (0.37%, 0.25%, ~0%), and ARC-AGI-2 leaderboards top out around 83-85%. This strongly suggests the case's headline claim is unverified or fabricated relative to the actual public benchmark record; there is no corroborating source for 'GPT-6 Astra' or a 99.95% figure anywhere in the evidence.

Why it matters to Scott

This is another unverified/likely-fabricated frontier benchmark claim in the exact shape the radar already tracks repeatedly (radar:schema-arc-agi-3-claim, radar:seed-iq-arc-agi-3d-doom, radar:concept.benchmark-integrity, radar:concept.arc-agi) โ€” a vendor-reported ARC-AGI-3 score with no independent corroboration. It illustrates Scott's harness-skepticism/trust-the-harness-not-the-model stance but adds no new claim or actor the radar hasn't already logged; it's one more repetition, not a development.
radar:concept.benchmark-integrityradar:concept.arc-agiradar:schema-arc-agi-3-claimradar:seed-iq-arc-agi-3d-doom
queries asked of Scott's wikis
  • benchmark gaming and metric validity in agent evaluation
  • provider adapter harness patterns in coding agents
  • abstract reasoning vs memorization in LLM benchmarks
  • trusting vendor-reported benchmark claims without reproduction
  • ARC-AGI or generalization claims in Scott's own notes
  • harness design and evaluation methodology for agentic systems

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

no chain yet โ€” the hourly chain pass fills this in

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hn โญOpenAI's GPT-6 Astra on ARC-AGI-3 Scored 99.95% with Provider Adapter HarnessHetPatel10611

Interpretation history

Decision trace