ARC Prize's published results claim OpenAI's GPT-6 Astra scored 99.95% on ARC-AGI-3 using a provider adapter harness, which would mark frontier-level generalization on abstract reasoning tasks if the methodology holds up.
state: expiredheat: lowuncertainty: highknownscott: lowgpt-6-astra reasoning-benchmark coding-agentsOpenAI
What is this?
The case claims OpenAI's 'GPT-6 Astra' scored 99.95% on ARC-AGI-3, but no web result actually mentions this model or score โ the supplied snippets instead show real ARC-AGI-3 results where frontier models (Gemini 3.1 Pro, Claude Opus 4.6, GPT-5.4) scored near zero (0.37%, 0.25%, ~0%), and ARC-AGI-2 leaderboards top out around 83-85%. This strongly suggests the case's headline claim is unverified or fabricated relative to the actual public benchmark record; there is no corroborating source for 'GPT-6 Astra' or a 99.95% figure anywhere in the evidence.
Why it matters to Scott
This is another unverified/likely-fabricated frontier benchmark claim in the exact shape the radar already tracks repeatedly (radar:schema-arc-agi-3-claim, radar:seed-iq-arc-agi-3d-doom, radar:concept.benchmark-integrity, radar:concept.arc-agi) โ a vendor-reported ARC-AGI-3 score with no independent corroboration. It illustrates Scott's harness-skepticism/trust-the-harness-not-the-model stance but adds no new claim or actor the radar hasn't already logged; it's one more repetition, not a development.
radar:concept.benchmark-integrityradar:concept.arc-agiradar:schema-arc-agi-3-claimradar:seed-iq-arc-agi-3d-doom
queries asked of Scott's wikis
- benchmark gaming and metric validity in agent evaluation
- provider adapter harness patterns in coding agents
- abstract reasoning vs memorization in LLM benchmarks
- trusting vendor-reported benchmark claims without reproduction
- ARC-AGI or generalization claims in Scott's own notes
- harness design and evaluation methodology for agentic systems
Measured heat
no measured readings yet โ the hourly heat pass fills this in
How the heat travelled
no chain yet โ the hourly chain pass fills this in
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-09-06T23:50:48Z
No new evidence since grounding confirmed this as an unverified/likely-fabricated claim (no such model/score exists in public ARC-AGI-3 records); engagement remains flat at 1 point/1 comment with zero corroboration. Nothing developed within horizon; closing as faded.
2026-09-06T23:48:11Z
grounded: known/low โ This is another unverified/likely-fabricated frontier benchmark claim in the exact shape the radar already tracks repeatedly (radar:schema-arc-agi-3-claim, rada
2026-09-06T23:39:52Z
case created โ Notable benchmark claim about GPT-6 Astra distinct from the existing safety-controls case, warranting its own tracking given hot ai-infrastructure/coding-agents context.
Decision trace
- 09-07 09:50expireNo new evidence since grounding confirmed this as an unverified/likely-fabricated claim (no such model/score exists in public ARC-AGI-3 records); engagement remains flat at 1 point/1 comment with zero
- 09-07 09:50alert_silentStale, uncorroborated single-post claim already assessed as likely fabricated; no first-party or credible secondary confirmation has emerged, and no new delta exists to justify alerting.
- 09-07 09:50alert_routeStale, uncorroborated single-post claim already assessed as likely fabricated; no first-party or credible secondary confirmation has emerged, and no new delta exists to justify alerting.
- 09-07 09:48alert_silentSingle low-engagement HN post with no body, no methodology, no independent corroboration of a 99.95% ARC-AGI-3 score โ this is the same repeated pattern of unverified vendor-shaped benchmark claims al
- 09-07 09:48surface_candidateSingle low-engagement HN post with no body, no methodology, no independent corroboration of a 99.95% ARC-AGI-3 score โ this is the same repeated pattern of unverified vendor-shaped benchmark claims al
- 09-07 09:48alert_routeSingle low-engagement HN post with no body, no methodology, no independent corroboration of a 99.95% ARC-AGI-3 score โ this is the same repeated pattern of unverified vendor-shaped benchmark claims al
- 09-07 09:48groundThis is another unverified/likely-fabricated frontier benchmark claim in the exact shape the radar already tracks repeatedly (radar:schema-arc-agi-3-claim, radar:seed-iq-arc-agi-3d-doom, radar:concept
- 09-07 09:39createNotable benchmark claim about GPT-6 Astra distinct from the existing safety-controls case, warranting its own tracking given hot ai-infrastructure/coding-agents context.