2026-10-11 17:11 UTC

Independent evaluation will determine whether Schema's process-only harness genuinely achieves 99% on the ARC-AGI-3 public set without model-weight changes or evaluation leakage.

state: expiredheat: lowuncertainty: highconvergesscott: mediumagent-harness arc-agi benchmark-verification test-time-computeAnthropicOpenAI

What is this?

Schema is a claimed agent harness for ARC-AGI-3, an interactive benchmark testing exploration, planning, adaptable world models, memory compression, and belief updating over time. Its site says the harness has frontier models construct executable models of game mechanics, test them against observations, and plan within them; it reports exact history reproduction in 14 of 25 games and lower action use than the human reference in those games. Social-post titles claim roughly 99% on the public set, but the supplied snippets do not establish an independent evaluation, rule out leakage, identify the harness creators, or substantiate involvement by Anthropic or OpenAI.

Why it matters to Scott

The Schema claim exemplifies Scott's specification-gaming and hidden-gates concerns: 99% on a public set without model-weight changes is exactly the pattern his frameworks anticipate, where a harness games the visible evaluation rubric. The claim reinforces his position that independent verification with mechanically different verifiers is essential, and that outcome capability is a product of harness × model, not model alone. The material does not establish whether the claim is valid, but the pattern itself converges with his critique of public-benchmark trustworthiness.
ip:framework.hidden-gates-frameworkip:concept.specification-gamingip:concept.correlated-checkers-pitfallip:concept.mechanically-different-verifiersip:source.give-the-agent-a-workshop-ebookip:framework.evaluation-driven-developmentip:concept.scaffolding-hypothesis
queries asked of Scott's wikis
  • agent harnesses versus model capability gains
  • executable world models for autonomous agents
  • test-time compute and iterative hypothesis testing
  • benchmark leakage public sets and holdout verification
  • agent memory compression and belief revision
  • harness portability across frontier and open models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐New Fable5/Opus4.8 harness called "Schema" claims 99% on ARC-3 [R]
MachineLearning
we_are_mammals013
🟠 redditARC AGI 3 could be gamed if Opus is a loop and not a pure model
singularity
sdnr8193109
🟠 redditARC-AGI-3 is 🗑️ Read before voting
singularity
No_Engineering_3223037

Interpretation history

Decision trace