Schema is a claimed agent harness for ARC-AGI-3, an interactive benchmark testing exploration, planning, adaptable world models, memory compression, and belief updating over time. Its site says the harness has frontier models construct executable models of game mechanics, test them against observations, and plan within them; it reports exact history reproduction in 14 of 25 games and lower action use than the human reference in those games. Social-post titles claim roughly 99% on the public set, but the supplied snippets do not establish an independent evaluation, rule out leakage, identify the harness creators, or substantiate involvement by Anthropic or OpenAI.
The Schema claim exemplifies Scott's specification-gaming and hidden-gates concerns: 99% on a public set without model-weight changes is exactly the pattern his frameworks anticipate, where a harness games the visible evaluation rubric. The claim reinforces his position that independent verification with mechanically different verifiers is essential, and that outcome capability is a product of harness × model, not model alone. The material does not establish whether the claim is valid, but the pattern itself converges with his critique of public-benchmark trustworthiness.
ip:framework.hidden-gates-frameworkip:concept.specification-gamingip:concept.correlated-checkers-pitfallip:concept.mechanically-different-verifiersip:source.give-the-agent-a-workshop-ebookip:framework.evaluation-driven-developmentip:concept.scaffolding-hypothesis
queries asked of Scott's wikis
- agent harnesses versus model capability gains
- executable world models for autonomous agents
- test-time compute and iterative hypothesis testing
- benchmark leakage public sets and holdout verification
- agent memory compression and belief revision
- harness portability across frontier and open models
2026-07-25T09:28:58Z
No independent reproduction, holdout evaluation, or implementation disclosure has appeared across many engagement-only reobservations; the discussion remains repetitive benchmark-design debate. Closing the window on this single-source claim; will reopen only if genuine verification evidence surfaces.
2026-07-25T08:23:20Z
The nominal evidence attachments contain no observations and add nothing to the exhausted benchmark debate. Schema remains an unverified, leakage-sensitive single-source claim; revisit only for an independent reproduction, holdout evaluation, or substantive harness disclosure.
2026-07-25T07:22:42Z
The only measurable change is a negligible score increase on an existing benchmark-design discussion; no independent reproduction, holdout test, or harness disclosure has appeared. Engagement-only reobservations are exhausted, so retain this as a cold verification watch and revisit only on substantive evidence.
2026-07-25T06:23:29Z
The newly attached reobservations contain no substantive evidence and do not change the single-source, leakage-sensitive status of Schema’s claim. Repeated engagement-only triggers are exhausted; revisit only if an independent reproduction, holdout evaluation, or implementation disclosure appears.
2026-07-25T05:22:44Z
The new attachment contains no substantive observation and does not alter the evidentiary picture: Schema’s score remains a leakage-sensitive, single-source claim without independent reproduction or holdout testing. Repeated engagement-only triggers are amplification, not verification.
2026-07-25T03:23:11Z
The nominally new attachment contains no substantive observation and adds no independent reproduction, holdout result, or implementation disclosure. The case remains a cold verification watch; further engagement-only triggers should not change its meaning.
2026-07-25T02:21:39Z
The latest trigger contains no substantive new evidence; Schema’s 99% result remains an unverified single-source claim without reproduction, holdout testing, or implementation disclosure. Repetitive engagement should not prompt another near-term review absent independent evaluation.
2026-07-25T01:21:04Z
The latest reobservation adds no independent reproduction, holdout evaluation, or implementation disclosure; it remains repetitive benchmark debate around an unverified single-source claim. Keep it as a cold verification watch pending genuinely new evidence.
2026-07-25T00:22:17Z
The reobserved material still provides no independent reproduction, holdout evaluation, or implementation disclosure for Schema’s claimed score. It is repetitive benchmark debate rather than verification, so the case remains a cold single-source watch.
2026-07-24T23:23:40Z
No substantive independent evidence has arrived: the activity still concerns benchmark design and looping rather than reproducing Schema’s score or testing it on a holdout. This remains a cold, single-source verification watch despite the surrounding agent-harness topic staying warm.
2026-07-24T22:28:41Z
The added activity remains debate about benchmark design and looping rather than independent verification of Schema’s score. With no reproduction, holdout result, or implementation disclosure, this is repetitive amplification of an unresolved single-source claim.
2026-07-24T21:24:13Z
The new material broadens the benchmark-validity critique but adds no independent reproduction, holdout result, or implementation evidence for Schema itself. Discussion remains speculative and repetitive, so the case stays a cold verification watch rather than advancing toward corroboration.
2026-07-24T21:21:13Z
evidence attached: reddit.post.1v5ne7n — The detailed critique directly contextualizes whether ARC-AGI-3 measures novel problem-solving or transferable priors and exploration heuristics.
2026-07-24T21:21:13Z
evidence attached: reddit.post.1v5nrvd — This raises a material benchmark-validity concern about whether ARC-AGI-3 rewards harness loops rather than intrinsic model capability.
2026-07-23T02:23:58Z
No independent evaluation or implementation evidence has appeared; the extraordinary public-set score remains a single-source, leakage-sensitive claim. With attention no longer moving, this is now a cold verification watch rather than an active signal.
2026-07-20T04:54:56Z
grounded: converges/medium — The Schema claim exemplifies Scott's specification-gaming and hidden-gates concerns: 99% on a public set without model-weight changes is exactly the pattern his
2026-07-19T11:24:20Z
case created — The extraordinary harness-only score is highly relevant to agent orchestration and straightforward to challenge or reproduce.