Reddit user we_are_mammals reports Kaggle's ARC-AGI-3 top scores jumped from 7% to 56% within 30 days — achieved by small local models in harnesses, the only compute Kagglers may use — crossing average-human performance on a benchmark designed to favor humans; disclosed methods and scores that hold under scrutiny confirm harness-driven rule-learning as a real generalization step on local models, while an exposed scoring exploit closes it as benchmark gaming.
state: watchingheat: mediumuncertainty: highconvergesscott: highagent-harnesses benchmark-saturation arc-agi local-modelswe_are_mammals
What is this?
ARC-AGI-3 is the ARC Prize Foundation's interactive reasoning benchmark: novel video-game-like environments where an agent gets no instructions and must explore, infer the rules, and act efficiently, calibrated so human testers solve 100% while frontier LLM agents scored under 1% at launch. The ARC Prize 2026 Kaggle competition runs on this benchmark under code-competition compute restrictions — meaning small open-weight models on provided hardware rather than frontier APIs — and the official ARC Prize blog already documents harness-style submissions as the leading approach (e.g. Tufa Labs' small open-source LLM driving a live Python REPL). The supplied sources confirm the setting and show scores moving extremely fast under harnesses — Goertzel reports ~33% human-normalized on public games, Symbolica ~36%, and a Claude harness claims 99% on the public set with a 'fixed fallback' caveat that smells of exploit — but none of the snippets independently verify the specific 7%→56% Kaggle leaderboard jump or its cause, so the case's either/or (real harness-driven generalization on local models vs a scoring exploit) remains open on this evidence.
Why it matters to Scott
A live field test of his model-plus-harness benchmark-unit thesis on the benchmark engineered to resist it, run on exactly the small local models he serves from his own stack (Ollama/gamepc) and echoing the interactive rule-learning he labs on (Snake DQN). The case forks into two canon landings: scores holding = dated field receipts that harness, not weights, is the unit of capability; an exposed exploit = a Specification Gaming/Hidden Gates case study — either branch feeds an argument he already makes, and it continues the radar's Schema 99% / Seed-IQ ARC-AGI-3 claim lineage and the Kaggle prize-integrity dispute.
ip:concept.model-plus-harness-benchmark-unitip:source.give-the-agent-a-workshop-ebookip:concept.specification-gamingdev:project.snakedev:technology.ollamaradar:concept.arc-agiradar:concept.agent-harnessesradar:concept.local-modelsradar:concept.small-modelsradar:concept.benchmark-saturationradar:concept.benchmark-integrityradar:schema-arc-agi-3-claimradar:seed-iq-arc-agi-3d-doomradar:kaggle-agi-prize-integrity-dispute
queries asked of Scott's wikis
- harness scaffolding vs raw model capability
- agent memory context eviction long-horizon sessions
- small local open-weight models closed via tool loops
- benchmark saturation eval integrity gaming
- coding agent REPL live environment iteration loop
- interactive environments on-the-fly rule learning agents
Measured heat
now 0 pts/hpeak 31 pts/hcomments 0/hpeers p50momentum: steady1 platformsage 174h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p77 vs 1188 stories at the 168h mark (now 174h old) — ahead of ship-harness-bench (1.0x), behind cloudflare-cf-agentic-cli (1.0x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-10-05T09:38:30Z
Discussion crossed from hype-echo to verification phase: leaderboard participant huikang revealed a milestone-prize mechanism rewarding disclosed solutions and posted a comparison of them (Kaggle discussion 744792) — the case's first concrete path through its harness-vs-exploit fork. Meanwhile the 'two-subreddit echo' turns out to be same-author crossposts, not independent corroboration, and the burst has cooled ~10x from peak.
2026-10-04T11:37:37Z
grounded: converges/high — A live field test of his model-plus-harness benchmark-unit thesis on the benchmark engineered to resist it, run on exactly the small local models he serves from
2026-10-04T11:28:32Z
case created — Same-day two-subreddit echo of an eightfold leaderboard movement on an anti-saturation benchmark attributed to local-model harnesses — a bounded capability-movement episode no open case carries.
Decision trace
- 10-08 10:40review_screenjev screen: no material development (noul=0.05)
- 10-06 07:22sensor_dirtycomment_update
- 10-05 20:38repriceDiscussion crossed from hype-echo to verification phase: leaderboard participant huikang revealed a milestone-prize mechanism rewarding disclosed solutions and posted a comparison of them (Kaggle disc
- 10-05 06:21sensor_dirtycomment_update
- 10-05 02:21sensor_dirtycomment_update
- 10-04 23:22sensor_dirtyvelocity_spike
- 10-04 22:37groundA live field test of his model-plus-harness benchmark-unit thesis on the benchmark engineered to resist it, run on exactly the small local models he serves from his own stack (Ollama/gamepc) and echoi
- 10-04 22:28createSame-day two-subreddit echo of an eightfold leaderboard movement on an anti-saturation benchmark attributed to local-model harnesses — a bounded capability-movement episode no open case carries.
- 10-04 22:25propose_attachCross-community echo (singularity crosspost) of the same Kaggle ARC-AGI-3 jump — same-day independent spread for the new case.