2026-10-11 16:38 UTC

Reddit user we_are_mammals reports Kaggle's ARC-AGI-3 top scores jumped from 7% to 56% within 30 days — achieved by small local models in harnesses, the only compute Kagglers may use — crossing average-human performance on a benchmark designed to favor humans; disclosed methods and scores that hold under scrutiny confirm harness-driven rule-learning as a real generalization step on local models, while an exposed scoring exploit closes it as benchmark gaming.

state: watchingheat: mediumuncertainty: highconvergesscott: highagent-harnesses benchmark-saturation arc-agi local-modelswe_are_mammals

What is this?

ARC-AGI-3 is the ARC Prize Foundation's interactive reasoning benchmark: novel video-game-like environments where an agent gets no instructions and must explore, infer the rules, and act efficiently, calibrated so human testers solve 100% while frontier LLM agents scored under 1% at launch. The ARC Prize 2026 Kaggle competition runs on this benchmark under code-competition compute restrictions — meaning small open-weight models on provided hardware rather than frontier APIs — and the official ARC Prize blog already documents harness-style submissions as the leading approach (e.g. Tufa Labs' small open-source LLM driving a live Python REPL). The supplied sources confirm the setting and show scores moving extremely fast under harnesses — Goertzel reports ~33% human-normalized on public games, Symbolica ~36%, and a Claude harness claims 99% on the public set with a 'fixed fallback' caveat that smells of exploit — but none of the snippets independently verify the specific 7%→56% Kaggle leaderboard jump or its cause, so the case's either/or (real harness-driven generalization on local models vs a scoring exploit) remains open on this evidence.

Why it matters to Scott

A live field test of his model-plus-harness benchmark-unit thesis on the benchmark engineered to resist it, run on exactly the small local models he serves from his own stack (Ollama/gamepc) and echoing the interactive rule-learning he labs on (Snake DQN). The case forks into two canon landings: scores holding = dated field receipts that harness, not weights, is the unit of capability; an exposed exploit = a Specification Gaming/Hidden Gates case study — either branch feeds an argument he already makes, and it continues the radar's Schema 99% / Seed-IQ ARC-AGI-3 claim lineage and the Kaggle prize-integrity dispute.
ip:concept.model-plus-harness-benchmark-unitip:source.give-the-agent-a-workshop-ebookip:concept.specification-gamingdev:project.snakedev:technology.ollamaradar:concept.arc-agiradar:concept.agent-harnessesradar:concept.local-modelsradar:concept.small-modelsradar:concept.benchmark-saturationradar:concept.benchmark-integrityradar:schema-arc-agi-3-claimradar:seed-iq-arc-agi-3d-doomradar:kaggle-agi-prize-integrity-dispute
queries asked of Scott's wikis
  • harness scaffolding vs raw model capability
  • agent memory context eviction long-horizon sessions
  • small local open-weight models closed via tool loops
  • benchmark saturation eval integrity gaming
  • coding agent REPL live environment iteration loop
  • interactive environments on-the-fly rule learning agents

Measured heat

now 0 pts/hpeak 31 pts/hcomments 0/hpeers p50momentum: steady1 platformsage 174h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-04 10:24⭐ origin directly observedTop ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]
we_are_mammals on r/MachineLearning
—
10-04 10:35first on r/singularity · published · +0.2hTop ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]
we_are_mammals
—
10-04 10:24amplified on r/MachineLearning 👑reddit.post.1wxcd4k
we_are_mammals
peak 129 · 74 comments · 80% of case engagement
10-04 10:35amplified on r/singularityreddit.post.1wxcjkc
we_are_mammals
peak 47 · 4 comments · 20% of case engagement
10-04 11:20our radar first saw it · +0.9hdiscovery anchor: reddit.post.1wxcd4k—
pace: p77 vs 1188 stories at the 168h mark (now 174h old) — ahead of ship-harness-bench (1.0x), behind cloudflare-cf-agentic-cli (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Top ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]
MachineLearning
we_are_mammals12874
🟠 redditTop ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]
singularity
we_are_mammals474

Interpretation history

Decision trace