2026-10-11 17:09 UTC

arc-agi

band: coolmomentum: stable score: 0.141
temperature history

Episodes (5)

Independent evaluation will determine whether Schema's process-only harness genuinely achieves 99% on the ARC-AGI-3 public set without model-weight changes or evaluation leakage.
expiredconvergesscott: medium
Independent evaluations will determine whether Seed IQ's reported 100% ARC-AGI 3 performance and 3D Doom II gameplay generalize to robust, generalizable 3D environment reasoning.
expirednovelscott: low
Independent evaluation will determine whether Orivael’s non-LLM reasoning system can reproduce its perfect ft09 result and generalize across additional ARC-AGI-3 task families.
expiredknownscott: low
The ARC-AGI Without Pretraining author claims ARC-AGI tasks can be solved competitively without pretrained foundation-model knowledge, challenging pretraining as a prerequisite for abstract reasoning.
expiredconvergesscott: medium
Reddit user we_are_mammals reports Kaggle's ARC-AGI-3 top scores jumped from 7% to 56% within 30 days β€” achieved by small local models in harnesses, the only compute Kagglers may use β€” crossing average-human performance on a benchmark designed to favor humans; disclosed methods and scores that hold under scrutiny confirm harness-driven rule-learning as a real generalization step on local models, while an exposed scoring exploit closes it as benchmark gaming.
watchingconvergesscott: high

Trajectory notes