2026-10-11 17:09 UTC

benchmark-saturation

band: coolmomentum: stable score: 0.191
temperature history

Episodes (4)

Independent replication and evaluator response will determine whether widely used AI benchmarks are saturated enough to materially distort model comparisons and drive adoption of harder, saturation-resistant evaluations.
resolvedknownscott: low
CAIS released HLE-Diamond, a revised Humanity's Last Exam, claiming restored headroom as the original HLE saturates; leaderboard and lab migration to it would make it the reference hard frontier-capability benchmark.
corroboratedconvergesscott: medium
Reddit user we_are_mammals reports Kaggle's ARC-AGI-3 top scores jumped from 7% to 56% within 30 days โ€” achieved by small local models in harnesses, the only compute Kagglers may use โ€” crossing average-human performance on a benchmark designed to favor humans; disclosed methods and scores that hold under scrutiny confirm harness-driven rule-learning as a real generalization step on local models, while an exposed scoring exploit closes it as benchmark gaming.
watchingconvergesscott: high
Epoch AI's FrontierMath chart shows Tier 3 fully saturated within two years of Fields medalists (Tao, Gowers, Borcherds) calling it 'exceptionally challenging' and Tao predicting it would 'resist AIs for several years' โ€” with Tier 4 reportedly saturated too โ€” confirming that expert-curated frontier benchmarks now lose screening power inside two years; corroboration of the Tier 4 result and identification of which models cleared it resolve the episode.
resolvedconvergesscott: high

Trajectory notes