2026-10-11 17:20 UTC

Independent replication and evaluator response will determine whether widely used AI benchmarks are saturated enough to materially distort model comparisons and drive adoption of harder, saturation-resistant evaluations.

state: resolvedheat: lowuncertainty: lowknownscott: lowagent-benchmarks benchmark-saturation

What is this?

A new arXiv systematic study examines benchmark saturation: the loss of meaningful differentiation when top AI systems converge on near-identical scores. Supplied sources indicate that this problem already affects widely used evaluations such as HumanEval and has prompted researchers to develop harder or more durable tests, but the snippets do not identify the study’s authors or provide its specific findings. Whether its conclusions materially change model comparisons or evaluator practice remains an open hypothesis requiring replication and response from the evaluation community.

Why it matters to Scott

Scott already argues that static, visible, or poorly scoped benchmarks can stop discriminating real system capability, and he builds progressive, replay-based evaluation harnesses as the alternative; see “Benchmarking the Wrong Unit” and “Progressive Evaluation Ladder.” The study could influence benchmark selection in his provider-execution work and strengthen that published position, but the supplied material gives no specific findings or evaluator response yet, while the radar already tracks benchmark saturation and integrity as established themes.
ip:concept.benchmarking-the-wrong-unitip:concept.progressive-evaluation-ladderip:framework.hidden-gates-frameworkdev:project.remote-execradar:concept.benchmark-integrityradar:concept.agent-benchmarksradar:concept.model-evaluation
queries asked of Scott's wikis
  • benchmark saturation in coding-agent evaluations
  • dynamic evaluations versus static benchmarks
  • benchmark gaming and test-set contamination
  • evaluation harnesses for real-world agent capability
  • model adoption decisions based on benchmark validity
  • saturation-resistant benchmarks and task generation

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (60) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnWhen AI Benchmarks Plateau: A Systematic Study of Benchmark Saturationdoppp104131
🟧 echo.paper ⭐A systematic study examines when AI benchmarks plateau and whether saturation undermines meaningful capability comparisons.paper authors——
🟧 hnLaunch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs ResearchMzzzzz3934
🟧 hnYour model already knows the answer: how benchmark answers leak into LLMsfran-mora100
🟧 hnHalf Our SLM Benchmark 'Failures' Contained the Right Answerrobmay20
🟧 hnGoodhart's Law Comes for Every Benchmark You Trustpseudolus10043
🟠 redditIntroducing BetterBench - more accurate PP and TPS measurement
LocalLLaMA
whodoneit12115
🟠 redditHow come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode
LocalLLaMA
Informal-Trouble21833197
🟠 redditThe current state of language models and human preference based rankings [R]
MachineLearning
adam_alpha_finetuner71
🟧 hnSciCode-Verified: How Benchmark Defects Underestimated LLM Scientific-Codingsbulaev10
🟧 hnAgentic test processes, LLM benchmarks, and other noteskqr20
🟧 hnShow HN: ThinkingType, an eval measuring if fonts change VLM judgmentsbwuckner-sf11
🟧 hnThe Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performancesbulaev10
🟠 redditTerminal Bench 3 has been released. It’s a new benchmark that hasn’t been included in model training sets yet. (I’m not showing the results from third-party harnesses to keep things fair.)
singularity
Distinct_Fox_635821238
🟧 hnAnthropic: Introducing The Conceptual Reasoning Indexoptimalsolver7652
🟠 redditAnthropic: Introducing The Conceptual Reasoning Index
singularity
EducationalCicada37464
🟠 redditAnthropic: Introducing The Conceptual Reasoning Index
ClaudeAI
Enough-Plantain27854030
🟧 hnShow HN: Artificial Analysis tool to create custom benchmarks for any use caseGcam60
🟧 hnAgents on Rails: The LLM Benchmark Projecthahahacorn40
🟧 hnShow HN: Self-bench – build SWE-bench style evals from private reposbyhong0320
🟧 hnAgents on Rails: The LLM Benchmark Projectksec10
🟧 hnSkills x2, 108 Runs: Optimising Efficacy and Token Efficiencydarvh10
🟧 hnShow HN: Stressing LLMs – Complexity Benchmark__alexander10
🟠 redditBenchmarks don't mean anything anymore.
LocalLLaMA
Nerfariox059
🟧 hnThe Benchmarkpocalypsecyndunlop17964
🟧 hnSix Ways an Eval Lies: structural defects that survive preregistrationThreadborne10
🟧 hnEvaluating AI Agents Live at the Grounded Reasoning Cupiwhalen20
🟠 redditWhen will scicode be saturated?
singularity
Worldly_Beginning647263
🟧 hnNvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmarkdsrtslnd236737
🟧 hnProgramBench Vetted: Reverse Engineering from a Runnable Binaryrigelbm271
🟧 hnGSM8K problems can't tell Haiku 4.5 from Opus 5Bhagichan20
🟧 hnReconstructing the benchmark behind Luc Julia's 64% LLM reliability claimmatthieu_bl20
🟠 redditGemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks?
LocalLLaMA
uncle_leon6293
🟠 redditLocal benchmarking isn't as easy as it seems
LocalLLaMA
KitchenAmoeba443800
🟠 redditA dataset with 52 Text to image model evaluation [P]
MachineLearning
dh7net36
🟠 redditYou can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm [R]
MachineLearning
eamonnkeogh49133
🟠 redditEvery benchmarks gets saturated after certain period of time, then why is HLE not yet saturated?
OpenAI
Lucky_Creme_520813
🟠 redditEvery benchmarks gets saturated after certain period of time, then why is HLE not yet saturated?
singularity
Lucky_Creme_520805
🟠 redditYour GNN is probably just an overcomplicated MLP (Tabular Leakage). We built SynthFin-AML to enforce strict causal boundaries. [P]
MachineLearning
Glabmayt207568
🟧 hn44% on ARC-AGI-1 in 67 centsporridgeraisin662149
🟠 redditIn regards to benchmaxxing...
LocalLLaMA
Nevermore1215615
🟠 redditLit Review on Running GUI Agents on phone: AndroidWorld
LocalLLaMA
East-Muffin-647220
🟧 hnFlavourbench: LLM Eval with Executable Culinary Ground Truthjosefchen10
🟠 redditHow an unsupported tool-call response could become “perfectly stable” in an LLM benchmark
artificial
docybo27
🟧 hnMatching Puzzle Pieces and Disappointing Benchmarksspeckx10
🟠 redditThe prevalent problem of misleading benchmark reporting (re: Astra)
singularity
PsychologicalSoup25110673
🟠 redditAny good alternatives to Artificial Analysis?
LocalLLaMA
metigue1227
🟠 redditAstra WITHOUT CoT gets 97% on ARC-AGI-3 and 86% on ARC-AGI-1
singularity
FeeAvailable377018962
🟠 redditArtificial Analysis Index is NOT Representative of real World Performance
LocalLLaMA
PerformanceRound7913625
🟠 redditWhat do you think ARC AGI 4 will be about?
singularity
ErmingSoHard2049
🟠 redditGPT-6 Astra gets 3% on the FrontierMath Erdős Benchmark, while every other Model(that was tested) got 0%
singularity
Every_Foundation5197672119
🟧 hnArtificial Analysis Intelligence Index v4.2nojs12953
🟠 redditAA Intelligence Index Changes
singularity
poigre5919
🟠 redditWhat the Artificial Analysis / GPT-6 Astra mess actually teaches us
LocalLLaMA
PerformanceRound7913024
🟠 redditWhat the Artificial Analysis / GPT-6 Astra mess actually teaches us
singularity
PerformanceRound791307
🟠 redditwhat benchmark to trust now?
singularity
TheReedemer69644
🟠 redditFrontier AI models are beginning to cross the human baseline on SimpleBench
singularity
Bojackin_Around26024
🟠 redditCoding benchmarks that are quickly showcasing deep capability
LocalLLaMA
Informal-Trouble21837731
🟧 hnRecreating Minecraft Is Not a Benchmarkkuberwastaken5641
🟧 hnMeasuring benchmark optimization in speech recognitiongmays20

Interpretation history

Decision trace