2026-10-11 17:11 UTC

Independent evaluations will determine whether corrected versions of GPQA, MMLU-Pro, and MMMU-Pro materially change frontier-model scores, rankings, or apparent performance ceilings.

state: expiredheat: lowuncertainty: highnovelscott: nonebenchmark-integrity model-evaluation frontier-modelsAdam Allcock

What is this?

An audit of GPQA, MMLU-Pro, and MMMU-Pro reportedly found malformed questions, incorrect answer keys, and ambiguous answers, leading to the removal of up to 12% of items and the release of cleaned datasets. The supplied search results establish that these benchmarks are used to compare frontier models and may be approaching saturation, but they do not provide independent evaluations of the corrected versions or show whether scores and rankings materially change. Adam Allcock is named in the case, though the supplied snippets do not establish his specific role in the audit or release.

Why it matters to Scott

No intersection found: the supplied Scott wiki and radar searches returned no hits connecting benchmark-integrity audits or these corrected datasets to Scott’s established positions, projects, or tracked stories.
queries asked of Scott's wikis
  • benchmark integrity and contaminated evaluation sets
  • evaluation harnesses for corrected benchmark datasets
  • frontier-model ranking sensitivity to benchmark errors
  • benchmark saturation and capability ceilings
  • dataset versioning and reproducible model evaluations
  • alternatives to static benchmark leaderboards

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[PAPER] GPQA, MMLU-Pro, and MMMU-Pro were audited for broken questions, and up to 12% of them had to be removed. New drop in clean versions released
LocalLLaMA
pawofdoom447
🟠 reddit[PAPER] Major benchmarks are found to be polluted, with up to 12% of questions broken
singularity
pawofdoom8513
🟧 echo.github ⭐Publishes an audit and cleaned datasets reporting that malformed questions, incorrect answer keys, and ambiguous answers required removing uAdam Allcock——

Interpretation history

Decision trace