An audit of GPQA, MMLU-Pro, and MMMU-Pro reportedly found malformed questions, incorrect answer keys, and ambiguous answers, leading to the removal of up to 12% of items and the release of cleaned datasets. The supplied search results establish that these benchmarks are used to compare frontier models and may be approaching saturation, but they do not provide independent evaluations of the corrected versions or show whether scores and rankings materially change. Adam Allcock is named in the case, though the supplied snippets do not establish his specific role in the audit or release.
No intersection found: the supplied Scott wiki and radar searches returned no hits connecting benchmark-integrity audits or these corrected datasets to Scott’s established positions, projects, or tracked stories.
queries asked of Scott's wikis
- benchmark integrity and contaminated evaluation sets
- evaluation harnesses for corrected benchmark datasets
- frontier-model ranking sensitivity to benchmark errors
- benchmark saturation and capability ceilings
- dataset versioning and reproducible model evaluations
- alternatives to static benchmark leaderboards
2026-07-29T12:36:38Z
No independent evaluations using the corrected datasets have appeared since the audit was released. The case has stalled on repetitive amplification of the original findings with no progress toward the hypothesis. Expired.
2026-07-29T11:25:31Z
Still no independent reruns or score comparisons using the corrected datasets. The attached evidence remains the same originating audit, recirculated. The case is stuck in repetitive amplification — no progress toward the hypothesis.
2026-07-29T10:29:34Z
The newly attached evidence still provides no independent corrected-dataset reruns or ranking comparisons, so the benchmark audit’s practical effect remains unknown. Continued circulation of the originating release is repetitive amplification, not corroboration.
2026-07-29T08:28:56Z
No independent corrected-set reruns or ranking comparisons have appeared; the attached evidence still only repeats the originating audit. The case remains unresolved but should cool to a slower watch until actual evaluation results emerge.
2026-07-29T06:27:17Z
The attached evidence still originates from the audit and cleaned-dataset release, with no independent reruns showing effects on scores, rankings, or ceilings. Repeated recirculation adds no corroboration, so the case remains a low-temperature watch pending actual evaluation results.
2026-07-29T05:24:21Z
The newly attached material still offers no independent corrected-set reruns or score comparisons, leaving the effect on rankings and performance ceilings unknown. Further recirculation of the originating audit is repetitive amplification rather than corroboration.
2026-07-29T04:24:56Z
The attached evidence still contains no independent reruns or corrected-set score comparisons, so the audit’s impact on frontier rankings and apparent ceilings remains unresolved. Continued recirculation of the originating claim adds no substantive corroboration.
2026-07-29T03:25:38Z
The added evidence remains another observation of the originating audit rather than an independent rerun, so it does not clarify whether corrected datasets change frontier rankings or ceilings. Repeated amplification without evaluation results warrants a slower cadence.
2026-07-29T02:29:52Z
The evidence remains confined to the original audit and cleaned-dataset release; no independent reruns establish an effect on frontier scores, rankings, or ceilings. This is repetitive amplification rather than substantive corroboration.
2026-07-29T00:22:35Z
The newly attached material still derives from the original audit and cleaned-dataset release, adding no independent reruns or ranking comparisons. The benchmark defects appear reproducible, but their effect on frontier scores and ceilings remains unknown.
2026-07-28T22:26:15Z
The attached evidence still traces back to the original audit and release rather than independent reruns. The central question—whether cleaned datasets materially change frontier scores, rankings, or apparent ceilings—remains unanswered.
2026-07-28T21:22:38Z
The new observation is only negligible engagement growth around the original audit; no independent reruns yet show whether the cleaned datasets alter scores, rankings, or ceilings.
2026-07-28T20:22:40Z
grounded: novel/none — No intersection found: the supplied Scott wiki and radar searches returned no hits connecting benchmark-integrity audits or these corrected datasets to Scott’s
2026-07-28T20:21:55Z
case created — The audit provides reproducible corrected datasets for widely used benchmarks, but its effect on model comparisons still requires independent reruns.