2026-10-11 18:00 UTC

benchmark-integrity

band: coolmomentum: stable score: 0.051
temperature history

Episodes (4)

Basalt Labs' claimed 99.44% HLE result will be shown to rely on a misleading model identity, evaluation setup, or served-model substitution.
expiredconvergesscott: medium
The DeepMind-sponsored Kaggle AGI benchmark prize will face a formal review or substantive rebuttal over allegations that the winning entry is nonsensical and unsupported.
expiredconvergesscott: low
Independent evaluations will determine whether corrected versions of GPQA, MMLU-Pro, and MMMU-Pro materially change frontier-model scores, rankings, or apparent performance ceilings.
expirednovelscott: none
LocalLLaMA user brainchillzZ reports that Gufo's headline GitHub benchmark โ€” 70.56 tok/s single-user Qwen3.8-27B Q4 on Strix Halo โ€” holds only under a degenerate prompt; Gufo correcting or clearly qualifying the benchmark methodology, or the community validating the number under representative prompts, resolves whether the project's marquee claim survives scrutiny or costs it adoption.
resolvedconvergesscott: medium

Trajectory notes