CAIS released HLE-Diamond, a revised Humanity's Last Exam, claiming restored headroom as the original HLE saturates; leaderboard and lab migration to it would make it the reference hard frontier-capability benchmark.
state: corroboratedheat: mediumuncertainty: mediumconvergesscott: mediumllm-evaluation benchmark-saturationCAIS
Surfaced 2026-09-25T18:02:03Z β CAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement). β The announcement burst is fully spent: at 48h both objects are inert (0 pts/h, 17th peer percentile, comment count frozen) and the velocity spikes were just the original spread replaying. No new evidence, no leaderboard or lab adoption β the case's meaning shifts from 'live release news' to 'dormant dated receipt for the benchmark-saturation radars, waiting on a slow adoption signal'. Engagement decay here is not evidence against the hypothesis.
What is this?
The Center for AI Safety (CAIS), which co-created Humanity's Last Exam (HLE) with Scale AI β a 2,500-question expert benchmark published in Nature in January 2026 β has released HLE-Diamond, a refined 1,000-question subset drawn from HLE's main pool and held-out reserve after a year-long cleaning process, hosted on Scale Labs' leaderboard with the stated aim of restoring measurement headroom as HLE saturates (third-party writeups put top-model scores past 50% in 2026, top rows below 65%, while MMLU and GPQA approach their ceilings). CAIS previously attempted a dynamic fix via an HLE-Rolling fork (Oct 2025); Diamond is instead a static cleaned set, and community reaction already frames HLE as 'benchmaxed.' The snippets corroborate that an operational Diamond leaderboard exists at Scale Labs β partially resolving the earlier 'who runs the leaderboard' caveat β but the lone Reddit report of GPT-6.1 Sol at third place remains unverified, and the snippets neither confirm the claimed Epoch AI critique of original HLE, the 'high reasoning' evaluation protocol, nor Diamond's grading method (the LLM-judge grading described on Artificial Analysis applies to original HLE).
Why it matters to Scott
CAIS's remedy for a 'benchmaxed' public gate β restore headroom by shipping a cleaned set drawn from the held-out reserve β is the exact held-out-gate discipline Hidden Gates canonizes, giving Scott a dated field-scale receipt for the ebook's argument; and the retreat from the dynamic HLE-Rolling fork to a static Diamond set hands him a real extension: a restocked answer key buys time without fixing the specification-gaming loop, which is the gap his process-gate alternative addresses. The GPT-6.1 Sol leaderboard report is single-source and unverified, so this corroborates that migration has begun without establishing reference-benchmark status β a publishing opportunity, not a change to what he builds.
ip:framework.hidden-gates-frameworkip:source.hidden-gates-ebookip:concept.specification-gamingradar:ai-benchmark-saturation-distortionradar:cleaned-benchmarks-frontier-rankingsradar:concept.benchmark-saturationradar:concept.benchmark-integrityradar:concept.llm-evaluation
queries asked of Scott's wikis
- benchmark saturation Goodhart metric gaming
- LLM-as-judge grading reliability eval harness
- training-on-test contamination canary strings
- held-out question reserve dynamic benchmark refresh
- eval protocol labels tool-assisted vs closed-book runs
- frontier model leaderboard comparison infrastructure
Measured heat
now 0 pts/hpeak 20 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 431h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p84 vs 1032 stories at the 336h mark (now 431h old) β ahead of baseten-harbor-admin-token-exposure (1.0x), behind github-hydrafusion-multi-model-orchestration (1.0x)
Evidence (3) β β canonical anchor
Interpretation history
2026-09-30T22:59:17Z
grounded: converges/medium β CAIS's remedy for a 'benchmaxed' public gate β restore headroom by shipping a cleaned set drawn from the held-out reserve β is the exact held-out-gate disciplin
2026-09-30T22:51:36Z
The awaited migration signal has begun: a frontier model (GPT-6.1 Sol) is reported third on a live HLE-Diamond leaderboard β single-source and low-traction, so it corroborates the release and confirms an operational leaderboard without establishing reference status; the case shifts from dormant dated receipt to early-adoption watch. The magnitude-valve reading reflects the spent announcement burst, not current spread, which is one quiet post β hence medium, not high, heat.
2026-09-30T21:37:57Z
evidence attached: reddit.post.1wudxzg β A frontier model posting a top-three Diamond result is direct evidence of lab/leaderboard migration, the exact resolution criterion of that case.
2026-09-25T18:00:33Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-23T22:43:43Z
grounded: known/medium β The radar already tracks this exact development in open cases: radar:ai-benchmark-saturation-distortion (saturated benchmarks driving adoption of harder, satura
2026-09-23T22:38:01Z
case created β High-spread (292/90) release of a credible benchmark revision with a clear resolution condition β adoption by labs and leaderboards.
Decision trace
- 10-01 08:59repriceThe awaited migration signal has begun: a frontier model (GPT-6.1 Sol) is reported third on a live HLE-Diamond leaderboard β single-source and low-traction, so it corroborates the release and confirms
- 10-01 08:59groundCAIS's remedy for a 'benchmaxed' public gate β restore headroom by shipping a cleaned set drawn from the held-out reserve β is the exact held-out-gate discipline Hidden Gates canonizes,
- 10-01 07:37attachA frontier model posting a top-three Diamond result is direct evidence of lab/leaderboard migration, the exact resolution criterion of that case.
- 10-01 07:32propose_attachA frontier model posting a top-three Diamond result is direct evidence of lab/leaderboard migration, the exact resolution criterion of that case.
- 09-26 04:02pushCAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement). β The announcement burst is fully spent: at 48h both objects are inert (0 pts/h, 17th peer
- 09-26 04:00repriceThe announcement burst is fully spent: at 48h both objects are inert (0 pts/h, 17th peer percentile, comment count frozen) and the velocity spikes were just the original spread replaying. No new evide
- 09-26 04:00alert_heldCAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement). β The announcement burst is fully spent: at 48h both objects are inert (0 pts/h, 17th peer
- 09-26 04:00alert_routeCAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement). β The announcement burst is fully spent: at 48h both objects are inert (0 pts/h, 17th peer
- 09-25 04:21sensor_dirtyvelocity_spike
- 09-24 18:20sensor_dirtyvelocity_spike
- 09-24 10:21sensor_dirtyvelocity_spike
- 09-24 08:43groundThe radar already tracks this exact development in open cases: radar:ai-benchmark-saturation-distortion (saturated benchmarks driving adoption of harder, saturation-resistant replacements) and radar:c
- 09-24 08:38createHigh-spread (292/90) release of a credible benchmark revision with a clear resolution condition β adoption by labs and leaderboards.