2026-10-11 16:37 UTC

CAIS released HLE-Diamond, a revised Humanity's Last Exam, claiming restored headroom as the original HLE saturates; leaderboard and lab migration to it would make it the reference hard frontier-capability benchmark.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: mediumllm-evaluation benchmark-saturationCAIS
Surfaced 2026-09-25T18:02:03Z β€” CAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement). β€” The announcement burst is fully spent: at 48h both objects are inert (0 pts/h, 17th peer percentile, comment count frozen) and the velocity spikes were just the original spread replaying. No new evidence, no leaderboard or lab adoption β€” the case's meaning shifts from 'live release news' to 'dormant dated receipt for the benchmark-saturation radars, waiting on a slow adoption signal'. Engagement decay here is not evidence against the hypothesis.

What is this?

The Center for AI Safety (CAIS), which co-created Humanity's Last Exam (HLE) with Scale AI β€” a 2,500-question expert benchmark published in Nature in January 2026 β€” has released HLE-Diamond, a refined 1,000-question subset drawn from HLE's main pool and held-out reserve after a year-long cleaning process, hosted on Scale Labs' leaderboard with the stated aim of restoring measurement headroom as HLE saturates (third-party writeups put top-model scores past 50% in 2026, top rows below 65%, while MMLU and GPQA approach their ceilings). CAIS previously attempted a dynamic fix via an HLE-Rolling fork (Oct 2025); Diamond is instead a static cleaned set, and community reaction already frames HLE as 'benchmaxed.' The snippets corroborate that an operational Diamond leaderboard exists at Scale Labs β€” partially resolving the earlier 'who runs the leaderboard' caveat β€” but the lone Reddit report of GPT-6.1 Sol at third place remains unverified, and the snippets neither confirm the claimed Epoch AI critique of original HLE, the 'high reasoning' evaluation protocol, nor Diamond's grading method (the LLM-judge grading described on Artificial Analysis applies to original HLE).

Why it matters to Scott

CAIS's remedy for a 'benchmaxed' public gate β€” restore headroom by shipping a cleaned set drawn from the held-out reserve β€” is the exact held-out-gate discipline Hidden Gates canonizes, giving Scott a dated field-scale receipt for the ebook's argument; and the retreat from the dynamic HLE-Rolling fork to a static Diamond set hands him a real extension: a restocked answer key buys time without fixing the specification-gaming loop, which is the gap his process-gate alternative addresses. The GPT-6.1 Sol leaderboard report is single-source and unverified, so this corroborates that migration has begun without establishing reference-benchmark status β€” a publishing opportunity, not a change to what he builds.
ip:framework.hidden-gates-frameworkip:source.hidden-gates-ebookip:concept.specification-gamingradar:ai-benchmark-saturation-distortionradar:cleaned-benchmarks-frontier-rankingsradar:concept.benchmark-saturationradar:concept.benchmark-integrityradar:concept.llm-evaluation
queries asked of Scott's wikis
  • benchmark saturation Goodhart metric gaming
  • LLM-as-judge grading reliability eval harness
  • training-on-test contamination canary strings
  • held-out question reserve dynamic benchmark refresh
  • eval protocol labels tool-assisted vs closed-book runs
  • frontier model leaderboard comparison infrastructure

Measured heat

now 0 pts/hpeak 20 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 431h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-23 22:38 (minted)⭐ origin echo-reconstructedCAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement).
CAIS on blog (echo) Β· attributed from reddit.post.1wobyjp Β· published time unknown
β€”
09-23 17:10first on r/singularity Β· published Β· lag ?Updated version of Humanity’s Last Exam: HLE-Diamond
socoolandawesome
β€”
09-23 17:10amplified on r/singularity πŸ‘‘reddit.post.1wobyjp
socoolandawesome
peak 359 Β· 101 comments Β· 85% of case engagement
09-30 19:13amplified on r/singularityreddit.post.1wudxzg
141_1337
peak 75 Β· 4 comments Β· 15% of case engagement
09-23 21:22our radar first saw it Β· lag ?discovery anchor: reddit.post.1wobyjpβ€”
09-25 18:00reached heat=high Β· lag ? Β· via ledgerβ€”β€”
pace: p84 vs 1032 stories at the 336h mark (now 431h old) β€” ahead of baseten-harbor-admin-token-exposure (1.0x), behind github-hydrafusion-multi-model-orchestration (1.0x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditUpdated version of Humanity’s Last Exam: HLE-Diamond
singularity
socoolandawesome351101
🟧 echo.blog ⭐CAIS announces HLE-Diamond, an updated version of Humanity's Last Exam (echoed via its X announcement).CAISβ€”β€”
🟠 redditGPT-6.1 Sol reaches third place in Humanity's Last Exam β€” Diamond
singularity
141_1337734

Interpretation history

Decision trace