Epoch AI's FrontierMath chart shows Tier 3 fully saturated within two years of Fields medalists (Tao, Gowers, Borcherds) calling it 'exceptionally challenging' and Tao predicting it would 'resist AIs for several years' — with Tier 4 reportedly saturated too — confirming that expert-curated frontier benchmarks now lose screening power inside two years; corroboration of the Tier 4 result and identification of which models cleared it resolve the episode.
state: resolvedheat: lowuncertainty: lowconvergesscott: highbenchmark-saturation agent-evaluation frontier-mathTerence TaoTimothy GowersRichard BorcherdsEpoch AI
What is this?
FrontierMath is Epoch AI's benchmark of unpublished, research-level mathematics problems built with 60+ mathematicians (lead mathematician Elliot Glazer) and organized into difficulty tiers; in Dec 2024 Fields medalists Terence Tao, Timothy Gowers, and Richard Borcherds publicly rated Tier 3 'exceptionally challenging,' with Tao predicting it would 'resist AIs for several years' — all corroborated by Epoch's own pages. The snippets show a steep trajectory since: leading systems went from <2% (late 2024) and <10% on Tier 4 (early 2026) to ~88% (Claude Fable 5) by June 2026, and multiple outlets (36kr, HTX) report Epoch officially concluding Tier 4 saturated after OpenAI's GPT-6 Astra scored 97.6%, solving every problem at least once. The snippets do not directly show the Tier 3 'fully saturated' chart the case's evidence asserts — Tier 4 saturation is the corroborated claim — and they add a validity wrinkle: the benchmark itself required a machine-assisted audit (GPT-5.5/Opus 4.7 screening plus mathematician review) that corrected or removed 19 of the 50 Tier 4 problems in a June 2026 v2, meaning the expert-curated problems contained errors that models helped surface.
Why it matters to Scott
Converges with his evaluation-validity canon and hands him dated receipts: the best-defended benchmark in the field (unpublished problems, expert-built, Fields-medalist-vetted) lost screening power inside two years of Tao's 'resist AIs for several years' prediction — direct receipts for his static-benchmark-decay / evaluation-shelf-life position (evaluation-driven-development), though the snippets corroborate Tier 4 saturation, not the asserted Tier 3 chart. Epoch's rescue — LLM screening plus mathematician review correcting 19 of 50 Tier 4 problems in a v2 — independently re-instantiates his LLM-assisted-audit / auditor-not-janitor pattern while raising a correlated-verifier question his mechanically-different-verifiers concept speaks to, and the v1→v2 regrade makes his version-bound-assessment discipline concretely necessary for any benchmark claim he cites in advisory work.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitip:concept.mechanically-different-verifiersip:concept.auditor-not-janitordev:concept.llm-rubric-gradingdev:concept.version-bound-ai-assessmentradar:ai-benchmark-saturation-distortionradar:hle-diamond-benchmarkradar:cleaned-benchmarks-frontier-rankingsradar:sous-physics-benchmark-regradingradar:openai-millennium-maths-claim
queries asked of Scott's wikis
- benchmark saturation evaluation validity shelf life
- agent eval harness design static benchmark decay
- contamination held-out unpublished benchmark problems
- expert-curated dataset quality errors LLM-assisted audit
- frontier math reasoning as capability signal
- coding agent regression eval suite
Measured heat
now 0 pts/hpeak 20 pts/hcomments 0/hpeers p29momentum: steady1 platformsage 68h
points/hour across evidence · reading as of 2026-10-08 10:28:17.051205+11:00 · deterministic, not a model opinion
How the heat travelled
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-10-07T23:40:22Z
The grounding pass at birth already delivered the case's stated resolution condition — multiple outlets report Epoch officially concluding Tier 4 saturated after OpenAI's GPT-6 Astra scored 97.6% — and everything since is repetitive comment churn (+2 points, flat at 77 comments, 0 points/hour at 68h), so the episode closes as a proved-out instance of expert-benchmark decay rather than a live story. Banked with it: the June 2026 v2 regrade where LLM-assisted screening corrected 19 of 50 Tier 4 problems — the benchmark needed models to fix it before models beat it.
2026-10-05T04:42:27Z
grounded: converges/high — Converges with his evaluation-validity canon and hands him dated receipts: the best-defended benchmark in the field (unpublished problems, expert-built, Fields-
2026-10-05T04:32:42Z
case created — A named expert-anchored benchmark's documented sub-two-year saturation is a distinct resolvable evaluation-validity episode — separate from the physics-regrading and OpenAI math-advisory cases it brushes — born into the hot agent-evaluation topic with the OpenAI math-release story likely to move it.
Decision trace
- 10-08 10:40resolveThe grounding pass at birth already delivered the case's stated resolution condition — multiple outlets report Epoch officially concluding Tier 4 saturated after OpenAI's GPT-6 Astra scored
- 10-07 06:24sensor_dirtycomment_update
- 10-06 10:22sensor_dirtycomment_update
- 10-06 04:21sensor_dirtycomment_update
- 10-05 23:21sensor_dirtycomment_update
- 10-05 19:21sensor_dirtycomment_update
- 10-05 16:20sensor_dirtyvelocity_spike
- 10-05 15:42groundConverges with his evaluation-validity canon and hands him dated receipts: the best-defended benchmark in the field (unpublished problems, expert-built, Fields-medalist-vetted) lost screening power in
- 10-05 15:32createA named expert-anchored benchmark's documented sub-two-year saturation is a distinct resolvable evaluation-validity episode — separate from the physics-regrading and OpenAI math-advisory cases it