2026-10-11 17:11 UTC

Epoch AI's FrontierMath chart shows Tier 3 fully saturated within two years of Fields medalists (Tao, Gowers, Borcherds) calling it 'exceptionally challenging' and Tao predicting it would 'resist AIs for several years' — with Tier 4 reportedly saturated too — confirming that expert-curated frontier benchmarks now lose screening power inside two years; corroboration of the Tier 4 result and identification of which models cleared it resolve the episode.

state: resolvedheat: lowuncertainty: lowconvergesscott: highbenchmark-saturation agent-evaluation frontier-mathTerence TaoTimothy GowersRichard BorcherdsEpoch AI

What is this?

FrontierMath is Epoch AI's benchmark of unpublished, research-level mathematics problems built with 60+ mathematicians (lead mathematician Elliot Glazer) and organized into difficulty tiers; in Dec 2024 Fields medalists Terence Tao, Timothy Gowers, and Richard Borcherds publicly rated Tier 3 'exceptionally challenging,' with Tao predicting it would 'resist AIs for several years' — all corroborated by Epoch's own pages. The snippets show a steep trajectory since: leading systems went from <2% (late 2024) and <10% on Tier 4 (early 2026) to ~88% (Claude Fable 5) by June 2026, and multiple outlets (36kr, HTX) report Epoch officially concluding Tier 4 saturated after OpenAI's GPT-6 Astra scored 97.6%, solving every problem at least once. The snippets do not directly show the Tier 3 'fully saturated' chart the case's evidence asserts — Tier 4 saturation is the corroborated claim — and they add a validity wrinkle: the benchmark itself required a machine-assisted audit (GPT-5.5/Opus 4.7 screening plus mathematician review) that corrected or removed 19 of the 50 Tier 4 problems in a June 2026 v2, meaning the expert-curated problems contained errors that models helped surface.

Why it matters to Scott

Converges with his evaluation-validity canon and hands him dated receipts: the best-defended benchmark in the field (unpublished problems, expert-built, Fields-medalist-vetted) lost screening power inside two years of Tao's 'resist AIs for several years' prediction — direct receipts for his static-benchmark-decay / evaluation-shelf-life position (evaluation-driven-development), though the snippets corroborate Tier 4 saturation, not the asserted Tier 3 chart. Epoch's rescue — LLM screening plus mathematician review correcting 19 of 50 Tier 4 problems in a v2 — independently re-instantiates his LLM-assisted-audit / auditor-not-janitor pattern while raising a correlated-verifier question his mechanically-different-verifiers concept speaks to, and the v1→v2 regrade makes his version-bound-assessment discipline concretely necessary for any benchmark claim he cites in advisory work.
ip:concept.evaluation-driven-developmentip:concept.model-plus-harness-benchmark-unitip:concept.mechanically-different-verifiersip:concept.auditor-not-janitordev:concept.llm-rubric-gradingdev:concept.version-bound-ai-assessmentradar:ai-benchmark-saturation-distortionradar:hle-diamond-benchmarkradar:cleaned-benchmarks-frontier-rankingsradar:sous-physics-benchmark-regradingradar:openai-millennium-maths-claim
queries asked of Scott's wikis
  • benchmark saturation evaluation validity shelf life
  • agent eval harness design static benchmark decay
  • contamination held-out unpublished benchmark problems
  • expert-curated dataset quality errors LLM-assisted audit
  • frontier math reasoning as capability signal
  • coding agent regression eval suite

Measured heat

now 0 pts/hpeak 20 pts/hcomments 0/hpeers p29momentum: steady1 platformsage 68h
points/hour across evidence · reading as of 2026-10-08 10:28:17.051205+11:00 · deterministic, not a model opinion

How the heat travelled

10-05 03:50⭐ origin directly observedFrom Dec 2024: Fields medalists Terence Tao, Timothy Gowers, and Richard Borcherds characterized [FrontierMath] Tier 3 as “exceptionally challenging,” with Tao predicting it will “resist AIs for several years.” They were fully saturated <2 years later.
Eliv_nurotic on r/OpenAI
—
10-05 03:50amplified on r/OpenAI 👑reddit.post.1wxyrkz
Eliv_nurotic
peak 114 · 77 comments · 100% of case engagement
10-05 04:20our radar first saw it · +0.5hdiscovery anchor: reddit.post.1wxyrkz—

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐From Dec 2024: Fields medalists Terence Tao, Timothy Gowers, and Richard Borcherds characterized [FrontierMath] Tier 3 as “exceptionally challenging,” with Tao predicting it will “resist AIs for several years.” They were fully saturated <2 years later.
OpenAI
Retrieved article excerpt

Open article · Retrieved 2026-10-05T04:23:37.725917+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Eliv_nurotic11677

Interpretation history

Decision trace