Epoch AI claims its new innovation benchmark shows LLMs significantly trail human researchers on novel problem-solving, establishing a measured capability gap on open-ended research.
state: seedheat: lowuncertainty: mediumconvergesscott: highai-assisted-mathematics agent-evaluation frontier-model-capabilitiesEpoch AIJason Li
What is this?
Epoch AI is an AI research organization that develops benchmarks to measure AI capabilities, notably FrontierMath (mathematical reasoning) and the Epoch Capabilities Index (ECI). Recent announcements include FrontierMath Erdős (68 unsolved Erdős problems formalized in Lean, where GPT-6 Astra scored 3%) and a benchmarking hub tracking model performance over time. A third-party source (ITNews Asia) reports Epoch AI claims 'even the most advanced LLMs have scored under two percent on their new benchmark,' but the specific 'innovation benchmark' named in the case hypothesis is not directly identifiable in the supplied snippets — it may refer to FrontierMath Erdős, EBR-bench, or a newer unannounced benchmark. The core claim (LLMs trail humans on novel research problems) aligns with Epoch's stated focus on 'open mathematical problems whose solutions can be automatically verified' to 'track progress in mathematical reasoning even as FrontierMath Tier 4 saturates.'
Why it matters to Scott
Epoch AI — a major benchmarking org — is independently adopting the methodology Scott's frameworks advocate: open-verifiable research-frontier problems (FrontierMath Erdős) with human baselines, automated Lean verification, and capability-gap measurement on novel problem-solving. This validates 'Benchmarking the Wrong Unit', 'Human Baseline', 'Verification Loops', and his agent-evaluation ladder (trace-backed comparison, claim-bounded adversarial verification, version-bound assessment). A consequential other party arriving at Scott's position creates a dated-receipts publishing opportunity.
ip:concept.benchmarking-the-wrong-unitip:concept.human-baselineip:concept.verification-loopsip:framework.ambition-frontier-rubricip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisondev:concept.claim-bounded-adversarial-verificationdev:concept.version-bound-ai-assessmentdev:project.remote-execwork:project.leverageairadar:frontiermath-tier3-saturationradar:epoch-price-of-thoughtradar:concept.ai-benchmarksradar:concept.benchmark-saturationradar:concept.ai-assisted-mathematicsradar:concept.agent-evaluationradar:concept.model-evaluationradar:kbr-ai-formalized-math-proofradar:lean-transformer-ai-proofsradar:proofatlas-moser-worm-lower-boundradar:anthropic-fermat-lean-formalization
queries asked of Scott's wikis
- Epoch AI benchmark strategy: open-verifiable problems vs. static benchmarks
- FrontierMath Erdős and research-frontier evaluation methodology
- AI-assisted mathematics: automated verification of novel proofs
- Agent evaluation on open-ended research tasks (not Q&A)
- Capability gap measurement: human baselines on research benchmarks
- Open-weight vs closed-model gap on research-level benchmarks
Measured heat
now 0 pts/hpeak 4 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 123h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p56 vs 1247 stories at the 96h mark (now 123h old) — ahead of agent-screenshot-data-leaks (1.0x), behind complex-kda-recurrent-release (1.0x)
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-10-08T07:13:26Z
grounded: converges/high — Epoch AI — a major benchmarking org — is independently adopting the methodology Scott's frameworks advocate: open-verifiable research-frontier problems (Frontie
2026-10-08T07:03:40Z
case created — Single Reddit post referencing Epoch AI X announcement; plausible but needs corroboration.
Decision trace
- 10-11 22:09review_screenjev screen: no material development (noul=0.16)
- 10-09 09:34sensor_dirtycomment_update
- 10-08 19:31sensor_dirtycomment_update
- 10-08 18:31attention_routeThe editor compared this story and chose to keep watching.
- 10-08 18:25attention_candidatecreate
- 10-08 18:13groundEpoch AI — a major benchmarking org — is independently adopting the methodology Scott's frameworks advocate: open-verifiable research-frontier problems (FrontierMath Erdős) with human baselines,
- 10-08 18:03createSingle Reddit post referencing Epoch AI X announcement; plausible but needs corroboration.