2026-10-11 16:38 UTC

Epoch AI claims its new innovation benchmark shows LLMs significantly trail human researchers on novel problem-solving, establishing a measured capability gap on open-ended research.

state: seedheat: lowuncertainty: mediumconvergesscott: highai-assisted-mathematics agent-evaluation frontier-model-capabilitiesEpoch AIJason Li

What is this?

Epoch AI is an AI research organization that develops benchmarks to measure AI capabilities, notably FrontierMath (mathematical reasoning) and the Epoch Capabilities Index (ECI). Recent announcements include FrontierMath Erdős (68 unsolved Erdős problems formalized in Lean, where GPT-6 Astra scored 3%) and a benchmarking hub tracking model performance over time. A third-party source (ITNews Asia) reports Epoch AI claims 'even the most advanced LLMs have scored under two percent on their new benchmark,' but the specific 'innovation benchmark' named in the case hypothesis is not directly identifiable in the supplied snippets — it may refer to FrontierMath Erdős, EBR-bench, or a newer unannounced benchmark. The core claim (LLMs trail humans on novel research problems) aligns with Epoch's stated focus on 'open mathematical problems whose solutions can be automatically verified' to 'track progress in mathematical reasoning even as FrontierMath Tier 4 saturates.'

Why it matters to Scott

Epoch AI — a major benchmarking org — is independently adopting the methodology Scott's frameworks advocate: open-verifiable research-frontier problems (FrontierMath Erdős) with human baselines, automated Lean verification, and capability-gap measurement on novel problem-solving. This validates 'Benchmarking the Wrong Unit', 'Human Baseline', 'Verification Loops', and his agent-evaluation ladder (trace-backed comparison, claim-bounded adversarial verification, version-bound assessment). A consequential other party arriving at Scott's position creates a dated-receipts publishing opportunity.
ip:concept.benchmarking-the-wrong-unitip:concept.human-baselineip:concept.verification-loopsip:framework.ambition-frontier-rubricip:concept.evaluation-driven-developmentdev:concept.trace-backed-agent-comparisondev:concept.claim-bounded-adversarial-verificationdev:concept.version-bound-ai-assessmentdev:project.remote-execwork:project.leverageairadar:frontiermath-tier3-saturationradar:epoch-price-of-thoughtradar:concept.ai-benchmarksradar:concept.benchmark-saturationradar:concept.ai-assisted-mathematicsradar:concept.agent-evaluationradar:concept.model-evaluationradar:kbr-ai-formalized-math-proofradar:lean-transformer-ai-proofsradar:proofatlas-moser-worm-lower-boundradar:anthropic-fermat-lean-formalization
queries asked of Scott's wikis
  • Epoch AI benchmark strategy: open-verifiable problems vs. static benchmarks
  • FrontierMath Erdős and research-frontier evaluation methodology
  • AI-assisted mathematics: automated verification of novel proofs
  • Agent evaluation on open-ended research tasks (not Q&A)
  • Capability gap measurement: human baselines on research benchmarks
  • Open-weight vs closed-model gap on research-level benchmarks

Measured heat

now 0 pts/hpeak 4 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 123h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

10-06 13:00⭐ origin echo-reconstructedEpoch AI's new innovation benchmark demonstrates LLMs still significantly trail human researchers on novel problem-solving.
EpochAIResearch on x (echo) · attributed from reddit.post.1x0j40e
—
10-08 05:53first on r/artificial · published · +40.9hEpoch AI shows LLMs still have a long way to go before matching human researchers on innovation
Eliv_nurotic
—
10-08 05:53amplified on r/artificial 👑reddit.post.1x0j40e
Eliv_nurotic
peak 7 · 21 comments · 100% of case engagement
10-08 06:36our radar first saw it · +41.6hdiscovery anchor: reddit.post.1x0j40e—
pace: p56 vs 1247 stories at the 96h mark (now 123h old) — ahead of agent-screenshot-data-leaks (1.0x), behind complex-kda-recurrent-release (1.0x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditEpoch AI shows LLMs still have a long way to go before matching human researchers on innovation
artificial
Retrieved article excerpt

Open article · Retrieved 2026-10-08T07:01:48.356592+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Eliv_nurotic721
🟧 echo.x ⭐Epoch AI's new innovation benchmark demonstrates LLMs still significantly trail human researchers on novel problem-solving.EpochAIResearch——

Interpretation history

Decision trace