2026-10-11 16:37 UTC

ai-assisted-mathematics

band: hotmomentum: stable score: 1.0
temperature history

Episodes (23)

Ensemble Prover’s maintainers claim their released open-source multi-agent Python system provides a usable workflow for autonomous theorem proving beyond isolated model-generated proof demonstrations.
expiredknownscott: low
ProofAtlas.ai and zero0_one1 claim a ProofAtlas harness using GPT-5.6 Pro improved the lower bound for Moser’s convex worm problem from 0.2322 to greater than 0.2374.
expiredconvergesscott: medium
Anthropic claims its released Lean 4 formalization of Fermat’s Last Theorem is a complete, machine-checkable artifact produced with substantial AI assistance, potentially establishing repository-scale formalization as a credible AI-assisted mathematics workflow.
corroboratedconvergesscott: high
Reddit user Artwastelander claims Claude Code (Fable 5.1) produced a valid proof that 31 three-point lines is the maximum for 15 points in the orchard problem, which would extend AI-assisted resolution of open combinatorial-geometry cases.
expiredknownscott: none
OpenAI reportedly claims a mathematical breakthrough on a Millennium Prize problem, potentially establishing substantive progress on a major open problem through AI-assisted research without yet establishing a complete solution.
acceleratingconvergesscott: high
Shengtao Guo, Ethan X. Fang, and Junwei Lu claim Odin discovered a proof giving a dimension-independent bound for the KomlΓ³s signing problem and square-root Beck–Fiala discrepancy, potentially establishing a major mathematical discovery by an AI research agent.
seednovelscott: low
HN user kbr- claims their published AI-assisted, formalized proof solves a 12-year-old mathematics problem, potentially establishing a new machine-checkable research result if the formal statement and proof substantiate the claimed solution.
seedknownscott: low
Dan Abramov claims his published AI-assisted Lean proof resolves Conway's refinement conjecture for omnific integers, potentially establishing a new mathematical result through a nonexpert-led agent workflow despite outstanding expert verification.
seedconvergesscott: medium
The newly formed independent Advisory Group on Mathematics and Artificial Intelligence says it is advising OpenAI on releasing a large batch of reportedly significant internal-model mathematics results and will publish recommendations, potentially establishing an externally visible release-governance process.
resolvedconvergesscott: medium
LinearSolveBench's maintainer claims the released benchmark measures whether coding agents can produce fast, accurate, and general C solvers for large sparse linear systems, extending agent evaluation beyond conventional repository tasks.
seedknownscott: low
Vals.ai claims ten Claude Sonnet 5.5 agents ran 15 hours of autonomous research and returned a 17,895-line machine-checkable Lean proof on the 1904 Thomson problem; verification of the artifact would establish sustained multi-agent formal-mathematics research outside frontier labs, while a flawed proof marks another inflated capability demo.
watchingnovelscott: high
Anthropic claims its Claude agent swarm produced a substantive numerical advance toward the Riemann Hypothesis β€” a claimed 50% result, since pushed to 67.25 under attempts to break it β€” now passing through Lean and expert-mathematician checking; verified acceptance of the bound establishes frontier agent teams as credible attackers of top-tier open problems, while a broken or retracted proof deflates the claim.
corroboratedconvergesscott: high
OpenAI releases a broad batch of mathematical results from its internal frontier model on GitHub with Lean proof formalizations, per-result compute estimates (~3 hours of ChatGPT Pro thinking on average), and release protocols developed with IAS's independent AGMAI β€” whether the math community verifies the results and other labs adopt the advised-release protocol resolves whether structured third-party-advised release of AI-generated mathematics becomes standard practice.
acceleratingconvergesscott: high
Epoch AI claims its new innovation benchmark shows LLMs significantly trail human researchers on novel problem-solving, establishing a measured capability gap on open-ended research.
seedconvergesscott: high
A. Dabrowski claims an AI-agent-produced paper and 238-module Lean 4 artifact prove the Hilbert-Smith conjecture, which would add a machine-checkable solution of a major mathematical problem if the formalization and underlying argument withstand expert review.
seednovelscott: medium
Flynn Bettens claims an AI-assisted flag-algebra proof determining the exact asymptotic constant for ErdΕ‘s Problem #1034 as (36βˆ’5√3)/66, with a Zenodo deposit including certificate, verifier, and independent Python verification script.
seednovelscott: low
ZWY-research releases Physics-to-Math Research, a bilingual agent skill for scientific-to-mathematical formulation, claim-integrity auditing, and bounded conditional derivation β€” a reusable component for AI-assisted mathematics workflows.
seedconvergesscott: high
Wenyu Du and Stephen Chung claim the Station environment with Supervisor and Meta Reflection mechanisms enables AI agents to rediscover 62.7% of criteria from held-out ICLR papers β€” if replicated, Station becomes a standard benchmark for open-ended scientific discovery by agents.
seedconvergesscott: high
Evolutionary LLM program search improves upper bounds for square packing, breaking 25 records including a 47-year-old record, demonstrating LLMs as optimization tools for combinatorial problems.
watchingconvergesscott: high
CNBC interview with NYU professor Tristan Buckmaster alleges OpenAI's AGMAI math results were trained on non-consensual user data, triggering privacy investigations or community rejection of the release protocol.
seedconvergesscott: high
A 9th-grade student claims to have used Claude to prove the rhombicosidodecahedron cannot pass through a copy of itself, using interval arithmetic over 12.3M boxes and 38 CPU hours to verify singular configurations at √5.
seednovelscott: low
Benjaminsen's SolveAtHome platform launches as a public venue for collaborative MD5 cryptanalysis using Claude Code, testing whether AI-assisted cryptanalysis can sustain a community-driven research effort.
seednovelscott: low
Terence Tao claims in his 'Math 2.0' presentation that AI will fundamentally transform mathematical research practice, and the talk's reception shapes community expectations and research directions for AI-assisted mathematics.
watchingconvergesscott: high

Trajectory notes