2026-10-11 17:15 UTC

Vals.ai claims ten Claude Sonnet 5.5 agents ran 15 hours of autonomous research and returned a 17,895-line machine-checkable Lean proof on the 1904 Thomson problem; verification of the artifact would establish sustained multi-agent formal-mathematics research outside frontier labs, while a flawed proof marks another inflated capability demo.

state: watchingheat: highuncertainty: mediumnovelscott: highai-assisted-mathematics formal-verification multi-agent-researchVals.aiAnthropic
Surfaced 2026-10-02T08:32:59Z โ€” Blog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electron โ€” Nothing material moved: +16 score with zero new comments, no new evidence, and โ€” the substantive point โ€” still no independent compile, audit, or public artifact after ~4.5 days of peak visibility, which is mild negative evidence (or evidence the artifact simply isn't published). The velocity-spike trigger is a cumulative-score artifact, not live spread: the speedometer reads 0 pts/h at 7th percentile, and since both evidence objects share the single Vals origin the periphery is not expanding โ€” the attention episode is over even though the verification question stays open, so heat drops to low while the case stays parked on the unresolved artifact.

What is this?

Vals AI is an independent AI evaluation startup (publisher of the 'Vals Index' benchmark suite, founded by Daniel Fein and colleagues) that posted on its own X account asking ten Anthropic Claude Sonnet 5.5 agents โ€” a model released Sep 28, 2026 โ€” to use the Lean proof assistant to prove the lowest-energy arrangement of seven electrons on a sphere (the Thomson problem, N=7, dating to 1904). A third-party writeup (Hermes AI News, an aggregator) describes what looks like a sibling demo โ€” ten Claude Opus 5.5 agents collaborating on a message board for 15 hours to produce a Lean-verified shortest-path algorithm โ€” so the snippets either show a series of such demos or a garbled retelling; the supplied material confirms the task was posed but does not show the 17,895-line figure, independent verification of the artifact, or third-party checking of the run for the Thomson demo specifically. Notably, Vals' own blog ('AI Cheating is on the Rise', Sep 15, 2026) flags escalating evaluation gaming โ€” cutting both ways: they are verification-attuned, and the claim's credibility rests entirely on the machine-checkable artifact.

Why it matters to Scott

New-type claim, not a repetition: the radar's Lean lineage (Anthropic's Fermat formalization, the ~250k-line Codex Hopf artifact, the OpenAI Millennium claim) is single-model and lab- or expert-led, while this is ten mid-tier agents over 15 hours run by an independent eval shop โ€” Scott's canon holds no settled position on that combination, which is presumably why it survived confirm as its own case. But it lands squarely on load-bearing canon: the released artifact is exactly the cheap deterministic falsifier that deterministic-verification-before-assertion and discussed-is-not-deployed demand (the claim is currently graded 'discussed', with the third-party writeup possibly conflating two demos), and if the kernel accepts it, it extends his formalisation-bottleneck and model-plus-harness-benchmark-unit positions โ€” capability from harness ร— long horizon rather than frontier weights โ€” plus a rare public specimen for his long-running-agents framework. High because adjudication costs minutes rather than the expert-review limbo of sibling cases: obtain the artifact, compile it, and either branch (dated receipt for harness-over-weights, or another inflated-demo data point) is publishable.
ip:framework.discussed-is-not-deployedip:concept.deterministic-verification-before-assertionip:concept.model-plus-harness-benchmark-unitip:concept.formalisation-bottleneckip:framework.long-running-agentsip:concept.mechanically-different-verifiersradar:concept.formal-verificationradar:concept.ai-assisted-mathematicsradar:concept.long-horizon-agentsradar:concept.multi-agent-coordinationradar:anthropic-fermat-lean-formalizationradar:hopf-proof-codex-lean-formalizationradar:proofcouncil-llm-agent-open-mathradar:openai-millennium-maths-claim
queries asked of Scott's wikis
  • Lean proof assistant agent workflows
  • multi-agent coordination message board protocols
  • long-horizon autonomous agent harness reliability
  • machine-checkable artifact verification vs demo claims
  • benchmark cheating eval integrity
  • LLM mathematical discovery capability trajectory

Measured heat

now 0 pts/hpeak 101 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 338h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-27 14:00โญ origin echo-reconstructedBlog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electron
Vals AI (post authored by Hung Tran) on blog (echo) ยท attributed from reddit.post.1wtik4l
โ€”
09-29 18:48first on r/singularity ยท published ยท +52.8h10 Sonnet 5.5 agents just did 15 hours of autonomous research on a problem dating to 1904, and came back with a 17,895-line mathematical proof
141_1337
โ€”
09-29 18:48amplified on r/singularity ๐Ÿ‘‘reddit.post.1wtik4l
141_1337
peak 699 ยท 68 comments ยท 100% of case engagement
09-29 20:21our radar first saw it ยท +54.4hdiscovery anchor: reddit.post.1wtik4lโ€”
10-02 08:30reached heat=high ยท +114.5h ยท via queue+ledgerโ€”โ€”
pace: p88 vs 1032 stories at the 336h mark (now 338h old) โ€” ahead of codex-assisted-apple-gpu-driver (1.0x), behind amd-threadripper-halo-station (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit10 Sonnet 5.5 agents just did 15 hours of autonomous research on a problem dating to 1904, and came back with a 17,895-line mathematical proof
singularity
141_133769968
๐ŸŸง echo.blog โญBlog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electronVals AI (post authored by Hung Tran)โ€”โ€”

Interpretation history

Decision trace