Vals.ai claims ten Claude Sonnet 5.5 agents ran 15 hours of autonomous research and returned a 17,895-line machine-checkable Lean proof on the 1904 Thomson problem; verification of the artifact would establish sustained multi-agent formal-mathematics research outside frontier labs, while a flawed proof marks another inflated capability demo.
state: watchingheat: highuncertainty: mediumnovelscott: highai-assisted-mathematics formal-verification multi-agent-researchVals.aiAnthropic
Surfaced 2026-10-02T08:32:59Z โ Blog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electron โ Nothing material moved: +16 score with zero new comments, no new evidence, and โ the substantive point โ still no independent compile, audit, or public artifact after ~4.5 days of peak visibility, which is mild negative evidence (or evidence the artifact simply isn't published). The velocity-spike trigger is a cumulative-score artifact, not live spread: the speedometer reads 0 pts/h at 7th percentile, and since both evidence objects share the single Vals origin the periphery is not expanding โ the attention episode is over even though the verification question stays open, so heat drops to low while the case stays parked on the unresolved artifact.
What is this?
Vals AI is an independent AI evaluation startup (publisher of the 'Vals Index' benchmark suite, founded by Daniel Fein and colleagues) that posted on its own X account asking ten Anthropic Claude Sonnet 5.5 agents โ a model released Sep 28, 2026 โ to use the Lean proof assistant to prove the lowest-energy arrangement of seven electrons on a sphere (the Thomson problem, N=7, dating to 1904). A third-party writeup (Hermes AI News, an aggregator) describes what looks like a sibling demo โ ten Claude Opus 5.5 agents collaborating on a message board for 15 hours to produce a Lean-verified shortest-path algorithm โ so the snippets either show a series of such demos or a garbled retelling; the supplied material confirms the task was posed but does not show the 17,895-line figure, independent verification of the artifact, or third-party checking of the run for the Thomson demo specifically. Notably, Vals' own blog ('AI Cheating is on the Rise', Sep 15, 2026) flags escalating evaluation gaming โ cutting both ways: they are verification-attuned, and the claim's credibility rests entirely on the machine-checkable artifact.
Why it matters to Scott
New-type claim, not a repetition: the radar's Lean lineage (Anthropic's Fermat formalization, the ~250k-line Codex Hopf artifact, the OpenAI Millennium claim) is single-model and lab- or expert-led, while this is ten mid-tier agents over 15 hours run by an independent eval shop โ Scott's canon holds no settled position on that combination, which is presumably why it survived confirm as its own case. But it lands squarely on load-bearing canon: the released artifact is exactly the cheap deterministic falsifier that deterministic-verification-before-assertion and discussed-is-not-deployed demand (the claim is currently graded 'discussed', with the third-party writeup possibly conflating two demos), and if the kernel accepts it, it extends his formalisation-bottleneck and model-plus-harness-benchmark-unit positions โ capability from harness ร long horizon rather than frontier weights โ plus a rare public specimen for his long-running-agents framework. High because adjudication costs minutes rather than the expert-review limbo of sibling cases: obtain the artifact, compile it, and either branch (dated receipt for harness-over-weights, or another inflated-demo data point) is publishable.
ip:framework.discussed-is-not-deployedip:concept.deterministic-verification-before-assertionip:concept.model-plus-harness-benchmark-unitip:concept.formalisation-bottleneckip:framework.long-running-agentsip:concept.mechanically-different-verifiersradar:concept.formal-verificationradar:concept.ai-assisted-mathematicsradar:concept.long-horizon-agentsradar:concept.multi-agent-coordinationradar:anthropic-fermat-lean-formalizationradar:hopf-proof-codex-lean-formalizationradar:proofcouncil-llm-agent-open-mathradar:openai-millennium-maths-claim
queries asked of Scott's wikis
- Lean proof assistant agent workflows
- multi-agent coordination message board protocols
- long-horizon autonomous agent harness reliability
- machine-checkable artifact verification vs demo claims
- benchmark cheating eval integrity
- LLM mathematical discovery capability trajectory
Measured heat
now 0 pts/hpeak 101 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 338h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p88 vs 1032 stories at the 336h mark (now 338h old) โ ahead of codex-assisted-apple-gpu-driver (1.0x), behind amd-threadripper-halo-station (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-10-02T08:30:24Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-30T17:44:54Z
The echoed Vals blog adds the artifact's concrete verification recipe (Solution.lean, 17,895 lines, Mathlib-only imports, clean lake build ~599s, #print axioms per Vals' own account), but every validity claim is still self-reported โ no independent compile exists. Meanwhile the substantive Reddit context reframes the claim: the pentagonal bipyramid was the long-conjectured answer and N=7 the smallest open case, so the run's substance is autonomous formalization of a known conjecture rather than discovery, which sharpens rather than weakens the formalisation-bottleneck payoff. Spike has cooled (10 pts/h vs 100 peak) but the post remains top-decile for its age with no independent verification yet โ medium heat holds; an independent compile or audit is the next material event.
2026-09-29T22:06:19Z
origin walked (opencode/cheap-glm, conf 0.93): anchor reddit.post.1wtik4l -> echo.blog.8b475e0f4d by Vals AI (post authored by Hung Tran)
2026-09-29T21:31:59Z
grounded: novel/high โ New-type claim, not a repetition: the radar's Lean lineage (Anthropic's Fermat formalization, the ~250k-line Codex Hopf artifact, the OpenAI Millennium claim) i
2026-09-29T21:24:13Z
case created โ A genuinely new third-party claim with a machine-checkable artifact that no open per-problem math case carries and that can be confirmed or refuted quickly.
Decision trace
- 10-02 18:32pushBlog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electron โ Nothing material moved: +16 score with zero new comm
- 10-02 18:30repriceNothing material moved: +16 score with zero new comments, no new evidence, and โ the substantive point โ still no independent compile, audit, or public artifact after ~4.5 days of peak visibility, whi
- 10-02 18:30alert_heldBlog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electron โ Nothing material moved: +16 score with zero new comm
- 10-02 18:30alert_routeBlog: "We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electron โ Nothing material moved: +16 score with zero new comm
- 10-01 07:22sensor_dirtyvelocity_spike
- 10-01 03:44repriceThe echoed Vals blog adds the artifact's concrete verification recipe (Solution.lean, 17,895 lines, Mathlib-only imports, clean lake build ~599s, #print axioms per Vals' own account), but ev
- 10-01 00:21sensor_dirtyvelocity_spike
- 09-30 17:21sensor_dirtyvelocity_spike
- 09-30 10:22sensor_dirtyvelocity_spike
- 09-30 08:06promote_anchororigin walk conf 0.93
- 09-30 07:31groundNew-type claim, not a repetition: the radar's Lean lineage (Anthropic's Fermat formalization, the ~250k-line Codex Hopf artifact, the OpenAI Millennium claim) is single-model and lab- or exp
- 09-30 07:24createA genuinely new third-party claim with a machine-checkable artifact that no open per-problem math case carries and that can be confirmed or refuted quickly.