2026-10-11 16:38 UTC

Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma claim Fable 5.1, driven through the Claude Science harness with only periodic 'keep going' prompts, computed the nine-loop six-particle (hexagon) amplitude in planar N=4 super Yang-Mills — an outstanding problem in scattering amplitudes — for roughly $1–2k of near-unattended inference, with the result verified by expert Lance; acceptance of the amplitude into the field and replication of the approach would establish frontier agent harnesses as demonstrated solvers of computational-physics barriers experts considered out of reach on academic budgets.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highai-research agent-harnesses ai-for-scienceMatt von HippelLiam FitzpatrickSiddharth Mishra-SharmaSong HeLanceAnthropic
Surfaced 2026-09-27T01:00:04Z — Von Hippel issued an August challenge for AI to solve N=4 SYM to nine loops on academic resources; Anthropic's Fitzpatrick and Mishra-Sharma — Fourth velocity_spike from the same aging Reddit thread; last look's 'receding' read was slightly off — the score resumed slow accrual (548→606 over ~9h) — but with zero comment velocity (0/h, 47 total, 0.98 ratio) it is a vote tail, not renewed discussion, and the periphery is flat for a fourth consecutive look (HN 103/60, resub 2/0; no outlets, publications, replications or derivatives). The spike channel for this object is a confirmed cohort artifact to be discounted as noise; case meaning unchanged: corroborated within the episode, quiescent, awaiting consolidation events that will arrive as new evidence, not attention.

What is this?

Anthropic published a guest post (Sep 25, 2026) by physicist-blog­ger Matt von Hippel reporting that two Anthropic physicists, Liam Fitzpatrick and Siddharth Mishra-Sharma, drove Claude Fable 5.1 inside the company's Claude Science harness to compute the six-particle (hexagon) scattering amplitude in planar N=4 super Yang-Mills at nine loops — one loop past the 2023 eight-loop record of SLAC's Lance Dixon — answering von Hippel's August 2026 public challenge to clear such a frontier calculation on an academic-scale budget. The run took a one-sentence problem prompt plus periodic 'keep going' updates every 4–6 hours while the humans slept, cost roughly $1–2k of inference (about $100 for the core bootstrap step on 96 CPUs over a week), and used the community's established bootstrap and form-factor recipes rather than new mathematics; Dixon spent two weeks validating the result and a Chinese Academy of Sciences group (Song He) independently reached most of the same result in parallel using GPT-6. Every substantive line still traces to the single first-party Anthropic post — von Hippel was compensated by Anthropic, Dixon's validation addendum came with Claude usage credits, and there is no independent publication or replication yet — and von Hippel's own takeaway deflates the 'impossible barrier' framing to known methods, more compute, and better engineering, leaving the near-unattended keep-going harness mode as the credibly novel demonstrated part.

Why it matters to Scott

Anthropic has now demonstrated at frontier scale the exact near-unattended 'keep going' heartbeat mode Scott's Long-Running Agents / Heartbeat Supervisory Program canon predates — von Hippel's own takeaway names that harness mode as the credibly novel part — making this the dated-receipts publishing opportunity for the framework plus a flagship $1–2k Cost-of-Cognition datapoint. It is simultaneously a live Verification-Paradox instance (a result too hard to self-check, validated by a credit-compensated Dixon — a correlated-checkers smell) with the 'just tell it to keep going' harness-vs-model question landing squarely on his Model-Plus-Harness benchmark unit, and the Schwartz follow-up shows Claude Science is a sustained program stream, not a one-off.
ip:framework.heartbeat-supervisory-programip:framework.long-running-agentsip:concept.verification-paradoxip:concept.cost-of-cognitionip:concept.model-plus-harness-benchmark-unitip:concept.correlated-checkers-pitfallradar:concept.long-running-agentsradar:concept.agent-harnessesradar:concept.ai-for-scienceradar:concept.verificationradar:concept.inference-economicsradar:agmai-openai-math-release-adviceradar:fable-astra-proof-replication
queries asked of Scott's wikis
  • long-running agent heartbeat keep-going supervision loop
  • cost of cognition inference cost per completed frontier task
  • verification paradox expert validation of AI scientific output
  • harness versus model attribution scaffolding capability claims
  • autonomous research agents AI-for-science public challenge as benchmark
  • agent-built tooling pipelines symbolic computation SymPy workflow

Measured heat

now 0 pts/hpeak 458 pts/hcomments 0/hpeers p18momentum: steady3 platformsage 410h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-24 14:00⭐ origin echo-reconstructedVon Hippel issued an August challenge for AI to solve N=4 SYM to nine loops on academic resources; Anthropic's Fitzpatrick and Mishra-Sharma
Matt von Hippel (guest post on Anthropic's research page) on blog (echo) · attributed from hn.story.49848033
—
09-25 18:11first on hacker news · published · +28.2hYes, Claude can do Nine Loops
tzury
—
09-25 22:37first on r/singularity · published · +32.6hClaude Fable 5.1 completed a frontier nine-loop particle-physics calculation experts had worked toward for years, with researchers mostly just telling it “keep going” while it built, debugged and ran the entire workflow
141_1337
—
09-25 18:11amplified on hacker newshn.story.49848033
tzury
peak 104 · 62 comments · 8% of case engagement
09-25 22:37amplified on r/singularityreddit.post.1wqa2u2
141_1337
peak 658 · 48 comments · 20% of case engagement
09-25 23:46amplified on hacker newshn.story.49851615
Hbruz0
peak 2 · 0 comments · 0% of case engagement
09-30 14:54amplified on r/singularityreddit.post.1wu72qu
drhenriquesoares
peak 771 · 181 comments · 27% of case engagement
10-01 18:09amplified on r/singularityreddit.post.1wv6q8k
badumtsssst
peak 14 · 0 comments · 0% of case engagement
10-01 22:51amplified on hacker newshn.story.49927980
acossta
peak 3 · 0 comments · 0% of case engagement
4 more amplifiers in ainews.case_chain
09-25 18:21our radar first saw it · +28.4hdiscovery anchor: hn.story.49848033—
09-27 00:58reached heat=high · +59.0h · via ledger——
pace: p96 vs 1032 stories at the 336h mark (now 410h old) — ahead of meta-muse-spark-13-release (1.0x), behind openai-rl-pause-sandbox-escape (1.0x)

Evidence (11) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnYes, Claude can do Nine Loops
Retrieved article excerpt

Open article · Retrieved 2026-09-25T18:28:14.826785+00:00

Science

# Yes, Claude can do Nine Loops

Sep 25, 2026

*In this guest post, physicist and science writer Matt von Hippel shares what happened when he issued a challenge to AI companies regarding a problem in his former subfield of theoretical physics.*

Illustration of nine-loops

It’s not often that you issue a challenge, only to see it beaten a month later. But we’re living in unusual times.

Let me introduce myself: I’m Matt von Hippel. I used to be a theoretical physicist; these days I’m a science writer. Throughout, I’ve been a blogger, writing weekly at [4gravitons.com](https://4gravitons.com/) about physics and the people who do it.

More and more, blogging about physics has meant blogging about AI. That’s a problem, because I’m definitely not an AI expert. I’ve [dabbled](https://arxiv.org/abs/2502.05121) in it, sure. I probably know more than your grandma. But I mostly have to step back and trust the experts. And frustratingly, the experts disagree! I’ve heard from smart, well-informed people who are confident that AI is a few years away from superintelligence, and that superintelligence will be capable of truly terrifying things. And I’ve heard from smart, well-informed people who are equally confident that LLM-based AI is close to a ceiling, that models like Claude won’t even be able to do impressive work in physics, let alone conquer the world.

I’ve been reluctant to make my own predictions. Before forming an opinion, I wanted to see an LLM make progress on something familiar, something I knew was hard to do because I’d tried to do something similar myself.

In addition to that, I wanted to see an LLM do something that I expected to be *computationally* hard. LLMs have made impressive strides in math, certainly, and [this month alone](https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/) has likely changed many peoples’ minds. But progress in math comes from new ideas, and ideas are mysterious things: one never quite knows how hard they are to find until they’re found. Computation felt more solid. I wanted to see an LLM tackle a challenge that seemed out of reach not because researchers didn’t know how to do it in principle, but because doing it seemed like the kind of thing that would take more computers and time than the researchers reasonably had access to. I wanted to see if those researchers were wrong: if a smarter, artificial researcher could use the same computers, and solve the problem anyway.

So, I issued [a challenge](https://4gravitons.com/2026/08/07/it-only-counts-when-ai-gets-to-my-field/):

> “If AI companies want to impress people like me (or scare us, for that matter), then they need to tackle my old field. Show that an AI can take the kinds of computer resources an academic has access to, and solve one of the scattering amplitudes field’s big outstanding problems. Show that a computational limit everyone expected to be a problem doesn’t actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops.”

In short: can AI solve a frontier problem in my former subfield of theoretical particle physics? And can it do it on a budget?

## **The challenge**

My old field is a branch of theoretical particle physics called amplitudeology. When other particle physicists predict new particles, they make sure they can do the calculations to test those predictions. They compute formulas called scattering amplitudes, which let physicists use the momenta and energies of subatomic particles to calculate how likely they are to react in particular ways. If physicists can make more accurate predictions for these reactions, they can check whether results from experiments like the Large Hadron Collider match those predictions. A mismatch could be evidence for a new theory, one that could explain some of physics’ big lingering mysteries, like the nature of dark matter, or the balance between matter and antimatter in the universe.

These scattering amplitude formulas are hard to compute, so hard that physicists almost always use approximations. They do partial calculations, cut off at a specific number of “loops,” a measure of how complicated interactions between particles are allowed to get. The more “loops” they include in their calculations, the closer they get to the real answer, and the harder, computationally, the calculation is to do.

In practice, most scattering amplitude formulas have only been calculated to two loops. A few have three. [The most precise prediction in particle physics you might have heard of used five](https://en.wikipedia.org/wiki/Anomalous_magnetic_dipole_moment#Electron).  
  
Amplitudeologists want to do better. They develop experimental new techniques, and test them on special “toy model” theories. By trying the technique with a toy model where the calculation is easier, rather than the more challenging particles of the real world, amplitudeologists can stress-test the new methods and see how far they can go.

I posted challenges for two of those toy models. The one the folks at Anthropic chose to tackle was to go up to nine loops with a particular toy model theory, called N=4 super Yang-Mills.

“Yang-Mills” is a technical name for a type of theory that explains most of the world around us. Three of the four fundamental forces of nature: electromagnetism, the strong nuclear force that holds the nuclei of atoms together, and the weak nuclear force that causes radioactive decay in things like bananas, are all Yang-Mills theories.

The “N=4 super” comes from supersymmetry. Physicists have speculated that each particle has a “supersymmetric partner,” a particle with the same charge, but of a different type, matching matter particles like electrons to force particles like photons. At one time they were optimistic these particles could explain dark matter, via undiscovered partners of more familiar particles. Those speculations used “N=1” supersymmetry. In “N=4,” each particle has *four* supersymmetric partners, not just one.

That surfeit of particles makes the theory very unrealistic. N=4 super Yang-Mills isn’t used as an explanation for dark matter, [or for anything in the real world](https://arstechnica.com/science/2013/05/earning-a-phd-by-studying-a-theory-that-we-know-is-wrong/). Instead, amplitudeologists use it to hone their techniques, because N=4 is paradoxically easier to calculate with. The delicate balance between the different particles means only certain combinations of variables are needed, streamlining calculations.

These calculations were done with an experimental technique called a bootstrap, which ended up bizarrely well-suited for use of AI. To bootstrap an amplitude, you don’t have to take into account every possible particle interaction. You just need to know roughly what the answer ought to look like, keeping track of every possibility in computer files in a specialized alphabet. Then you start checking everything you know: predictions from other calculation techniques, rules the answer has to obey, links to related problems where the answer was easier to find. It’s a bit like Sudoku, where you begin with a grid with all possible numbers, then cross them out as you go. In the end, you’re hoping to find that only one possibility satisfies all the checks, while having enough checks left over to make sure you didn’t make a mistake.

That meant that Lance was already well set up to check if someone had handed him the next amplitude formula, with nine loops. It would be an interesting answer, not just as a validation of the bootstrap technique, but as a rare example of an amplitude with that many loops of complexity, an answer that could be worth studying in its own right.

But he hadn’t computed it, and neither had anyone else in the field. The way he found the eight-loop answer was already a bit indirect, via a [surprising link](https://www.quantamagazine.org/particle-physicists-puzzle-over-a-new-duality-20220801/) to a different but related formula called a form-factor, a kind of partial amplitude involving different particles that turns out to be a bit easier to calculate. He was expecting to find the next loop even more indirectly, potentially by a different kind of AI method. If people thought it was possible to just run the usual bootstrap method for one more loop, someone would have done it.

## **Then people did it**

Apparently, there are folks at Anthropic who read my blog.

At the end of August, Liam Fitzpatrick and Siddharth Mishra-Sharma, two physicists at Anthropic, reached out to me to say they had tackled one of the challenges in my post. After verifying the result with Lance, they talked me through how they got it.

True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within [Claude Science](https://claude.com/product/claude-science), a platform scientists can pay to use. Claude Science is what folks in the biz call a “harness,” a program that uses the Claude LLM with structured rules and prompts in order to get more robust and scientifically useful behavior.

Apparently, after asking Claude which problem it was most likely to be able to tackle, they gave it a simple prompt:

“The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.”

From there, they just kept telling it to keep going, with comments like:

“I'm going to sleep and won't be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours.”

Claude ended up doing the calculation two different ways: the original bootstrap, and the indirect form-factor approach. Either approach would have cost an end-user around one or two thousand dollars, mostly due to the expense of running Claude for so long. The bootstrap calculation, done with the Python programming language with package SymPy, took around $100 of the budget, corresponding to running 96 CPUs for a week.

Running 96 CPUs for a week might have felt like a lot when I was doing this kind of work ten years ago, but it’s pretty affordable now if you have a good reason.

As it turned out, the result wasn’t all that far away for humans either. A few days after I heard from Anthropic, we heard from Song He, an amplitudeologist at the Chinese Academy of Sciences in Beijing. Song’s group had already gotten the majority of the result. They’d used some AI assistance, based on GPT-6, but not the kind of one-shot almost human-less approach Anthropic used.

Everyone has been friendly here, which is a bit of a relief. The humans, Lance and Song and their collaborators, will get to publish the results, taking time to explain them and analyze them for the benefit of future researchers. Claude’s role is done, for now.

## **So, problem solved?**

I set my challenge because I wanted a better sense of what current AI can do, and where it could go from here. So what have I learned?

I’d thought this could be a chance to see AI overcome a computational barrier in a surprising way. Instead, it did something it turned out humans were also able to do. Claude used known methods, with a bit more compute than people had tried to use before. It may have gotten a boost from using Python, and not Maple (Lance’s favorite program for math) or Mathematica (mine), and it may have used much better software engineering practices than we would have, but not super-intelligently so.

My biggest takeaway is that there is more low-hanging fruit out there than you’d expect. Even when a goal is simple and well-defined, sometimes it’s going to look much less achievable to experts than it actually is. There are people with a computer science background who’ve been telling me for years that amplitudeologists could make a lot more progress just by hiring a few programmers. They should feel vindicated.

It’s also noteworthy that Claude Science accomplished this in one shot, without any scientifi
tzury10462
🟧 echo.blog ⭐Von Hippel issued an August challenge for AI to solve N=4 SYM to nine loops on academic resources; Anthropic's Fitzpatrick and Mishra-SharmaMatt von Hippel (guest post on Anthropic's research page)——
🟠 redditClaude Fable 5.1 completed a frontier nine-loop particle-physics calculation experts had worked toward for years, with researchers mostly just telling it “keep going” while it built, debugged and ran the entire workflow
singularity
141_133765548
🟧 hnIt Got to My FieldHbruz020
🟠 redditClaude solves an important mathematical problem.
singularity
drhenriquesoares769181
🟠 redditWorking on Claude-shaped problems with BootLoops, a toolkit for exact calculations in quantitative science.
singularity
badumtsssst140
🟧 hnClaude-Shaped Scienceacossta30
🟠 redditA Harvard physicist spent 3 months doing research with Claude Fable 5: it reproduced weeks of work in 20 minutes, completed 15 never-before-solved physics calculations, and contributed to 36 papers across 18 fields
singularity
141_13371222192
🟧 hnClaude-Shaped Sciencefamouswaffles2912
🟠 redditUsing Claude Science to produce the first complete map of the sky in UV light
singularity
ResultBackground2450764
🟧 hnThe Missing Map of the Skyjmague22

Interpretation history

Decision trace