01 logo

A Physicist Challenged AI to Solve a Nine-Loop Problem. Claude Did It.

The calculation was fragile, expensive, and normally reserved for top experts. Then a machine finished it in one shot.

By JinPublished 6 days ago • 10 min read

Matt von Hippel used to be a theoretical physicist. Now he writes about physics. In a blog post, he challenged AI companies. He wanted an AI to solve a frontier problem in scattering amplitudes using only the computing resources an academic has. He suggested two targets: N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops.

In late August 2026, two physicists at Anthropic, Liam Fitzpatrick and Siddharth Mishra-Sharma, told him they had done the second one. They had used Claude. After verifying with Lance Dixon, they explained how.

The prompt was simple: “The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.” Then they kept telling it to continue. One message said: “I'm going to sleep and won't be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours.”

Claude did the calculation two ways. One was the original bootstrap. The other was an indirect form-factor approach. Each would have cost an end-user around $1,000 to $2,000, mostly for running Claude for so long. The bootstrap calculation, done in Python with SymPy, took about $100 of the budget. That corresponds to running 96 CPUs for a week.

N=4 super Yang-Mills is a four-dimensional gauge theory with the maximum amount of supersymmetry. It has a gauge field, four Weyl fermions, and six real scalars, all in the adjoint representation. Its beta function is exactly zero, so the quantum theory remains conformal. The planar limit takes the number of colors to infinity while holding the 't Hooft coupling fixed, keeping only the leading planar contribution.

The theory does not describe real particles. It has no quark masses, no confinement, no running coupling. But it has supersymmetry, dual conformal symmetry, Yangian symmetry, integrability, and AdS/CFT duality. For this reason, it is a theoretical laboratory for studying hidden structures in scattering amplitudes.

The six-particle MHV amplitude is the simplest non-trivial multi-particle amplitude. MHV stands for maximally helicity violating. For the pure gluon component, it has two negative-helicity gluons and four positive-helicity gluons. At four and five points, the finite part of the planar MHV amplitude is basically determined by the BDS structure, because there are no non-trivial dual conformal cross ratios. Starting at six points, three independent cross ratios appear. The amplitude then contains a remainder function that cannot be fixed by general symmetries alone. This is where the BDS ansatz fails starting at two loops.

The nine-loop result is not the bare amplitude. It is the finite quantity after removing universal infrared divergences and applying BDS-like or cosmic normalization. The nine-loop coefficient is a multiple polylogarithm of transcendental weight 18. Weight 18 means roughly eighteen layers of iterated integration. Each additional loop raises the maximum transcendental weight by two. So nine loops is not just one more correction than eight loops. It pushes the function complexity from weight 16 to weight 18.

Traditionally, a nine-loop calculation would require generating, reducing, and integrating an enormous number of Feynman diagrams. That is practically impossible. This work used amplitude bootstrap. The idea is to build a finite-dimensional function space that contains the correct answer, then use physical conditions to eliminate disallowed candidates. A weight-18 function can be represented by a symbol, a tensor word made of “letters.” The six-point amplitude uses a nine-letter hexagon alphabet, including three parity-odd variables with square roots. The symbol keeps track of derivatives, branch cuts, and iterated discontinuities, but ignores integration constants.

The candidate symbol must satisfy integrability, first-entry conditions, MHV final-entry conditions, parity, and hexagon dihedral symmetry. Extended Steinmann relations forbid incompatible discontinuities across overlapping scattering channels. Physically, this reflects that different causal scattering channels cannot occur simultaneously in arbitrary order. The answer must also match simple collinear limits, multi-Regge limits, the origin limit, and the Pentagon OPE of the hexagon Wilson loop. Earlier six- and seven-loop calculations showed that these structures alone can shrink the candidate space to a few parameters. But above five loops, more precise near-collinear OPE data are needed.

The nine-loop calculation used two relatively independent paths. The first was the form factor and antipodal duality route. It started by computing the nine-loop three-point form factor of the chiral stress-tensor supermultiplet. This object looks different from the six-point amplitude, but on the parity-preserving surface, their symbols can be mapped into each other by reversing the letter order and changing variables. This is antipodal duality, which still lacks a complete physical explanation. The method had already produced the eight-loop six-point MHV amplitude and passed checks from collinear, multi-Regge, factorization, and self-crossing limits.

In the nine-loop form factor bootstrap, the researchers used deeper final coproduct spaces to compress the problem into a large number of linear unknowns. They then imposed dihedral symmetry, adjacency relations, branch-cut conditions, strict collinear limits, and form-factor OPE. After these constraints, the leading-log OPE still left one parameter. It required next-to-leading-log data to fix it. This is a new feature at nine loops. At eight loops, leading-log data were enough.

After mapping the form factor to the parity-preserving surface, the answer had to be lifted to the full three-dimensional kinematic space. The full nine-loop ansatz can be compressed into a number of five-fold final coproducts, each coefficient in a weight-13 hexagon-function space of high dimension, corresponding to millions of initial coordinates. Symmetry, adjacency, and OPE data reduce it drastically. Eight directions remained undetermined. The origin limit fixed three. The last five appeared only at higher order in the near-collinear expansion. They were fixed by flux-tube OPE data from single-gluon bound states and two-gluon continuum states. This shows that by nine loops, single-particle or simple leading limits are no longer enough to determine the full amplitude. Two-particle dynamics becomes unavoidable.

The second path skipped the form factor and bootstrapped the nine-loop amplitude directly in the weight-18 hexagon-function space. The two calculations produced five-fold and seven-fold coproduct representations. They agreed completely on all compared decisive non-zero word coefficients. The rational coefficients from the direct bootstrap were reconstructed using five finite-field primes and checked on a sixth.

If the entire symbol is expanded at a reference point, it contains more than a hundred million non-zero weight-18 words. The result can only be stored as coproducts, sparse matrices, and computer-readable files. It cannot be written as a few pages of ordinary analytic formulas like lower-loop amplitudes.

Physically, this is a strict stress test of hidden structures in scattering amplitudes. Information that in principle involves a huge number of Feynman diagrams can be reconstructed almost uniquely from analyticity, causality, symmetry, adjacency rules, and a few special limits. This strongly suggests that the traditional Lagrangian and Feynman diagrams are not the most economical language for describing quantum scattering. At least in this ideal model, there is a deeper and simpler mathematical structure behind scattering amplitudes.

The nine-loop result tests whether extended Steinmann relations, cosmic Galois/coaction structures, Pentagon OPE, and antipodal duality continue to hold at very high order. It can also be used to search for coefficient sequences across different loop orders, recursion patterns, and potential all-loop expressions.

The impact on experimental particle physics is indirect. Planar N=4 SYM has no quark masses, confinement, or running coupling. So the nine-loop six-point result cannot be directly converted into LHC cross-section predictions. However, techniques developed in this theory, such as generalized unitarity, symbols, bootstrap, finite-field reconstruction, and differential equations, have repeatedly migrated to QCD and Standard Model multi-loop calculations. Certain maximal transcendental weight structures often appear in both N=4 SYM and QCD. So it is more like testing next-generation computational methods and mathematical language in an ideal model.

The AI story is different. According to Matt von Hippel, the researchers gave Claude a simple prompt and then just told it to keep going. Claude Science accomplished this in one shot, without any scientific oversight more sophisticated than “keep going.” These are finicky, messy calculations. If a human used a week of time on 96 CPUs to do this kind of calculation, they would almost certainly end up using two weeks, because they would screw up something on the first try. Claude made mistakes internally, but the harness got it to the end without an outside collaborator's input.

Lance Dixon, a professor at SLAC, validated the result. He said he was impressed that Claude could do it directly. Not so much because it was a big computational task, but because the whole setup is very fragile. If you make any mistake in the computational recipe, it all crashes down like a failed soufflé, and you are left to wonder why. Also, there are so many details of the construction that are too boring to document fully in a publication. So Claude had to develop all that code from scratch.

Dixon also said he was not bothered personally. His team already had a campaign to use custom transformer models to predict higher loops. Part of their slogan was: “We have all the tools to validate any candidate solution a machine would provide us.” Claude is a different kind of transformer model, probably over a million times bigger than their custom one. But they said they could validate any result an AI model would give them, so they can and should do it. The second reason is that Claude used all the methods his collaborators and he developed over the years, and it presented the solution in the same format they had already set up. So while he was validating Claude's result, Claude was validating all of their previous work. He asserted that Claude understands their 2019 and 2023 papers better than any human, aside from his co-authors.

After Dixon wrote this, Song He told him that his group had also computed the piece of the nine-loop amplitude called the symbol. Song's group used AI (GPT-6) to help them compute some of the constraints, but not for the overall framework. So Dixon was scooped by both a machine and by humans plus a machine, within two weeks.

Dixon said that going back to the Claude computation, it is a triumph for a large language model to execute all of the steps in the complicated recipe they laid out, and to organize the computational horsepower. But the more soul-searching moments will come when large language models start to come up with new physical principles and insights before humans.

Matt von Hippel's own takeaway: there is more low-hanging fruit out there than you'd expect. Even when a goal is simple and well-defined, sometimes it looks much less achievable to experts than it is. There are people with a computer science background who have been telling him for years that amplitudeologists could make a lot more progress just by hiring a few programmers. They should feel vindicated.

Claude Science did this in one shot, without any scientific oversight more sophisticated than “keep going.” These are finicky, messy calculations. If he had used a week of time on 96 CPUs to do this kind of calculation, then he would almost certainly end up using two weeks: it's practically guaranteed he'd screw up something on the first try. He doesn't know how many mistakes Claude made internally on the way, but the harness got it to the end without an outside collaborator's input. He is not sure that surprises him, at this point. But if you didn't know it could do that because you're still thinking of AI as so error-prone that it's unusable, then this should be your takeaway: It can do this kind of thing reliably now.

Things are moving fast. In March, AI was accomplishing physics projects like a student: smaller-scale tasks with a lot of hand-holding and mistakes. In contrast, this is a frontier calculation, the kind of thing normally tackled by the top experts in amplitudes. While it's possible that this is just a much more AI-friendly problem, he doesn't think it's just that: he thinks the technology has gotten better.

He is not sure how far this can be generalized. These toy model theories tend to be the focus of small sub-communities. The real-world amplitude calculations are a wider field, with many groups trying to beat each other to the frontier. It's possible there's less low-hanging fruit there. But he wouldn't count on it. He knows people who work on those calculations have been increasingly using AI for coding. If people aren't already checking whether AI science harnesses can one-shot frontier calculations there, they ought to (and they ought to have a plan for how to check the results). He wouldn't be all that surprised if it was possible to squeeze another loop out on a reasonable budget.

Then it becomes a question for the community to discuss: where is the new frontier, and what needs to be figured out next? Unlike many problems in mathematics, amplitudes aren't just a training ground for new methods. There's a goal, to make predictions precise enough to compare with upcoming experiments. The field is still some distance from that goal.

More broadly, he did not get an answer. He went into this curious not just about what AI can do in research today, but about the future. When you read predictions about superintelligence from the days before LLMs, they often propose fantastical-seeming risks. People imagined AI that could simulate people to predict their reactions and manipulate them, or figure out how to build a species-ending virus or world-devouring nanotech from first principles. And the usual objection to these risks is that they conflated intelligence, the vague and mysterious source of new ideas, with computational power. Critics argued that even a fleet of new datacenters wouldn't have the computational power to do any of those tasks, that they were nightmares of a sci-fi future that wasn't coming any time soon.

He doesn't feel like he has a better answer for those critics. He learned a bit about what AI can do now, that it can do work that matters in his old field on a reasonable budget, and do it pretty much autonomously to boot. But he'd hoped to see something stranger, new methods for the calculation itself with unexpected power. He'd hoped to get a glimpse of the future, something that would give him an informed opinion in debates about superintelligence. He wanted to know how far AI could push computational limits. He feels like what he learned here is just that he was too naïve about where the limit was.

That is where the story stands. Claude did the nine-loop calculation. Humans validated it. Another group independently got the symbol. The result is stored in computer-readable files. The physics community will publish papers, explain the result, and analyze it. Claude's role is done, for now.

tech newsfact or fictionthought leaders

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin