One Sentence, Nine Loops, Two Thousand Dollars: How Claude Passed the Human Record in Physics
Anthropic’s model ran for days with almost no supervision. The result was verified. The method was old. The execution was not.

If you read only the task, it sounds ordinary:
“Calculate the six-particle nine-loop amplitude in planar N=4 supersymmetric Yang-Mills theory.”
It is not ordinary. No human team had calculated that result before. This time, the thing that did it was not a research group, a years-long PhD project, or a supercomputing center. It was Anthropic’s Claude, running for several days with almost no human supervision, at a cost of about one to two thousand dollars.
Two different calculation routes agreed completely. The run also reproduced the published eight-loop result. Lance Dixon, who held the eight-loop record, independently verified the result. Around the same time, Song He’s team at the Institute of Theoretical Physics, Chinese Academy of Sciences, independently obtained the symbol for the nine-loop amplitude with GPT-6 and released datasets for the two- through nine-loop amplitudes.
The event drew attention in theoretical physics and AI circles. The reason is simple: it shows AI taking on the longest, most tedious, and most failure-prone part of frontier scientific computation.
What nine loops means
Scattering amplitudes are the mathematical objects particle physicists use to predict what happens when particles collide. They determine which particles can be produced, with what probability, and with what distribution. More precise calculations test the Large Hadron Collider and other experiments more sharply. They also make it easier to spot signs of physics beyond the Standard Model.
Complexity explodes with each loop.
In Feynman diagrams, one loop means more virtual particles, more momentum integrals, more divergent terms, and more intermediate expressions. One or two loops can be done by hand. Three or four require specialized algorithms and computer algebra systems. Beyond five loops, the number of diagrams becomes unrealistic and the intermediate expressions grow past what ordinary computer memory can hold. At nine loops, enumerating Feynman diagrams is an astronomical exercise.
High-order amplitude calculation stopped being “write the formula, substitute, simplify” a long time ago. It relies on specialized methods. One of the most important is bootstrap.
Bootstrap does not start from Feynman diagrams. It starts by guessing what class of functions the final result belongs to. Then it applies physical constraints to force out a unique answer. Those constraints include collinear limits, soft limits, Steinmann relations, symmetry, and known lower-order results. Each constraint removes unknowns. The process resembles a large Sudoku: you do not calculate each number directly. You eliminate possibilities until only one answer remains.
Human physicists spent more than a decade refining this method. It made previously impossible high-order calculations possible. It is still slow. Function spaces can be huge. Constraints can number in the thousands. In between sit repeated trial and error, organization, and verification. A team often needs years to advance one loop.
In 2023, Lance Dixon’s team used antipodal duality to calculate the eight-loop amplitude. That was the human record. Dixon’s team had prepared for years. Nine loops was assumed to be a long-term target.
Then Claude arrived.
The task was one sentence
Two Anthropic physicists, Liam Fitzpatrick and Siddharth Mishra-Sharma, first asked Claude which theoretical physics problem it was most confident it could solve.
Claude gave its answer. Then they gave it one prompt:
“The problem is to calculate the six-particle (hexagon) nine-loop amplitude in planar N=4 supersymmetric Yang-Mills theory.”
After that, they sent mostly “continue” instructions.
One instruction reportedly said:
“I’m going to sleep and will be away for the next few hours. Keep calculating until I stop you, and report every 4 to 6 hours.”
That is what “almost no human supervision” means here. Humans did not guide each step. They did not choose the method or decide what to do next. They gave a goal. Claude organized the calculation, generated code, scheduled intermediate results, and worked through constraints until the linear equations could be solved.
Claude ran for several days.
It produced the six-gluon nine-loop scattering amplitude by two independent routes. The routes agreed. It also reproduced the published eight-loop result. The pure computing cost was on the order of a hundred dollars. The main cost came from long calls to Claude. One complete route cost about one thousand to two thousand dollars.
That cost matters.
One to two thousand dollars is not luxurious for frontier theoretical physics. The expensive part has always been human time. A graduate student doing similar work might spend weeks or months on one attempt. When a single trial costs a few cents, the economics of the calculation change.
Two routes, one answer
Claude did not have a flash of inspiration. It used bootstrap, a method humans developed.
In N=4 super Yang-Mills theory, amplitudes can be described with symbols and an alphabet. The final expression is like a word made of letters. Physical constraints describe how that word behaves in certain limits. Over the past decade, human physicists guessed the alphabets, checked their structure, built function spaces, and developed a constraint-solving process.
Claude automated that process through nine loops.
It generated code, organized large intermediate results, handled high-dimensional linear systems, and tried paths until constraints closed. It had to judge which paths worked and which did not, then adjust when it hit dead ends. None of this looks smart in isolation. Together it is tedious and unforgiving. One symbolic error or one missed boundary condition can ruin the whole calculation.
Claude finished with two different routes. The routes differed in method organization, intermediate steps, and code. The results agreed. It also reproduced the published eight-loop result. That is an internal cross-check. If it could not get the known eight-loop result right, the nine-loop result would mean nothing.
That is also why the field recognized the result quickly. Bootstrap calculations have a strict verification culture. Results must satisfy all known constraints. They must reduce to known lower-order results in certain limits. They must be reproducible by independent methods. Claude’s result passed.
The eight-loop record holder checked it
Lance Dixon independently checked the result with his own verification tools.
Dixon had thought the nine-loop amplitude would be very hard to calculate directly. His team had prepared for several years. Claude solved it head-on.
Dixon’s assessment was candid. For a large language model to execute every step of this complex process and organize the computing power was, in his words, “a quite remarkable victory.” He also said that, apart from his own collaborators, Claude understood his team’s 2019 and 2023 papers better than anyone he knew.
That sentence may matter more than the nine-loop result itself.
It means Claude was not just running code. It understood the literature, the method, and the constraint system well enough to organize them into a working calculation. It did not propose new physics or invent a new method. It pushed an existing human method to a precision humans had not reached.
A person can solve a new problem with old tools. The tools are old. The answer is new.
Another team arrived at the same time
Around the time Claude finished, Song He’s team at the Institute of Theoretical Physics, Chinese Academy of Sciences, advanced to nine loops with GPT-6.
The team used GPT-6 for some constraints. The overall framework was still built by human researchers. They independently obtained the symbol for the nine-loop amplitude and released datasets for the two- through nine-loop amplitudes.
That matters.
If only Anthropic had announced Claude’s result, outsiders might wonder whether the demonstration was staged, whether humans intervened heavily, or whether it was a one-off. Song He’s team used a different AI model and a different organizational approach. They arrived at the same position. The result is not isolated.
It looks like an early signal: large models are becoming usable executors in high-order theoretical calculations.
The public datasets also have long-term value. They are results, verification material, training material, and a starting point for follow-up work. If someone wants to push to ten or eleven loops with AI, those data are infrastructure.
The automation is the story
It is easy to misread the event as “AI solved a frontier problem in theoretical physics.”
A more accurate description: with almost no human supervision, AI completed an extremely complex calculation in frontier theoretical physics. Human teams had reached eight loops after years of work. AI reached nine.
It did not propose new physical principles. It did not discover new symmetries. It did not invent new mathematical structures. It used mature human methods. It executed.
Execution is one of the largest costs in research.
Modern frontier theoretical physics is not one person, one sheet of paper, one pen. It requires code, debugging, data management, intermediate expressions, verification, and organization. This work consumes PhD students and postdocs. The failure rate is high. Many attempts produce nothing but the knowledge that a road does not work.
When AI takes on that work, research organization changes.
Human researchers can focus on asking questions, designing constraints, judging direction, and verifying meaning. AI can handle the long, tedious, failure-prone execution. This is not replacement. It is a reorganization of labor.
That also explains why the shock inside theoretical physics may be larger than the public reaction. To outsiders, “six-gluon nine-loop scattering amplitude” is gibberish. To insiders, it means a calculation that once required years of team effort can now be completed in one to two weeks, at a cost of one to two thousand dollars, by AI running for several days.
The contrast is sharp.
No new physics, no new method
Some physicists have warned against overreading the result.
Theoretical physicist von Hippel noted that Claude used ready-made methods. The computing power was not extravagant. The new part was that a human was no longer doing the work. Claude did not invent new letters or discover new symmetry structures. It extended a road on a map humans had already drawn.
The nine-loop result itself produced no new physical conclusions. N=4 super Yang-Mills theory is an idealized toy model, not the real world. Its value is in testing methods and tools, not in describing nature directly.
Scientific progress has two modes.
One pushes into the unknown: new particles, new symmetries, new principles. The other pushes known methods farther and finds where they break. The second mode also produces data. It shows the limit of this automated process and where human methods remain irreplaceable.
Von Hippel also said he had expected AI to break calculation bottlenecks with methods humans had not thought of. What happened was different. That difference is still important. It shows AI can reliably execute one of humanity’s most complex calculation processes.
The next question is not whether it can calculate one loop higher. The question is whether it can propose a new method itself.
That is the dividing line.
What happens next
In the short term, researchers will continue to sort out and verify the nine-loop result. Dixon’s team has finished its verification. Broader independent reproduction, cross-checking with other methods, and dataset organization will continue.
In the medium term, ten and eleven loops may become targets. If nine loops takes one to two weeks and one to two thousand dollars, ten loops will cost more, but it may be affordable. AI also changes the number of attempts. Human teams try the routes they are most confident in because failure is expensive. AI can try several routes at once because a single failure is cheap.
In the long term, this may change how theoretical physics calculation works.
Graduate students will still need to learn bootstrap. They may need to learn earlier how to collaborate with AI, how to design constraints, how to verify results, and how to judge whether a calculation is worth doing. Tedious code and intermediate result organization may increasingly go to AI. Human roles will move toward questioner, designer, verifier, and judge of meaning.
Authorship, credit, and reproducibility standards will face new problems. If AI completes a calculation, who is the author? Does the model provider count as a collaborator? How is the AI’s process recorded? How is the result made reproducible? These are institutional questions, not technical ones.
The map and the road builder
One sentence, nine loops, two thousand dollars.
The combination is striking because it turns something that once required years of team effort into an automated task that runs for several days.
The boundary remains. Claude did not draw a new map. It extended a road on a map humans had already drawn. It did not discover new physics, propose new principles, or invent a new method. It executed. The execution was complex, tedious, and easy to get wrong.
The equation did not change. The solver did.
Humans still ask questions, design constraints, judge meaning, and verify results. AI has begun to take on the long, dull, failure-prone execution. When execution stops being the bottleneck, the bottleneck moves to taste, direction, and verification.
Nine loops is not the end. It is a signpost. It says AI can extend a road on a map humans have drawn.
The next question: can it draw a new map itself?
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.