A theoretical physicist issued a public challenge, betting that AI could not crack one of his old field's toughest open problems. Anthropic's research team took up the bet using its Claude model, and the result rewrote a decades-old computational record in particle physics for a cost of roughly 100 to 2,000 dollars.
A Challenge From a Former Physicist
The story began with Matt von Hippel, a blogger who once worked as a theoretical physicist before moving into science writing. On his blog, he laid down a challenge to AI companies, picking a target from a subfield of particle physics known as amplitudeology, which deals with the formulas physicists use to predict how subatomic particles interact.
These formulas, called scattering amplitudes, are notoriously hard to compute in full. Physicists typically work with approximations that stop at a certain number of "loops," a measure of how complex the particle interactions are allowed to get. Most published amplitude calculations only reach two loops, and only a handful reach three. Even the most well-known high-precision prediction in particle physics relies on a calculation carried out to five loops.
Von Hippel's challenge asked for nine loops in a toy model theory called N=4 super Yang-Mills. The theory is not meant to describe the real world; it assumes each particle has four hypothetical supersymmetric partners, a setup that makes the physics unrealistic but mathematically more tractable. That is precisely why amplitudeologists use it as a proving ground for new computational techniques before applying them to real-world problems.
Two Independent Methods, One Result
Two physicists at Anthropic picked up the challenge in late August. They ran the calculation with a model called Claude (Fable 5.1) inside Claude Science, a research-oriented environment built to give the model structured tools for scientific work. Their prompt was strikingly simple: compute the six-particle amplitude at nine loops, then keep going with periodic status updates.
Claude worked the problem two separate ways. The first was the traditional "bootstrap" method, narrowing down the shape of the answer by checking it against everything already known about the constraints the result has to satisfy. Run in Python with the SymPy library, this portion of the calculation needed the equivalent of 96 CPUs running for a week, at a cost of roughly 100 dollars (about 15,600 yen). The second approach went through a related, more tractable quantity called a form factor. Counting the model-inference time needed to run Claude for the full task, the total bill for an end user came to somewhere between 1,000 and 2,000 dollars (about 156,000 to 312,000 yen).
The finished result contains more than 107,000 nonzero coefficients, spread across eight files that each exceed 100 megabytes.
Verified by the Researcher Who Held the Previous Record
The result was checked by Lance Dixon, a professor of particle physics and astrophysics at SLAC National Accelerator Laboratory and Stanford University, and one of the pioneers of the technique. Dixon himself had held the previous human record, reaching eight loops. After Anthropic reached out on September 1, he spent roughly two weeks validating Claude's nine-loop amplitude by converting it into the form-factor format his own group has worked with for years.
Dixon noted that the calculation's framework is extremely delicate: a single mistake anywhere in the process can cause the whole thing to collapse, and much of the fine technical judgment involved is never fully written down in any publication. That meant Claude had to build the supporting code essentially from scratch. Working through the verification, Dixon came away with the impression that Claude understood his group's 2019 and 2023 papers better than anyone outside the group of co-authors who wrote them.
Separately, just days after Anthropic contacted von Hippel, a team led by Song He at the Chinese Academy of Sciences in Beijing announced it had independently worked out part of the nine-loop result, a piece known as the symbol. Song's team used the AI model GPT-6 to help compute some of the constraints, though the overall framework was built by the researchers themselves. Human experts and AI systems were, in effect, converging on the same frontier problem at almost the same time.
Summary
Von Hippel's own reading of the result is measured: Claude did not discover some entirely new computational principle, he says, but simply ran established methods with more computing power than researchers had previously thought to try. Even so, he argues that showing a problem long assumed to be out of academic reach was in fact solvable carries real weight. He points to a striking pace of change: as recently as March, AI was tackling physics problems the way a student would, with plenty of hand-holding and mistakes; by September, it completed a genuine frontier calculation with nothing more than instructions to keep going. Dixon, for his part, says he was not rattled by the outcome, since his group had already built the tools needed to check any AI-generated candidate solution. But he is candid about what it signals: AI has arrived as a working presence in physics research itself.
Source: [1] https://www.anthropic.com/research/yes-claude-can-do-nine-loops
※1 USD = 156 JPY (as of September 26, 2026) ※Thumbnail image is an AI-generated illustration
