Education logo

The AI That Did 37 Years of Math in 36 Hours

Claude didn't solve the Riemann Hypothesis. But it just turned a 167‑year‑old mountain into a very different climb.

By JinPublished 2 months ago • 6 min read

Prologue: 25.6 Percentage Points, Bought with 31 Million Tokens

On August 10, 2026, Anthropic published a research blog post titled "Learning more about Claude's mathematical capabilities." The headline was restrained. But the content drew attention across the mathematics community.

An unreleased research version of Claude was given a task by an employee: take a serious crack at the Riemann Hypothesis.

It worked for a day and a half. It coordinated 60 sub‑agents, ran 2,400 shell commands, wrote hundreds of Python scripts, and burned through 31 million output tokens. The result? The Riemann Hypothesis itself remained untouched. But along the way, Claude raised the lower bound on the proportion of non‑trivial zeros of the Riemann zeta function lying on the critical line. That bound had stalled at 41.6% for over three decades; Claude pushed it to 67.2%.

Thirty‑seven years: humanity had advanced that metric by 0.8 percentage points. A day and a half: Claude advanced it by 25.6.

The contrast is stark. Social media and comment sections are predictable: "Is AI about to conquer mathematics' highest peak?"

The answer is no. But the story itself is more complex and more interesting than "AI proves Riemann Hypothesis."

What Is the Riemann Hypothesis?

To understand the weight of this event, we need to go back to Berlin, 1859.

That year, the German mathematician Bernhard Riemann published an eight‑page paper titled "On the Number of Primes Less Than a Given Magnitude." In it, he proposed a conjecture about the distribution of prime numbers, later recognized as one of the most important unsolved problems in all of mathematics.

Prime numbers — such as 2, 3, 5, 7, 11, 13 … — are divisible only by 1 and themselves. Euclid proved over two thousand years ago that there are infinitely many primes, but their distribution remains a deep mystery. Mathematicians want to know: how many primes are there below a given number? How are they spread across the number line?

Riemann found a crucial clue. The distribution of primes is closely connected to a function, now known as the Riemann zeta function:

https://www.youtube.com/watch?v=cNEfMlYbazU

Here s is a complex number, written as a + bi. When the real part a is greater than 1, the series converges. For real parts less than or equal to 1, the series diverges, but mathematicians extend the function to the entire complex plane using a technique called analytic continuation.

After continuation, the zeta function has certain points where it equals zero. When s is a negative even integer (−2, −4, −6, …), the function is zero; these are the trivial zeros. But there are other zeros, all with real parts between 0 and 1, known as the non‑trivial zeros.

Riemann's conjecture was this: all non‑trivial zeros of the zeta function have real part exactly 1/2. In other words, they all lie on the vertical line in the complex plane called the critical line (Re(*s*) = 1/2).

That is the Riemann Hypothesis.

The Clay Mathematics Institute listed it as one of the seven Millennium Prize Problems, with a $1 million reward. Mathematicians call it the "Holy Grail" of mathematics, and that is not hyperbole. No one has claimed it in 167 years.

Every Step of Those 167 Years

The statement of the Riemann Hypothesis is not complicated, but for 167 years, the finest mathematicians have stumbled against it.

Since proving that 100% of zeros lie on the critical line proved intractable, mathematicians shifted the question: can I at least show that a significant proportion of zeros are on that line?

This "lower bound" problem became the main yardstick of progress. For decades, mathematicians ran a relay race, each leg painfully short:

  • 1942: Norwegian mathematician Selberg proved that a positive proportion of non‑trivial zeros lie on the critical line. Not 0%, not infinitesimal — a definite positive number. But the proof gave no concrete value.

  • 1974: Levinson pushed it to at least 1/3 (≈33.3%).

  • 1989: Conrey raised it to 40%.

  • 2011: Bui, Conrey, and Young together pushed it to 41.05%.

  • 2012: Feng reached 41.28%.

  • 2020: Pratt, Robles, Zaharescu, and Zeindler jointly pushed it to about 5/12, or 41.6%.

From 1989 to 2020, a span of 31 years, the bound advanced by 1.6 percentage points. From 1974, that is 46 years for about 8.3 points.

Every step was backed by detailed analysis, refined methods, and years of deduction. The problem is inherently deep; it demands not computation, but profound mathematical insight.

What Did Claude Actually Do?

Claude's methodology was very different from that of human mathematicians.

It started with large‑scale trial and error. According to the report, the research version first generated 650 possible proof strategies, all failures. A human researcher facing 650 consecutive failures would be demoralised. Claude was not. After repeated prompts of "keep going," it entered a deep‑thinking mode, silently coordinating 60 sub‑agents to cross‑review each other's work, run numerical verifications, and search the literature. It downloaded 54 papers from arXiv to cross‑check its results.

Two thousand four hundred shell commands. Hundreds of Python scripts. 31 million output tokens. In the end, the lower bound moved from 41.6% to 67.2%.

The technical path was not Claude inventing new mathematics. More precisely, it did something human mathematicians find extraordinarily difficult: it stitched together existing technical routes at scale, combined them, tested them, and found the optimal path through massive computation.

Anthropic's blog openly stated that Claude's work is a combinatorial innovation on known frameworks, not the creation of new mathematics, which is exactly why it did not prove the Hypothesis itself. But being able to absorb the core results of its predecessors in a day and a half, and find a combinatorial path that eluded human mathematicians for decades — that alone has made many researchers rethink their tools.

The result was formally verified using Lean 4 and reviewed by two leading analytic number theorists, Brian Conrey and Dan Goldston. The math is solid.

The Gap Between 67.2% and 100%

The gap between 67.2% and 100% is not 32.8 percentage points; it is a qualitative difference.

The Riemann Hypothesis demands: all non‑trivial zeros lie on the critical line. This is a yes/no proposition. Proving 10% lie on the line is logically no closer to "all" than proving 90%, because you still cannot rule out the remaining 10% (or 0.1%) as counterexamples.

Anthropic itself made this clear in the blog: they do not expect this technical route to solve the Riemann Hypothesis.

Terence Tao, in an earlier interview, touched on this. He argued that solving the Hypothesis would almost certainly require creating a new kind of mathematics, or forging a new connection between two previously unrelated fields. Current AI capabilities, even Claude 4, still operate within the bounds of existing mathematical tools, far from "creating new mathematics."

67.2% makes for a sensational headline. To mathematicians, it is a big step forward, but the summit remains far away.

Where Is the Change?

So where does the groundbreaking significance lie?

Perhaps not in mathematics itself, but in how AI does research.

Claude's performance revealed several capabilities that human researchers lack. It conducted tireless large‑scale trial and error: after 650 failed ideas, it persisted, coordinating 60 sub‑agents and executing thousands of commands without complaint. It digested literature at lightning speed, reading and cross‑checking 54 arXiv papers in a day and a half, a task that would take a typical PhD months or years. Its reasoning was systematic, not inspired; it worked like a combinatorial optimizer, trying every arrangement of known tools. And it closed the loop with formal verification: Lean 4 checked each step, leaving almost no room for human error.

Anthropic's blog was not declaring "AI conquers the Riemann Hypothesis." It was demonstrating Claude's collaborative potential in mathematical research. Claude is more like a super‑postdoc that never sleeps, never gets frustrated, and processes information orders of magnitude faster than any human; its depth and creativity are still limited, but its breadth and efficiency are themselves a new kind of productivity.

Tao once predicted that solving the Riemann Hypothesis might require a kind of human‑AI collaboration "that does not yet exist." Claude's experiment may be a prototype of that future: AI handles large‑scale exploration, combination, and verification; humans set direction, interpret results, and provide deep insight.

Epilogue

For 167 years, the Riemann Hypothesis has stood like a high mountain, drawing generation after generation of mathematicians. Claude did not reach the summit. But it drove a new kind of piton into the steep rock face.

The symbolic importance of that piton may exceed its immediate mathematical value. It indicates that AI is becoming more than a calculator or literature search tool; it is becoming a collaborator that can push the frontier itself. It still makes mistakes. It still cannot "create new mathematics." But its speed, breadth, and tenacity are already changing the very process of doing mathematics.

The Holy Grail remains distant. The $1 million prize is still unclaimed.

But the path to that Grail, starting August 10, 2026, looks different now.

interviewdegreecoursescollege

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin