Education logo

The $2,000 Problem

OpenAI's Astra just cracked ten unsolved math problems for less than a PhD student's monthly rent. The future of research just got a lot more expensive — for humans.

By JinPublished 2 months ago • 10 min read

On August 1, 2026, OpenAI posted an announcement titled Ten Advances in Mathematics. Their next-generation model, Astra, had solved ten problems across high-dimensional geometry, coding theory, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. Several of these—constructing non-sofic groups, disproving the Connes rigidity conjecture, proving an exponential parallel repetition theorem for general two-player quantum games—had been open for two to three decades, some for half a century.

One number from the announcement kept getting cited: the token cost for generating these solutions, priced at OpenAI's Sol API rate, came to roughly $2,000.

What does $2,000 look like on the books of the U.S. National Science Foundation? NSF annual stipends for a mathematics PhD student run between $35,000 and $45,000. $2,000 is less than two and a half weeks of that. In China, the National Natural Science Foundation's Mathematics Tianyuan Fund averages 200,000 to 400,000 RMB per project over two to three years. $2,000 converts to about 14,000 RMB—in Beijing, that covers four to five months of a PhD student's research assistantship.

Here is the gap: even if you gathered top mathematics departments from around the world, recruited senior faculty across eight subfields, and gave them five years, no one would bet on cracking all ten. A single human researcher cannot go deep in all these areas over a lifetime.

So this announcement is not primarily about whether AI can do math. It is about what happens when the core intellectual labor of a research project costs $2,000. What is left of "doing research"?

What follows is a breakdown across three dimensions: technical verification, academic ecosystem, and historical placement.


I. What the Announcement Said—and What It Didn't

OpenAI's announcement came with a 249-page preprint. The ten problems roughly break down as follows:

  • Group theory: explicit construction of non-sofic groups. Open ~25 years.

  • Operator algebras: disproof of the Connes rigidity conjecture. Dates to the 1970s.

  • Quantum information: exponential parallel repetition for general two-player quantum games. Prior results only covered special cases.

  • High-dimensional geometry & coding theory: improved upper bounds on sphere packing densities and several asymptotic coding bounds.

  • Arithmetic circuit complexity: progress on permanent lower bounds. A problem stuck for over forty years.

  • Lattice cryptography: improved hardness lower bounds for the shortest vector problem.

  • Extremal combinatorics: solutions to several Erdős problems, some traceable to the 1930s.

OpenAI itself noted that "the importance of the results varies." If the Connes disproof holds, operator algebras get restructured. If non-sofic groups are real, geometric group theory textbooks get rewritten. Some of the Erdős solutions, while genuine advances, have narrower reach. No research program produces ten equal-sized breakthroughs—the gradient is itself unremarkable.

Each result came with a Lean formal proof certificate. Lean is a proof assistant that checks every step from axioms to conclusion for gaps, type errors, and broken inferences. In that narrow sense, Astra's arguments are machine-verifiable.

But there are two gaps between formal verification and "the paper is correct."

First: does the formal system's definition of "sofic group" match exactly what the mathematics community means by the term? If the formal version deviates at some boundary condition from the natural-language version, what gets verified is "that formal proposition," not the original problem. OpenAI acknowledged this in their FAQ: Astra generated the mathematical arguments; humans translated them into readable papers; then the model generated Lean certificates. The human role is translator and verification architect. The translation has not yet been peer-reviewed line by line.

Second: a logically correct proof can be inelegant. It can work. It can also fail to embed into a broader framework, point toward deeper structure, or make anyone who reads it feel "of course." The mathematical community's judgment of value sits in a different dimension—formal certificates do not address it.

By early August 2026, few mathematicians worldwide had found time to read Astra's papers closely. An independent blog tracking AI mathematics progress collected scattered reactions: operator algebra researchers found the Connes disproof path "unexpected but technically plausible," while flagging a key lemma for closer scrutiny. Group theorists noted that the non-sofic construction relies on a previously understudied combinatorial object whose existence proof needs independent verification. On the extremal combinatorics side, several Erdős solutions were called "technically correct but not particularly clean."

On the same day, Anthropic researchers went public, saying they had independently re-derived five of the ten results in 24 hours using their publicly available model Claude Fable. If confirmed, this suggests a different picture: perhaps some of these problems are not as hard as OpenAI's framing suggests, or perhaps current AI models are more broadly capable in this domain than we assumed—Astra is not a unique exception.

But even if we accept a much weaker conclusion—that an AI can produce partially correct, partially meaningful mathematical results at extremely low cost—the basic assumptions about "doing mathematics" have already shifted.


II. Two Thousand Dollars and What It Cannot Buy

The $2,000 figure is frequently cited without a crucial qualification: OpenAI counted inference tokens only. They did not include training cost—industry estimates for a model of Astra's scale run into the hundreds of millions of dollars. They also did not include the human hours spent translating arguments into papers and running Lean verification. OpenAI said "humans organized the arguments"—that likely consumed weeks or months of their mathematics team's time.

But adding those costs back does not reverse the direction of the tilt. Once a model is trained, the marginal cost of research drops toward zero. A human researcher's marginal cost is time, energy, and salary—and while experience accumulates with each additional year, the curves for physical stamina and attention slope downward. The two cost structures are fundamentally different.

OpenAI did not release another set of numbers: how many attempts Astra made on these ten problems, which ones failed, and what criteria selected the final ten. If Astra succeeded on ninety out of a hundred randomly chosen open problems, that is one story—it demonstrates solid general research capability. But if these ten were winnowed from tens of thousands of attempts, the picture looks closer to extremely efficient heuristic search—impressive, but not yet at the level of "understanding mathematics."

An analogy: when AlphaGo beat Lee Sedol, its training process involved tens of millions of self-play games. No one dismissed AlphaGo's Go ability because it was later repeatedly validated in public matches. But mathematical research has no "repeat matches"—each problem is a unique, one-off experiment. Success rate becomes a hidden variable that cannot be inferred from public information. This is the fundamental epistemic dilemma the academic community faces when evaluating this class of AI output.


III. PhD Students and the Cost-Benefit Ledger

Consider the argument that recruiting PhD students in theoretical directions is already a losing proposition. Its premise is that a professor's core performance metric is the quantity and speed of results. If the measure is papers and project outputs, then running AI with the same funds and time may indeed yield higher expected output than training a PhD student.

But utilitarianism operates differently across institutional contexts.

At U.S. research universities, faculty evaluation depends heavily on PhD training itself—where graduates place, what they publish, their subsequent academic influence. These are core dimensions of promotion and reputation. PhD students are not research tools; they are the basic unit of academic reproduction. If professors stop training them, labs break, the pipeline breaks, and the evaluation metrics tied to mentorship collapse.

At some Chinese institutions, the situation is more delicate. If evaluation emphasizes paper counts, grant money, and talent titles, while the "educational" dimension of PhD training is underweighted, then there is a theoretical incentive to replace some human research labor with AI.

But there is a deeper conflation here: producing results and training researchers are two different activities. The core product of doctoral education is not papers—it is people. Even if AI solves every existing mathematical problem, society still needs people who can grasp the depth of these solutions, judge their significance, and integrate new knowledge into the broader intellectual landscape. If professors stop producing such people, mathematics as an intellectual activity sustained by human understanding faces a crisis deeper than "AI replaces humans"—the knowledge remains, but the people who can digest it are gone.

If this diagnosis holds, the goals of PhD training in theoretical directions need to shift.

One shift: from problem-solving to problem-asking. If AI can efficiently solve given problems, the human comparative advantage shifts upstream—to discovery, formulation, and priority-setting. Training the ability to judge "what problems are worth solving" and "which directions might open deep insights" becomes more critical than training technical dexterity alone.

Another shift: from proof construction to proof interpretation. Formal proofs generated by AI need to be translated into readable mathematical narratives and embedded in broader theoretical contexts. The next generation of mathematicians needs training as interpreters and integrators of mathematics, not just theorem-provers.

A third shift: from lone work to human-AI collaboration architecture. How to formulate a problem AI can take on, how to design search spaces and validation strategies, how to judge whether an AI output has genuine mathematical meaning—these become core skills.

None of these shifts makes doctoral training easier. On the contrary, asking a good question demands higher-order insight and creativity than solving an existing one. The mathematician's intellectual labor moves from the execution layer to the strategic layer. The difficulty does not decrease; it increases.

If AI-assisted research becomes routine, the funding system will also adjust. Agencies like NSF and ERC may need to add criteria around "AI reproducibility" and "human-AI collaboration transparency." Reviewers will ask not only "is this method novel," but also "if AI can do this too, what is the human researcher's distinctive contribution." Journal peer review will face pressure—if every submission comes with a Lean certificate, reviewers no longer need to check line-by-line logic, but their work shifts toward conceptual review: judging importance and naturalness. This is not downgrading; it is upgrading. Open science may gain momentum—OpenAI released papers, formal proofs, and reasoning traces. If this becomes standard, black-box AI research will face disclosure pressure, requiring researchers to reveal AI involvement and failed attempts.


IV. Astra's Place in AI History

Place Astra on the timeline:

  • 1997: Deep Blue beats Kasparov. AI surpasses humans in a closed game.

  • 2016: AlphaGo beats Lee Sedol. AI holds its own in Go, a game of vast state space beyond brute force.

  • 2020: AlphaFold 2 solves protein folding. AI breaks through on a concrete scientific problem.

  • 2024: OpenAI's o1 series exceeds human average on math benchmarks. AI approaches human-level formal reasoning.

  • 2026: Astra solves ten long-open mathematical problems. The difference: this is not a high score on a predefined benchmark. This is new ground opened at the actual frontier of knowledge.

Before Astra, AI's achievements in mathematics fell into two categories. One: formalizing existing proofs—the Lean community has been doing this. Two: solving contest-level problems in controlled settings like the MATH benchmark and IMO problems. Astra is the first AI to produce potentially significant results on problems where even researchers did not know the answer.

You sketched a scenario earlier: "at that point, you might only need to specify a broad direction, and a powerful AI could run continuously and accumulate large amounts of new knowledge." Astra's demonstration already shows a prototype—a system capable of producing verifiable mathematical knowledge across fields at near-zero marginal cost. But a "continuously running automated research system" still lacks three things.

First, autonomous problem discovery. Astra solved human-posed problems. Genuine automated research requires a system that can autonomously identify knowledge gaps and prioritize directions. Second, diagnosis and strategy adjustment from failure. Human researchers learn from failure and adjust paths. Current AI does not. Failure is just failure—it produces no meta-cognition. Without this, so-called automated research can only brute-force search, far less efficient than human-guided exploration. Third, knowledge evaluation and integration. How newly generated knowledge gets valued, integrated with existing systems, and translated into teachable forms—these still depend heavily on human labor.

So "fully automated research" remains distant. But "highly automated human-AI collaborative research"—AI handling 90% of exploration and verification, humans handling strategic direction and deep interpretation—will likely become routine in mathematics and theoretical computer science within three to five years.

If this projection holds, the work of mathematicians and theoretical computer scientists will change. One end state: AI handles breadth—scanning vast regions of mathematical space at extremely low cost, finding connections, counterexamples, constructions. Humans handle depth—judging which discoveries truly matter, which point to deeper structures, which can anchor new theories. Both handle transmission—converting new knowledge into something teachable, ensuring mathematics remains human intellectual wealth, not machine internal state.

In this picture, mathematicians are not "unemployed." They are moved to a higher level of abstraction. Just as calculators did not put mathematicians out of work—they outsourced arithmetic and let mathematicians move upward—AI is outsourcing proof search and letting humans continue upward.


V. What Remains

Back to the opening question: how to evaluate OpenAI Astra's announcement?

This is neither a declaration that AI has fully surpassed human mathematicians, nor is it merely a lucky brute-force find. It is a signal pointing toward a deep transformation in how mathematics gets done.

On the positive side: Astra demonstrates the possibility of high-intensity exploration at the mathematical frontier at vanishingly low cost. If these results are confirmed and have lasting impact, they will push forward multiple subfields within weeks. For a discipline whose normal tempo is slow accumulation, "slow" will no longer be inevitable.

On the cautious side: independent verification must hold. Formal certificates lower the barrier to logical checking, but they cannot substitute for human judgment about "mathematical meaning." OpenAI also needs to disclose more experimental details—failure rates, selection criteria, human intervention levels—for outsiders to make reliable assessments of its true research capability.

Finally, what this event truly presses against is the tension between utilitarian calculation and value commitment. Beyond the "$2,000 versus PhD salary" comparison, training the next generation of researchers who can understand, interpret, and advance mathematical knowledge—that value cannot be entered into a cost-benefit spreadsheet. Even if AI eventually solves every mathematical problem, understanding those solutions is itself an irreducible part of human intellectual life.

For young scholars entering or about to enter theoretical research: your skill set will need to change, but your task has not disappeared. The future mathematics community will consist of three kinds of people: those who build AI, those who use AI, and those who decide what problems AI should solve—and how to understand its answers. The last kind will determine the direction of mathematics as a human intellectual activity.

Late on the night of August 1, 2026, a mathematician at Peking University wrote on social media: "Tomorrow I'll open Lean and go through that non-sofic construction line by line. If it's wrong, this is over. If it's right, I need to rethink my dissertation topic."

Mathematics will not disappear, nor will the people who do it. They are just learning to become another kind of mathematician.

studentstemdegreehigh school

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin