01 logo

OpenAI Just Solved Hundreds of Math Problems. The Fight Over How It Did That Is Just Beginning.

A TCS researcher breaks down the derandomization proof, the 3SUM collapse, and the ethics fight splitting the math community.

By JinPublished about 14 hours ago • 6 min read

On October 7, 2026, OpenAI published 722 mathematical manuscripts in a single GitHub repository. The papers cover 372 result families and reach into theoretical computer science, number theory, geometry, and mathematical physics. An unreleased internal model produced them. OpenAI gave the model about 4,000 problems. The average result used roughly three hours of ChatGPT Pro compute.

For people outside math, this looks like another round of AI publicity. For people inside theoretical computer science, the release marks several breaks in method. Some fields are being punched through fast. Others remain untouched. The research culture around them is shifting.

What Actually Landed

The first result to stop TCS people cold is item 103: L=RL=BPL. That completes log-space derandomization. The field has been circling this target since Saks and Zhou in 1999. Later work pushed closer to log⁡nlogn space but never closed the last gap. Hoza’s ECCC survey tracks that line. Item 103 ends it.

The result matters because it removes a long-standing conditional hedge. Bounded-error logarithmic-space computation now has a deterministic simulation with explicit polynomial time. The proof also appeared in Lean, so part of the verification is formal. Formal proof does not guarantee that the technique will generalize. TCS often cares as much about the method as the theorem.

OpenAI did not publish a proof of P=BPP. That would require general circuit lower bounds. Circuit complexity is absent from the release. The model did give arithmetic circuit lower bounds: Ω(n4/log⁡n) for the permanent and Ω(n2log⁡log⁡n) for division-free circuits. Those are arithmetic, not Boolean. General Boolean circuit lower bounds, the kind strong enough to imply P≠NP, remain open.

3SUM and APSP Lose Their Footing

The 3SUM and APSP disproofs hit harder. Alman and Vassilevska Williams thank Claude in the paper. The model found algorithms that refute 3SUM, APSP, and the Exact Triangle hypothesis. The authors then understood, simplified, and extended those algorithms. The results are deterministic: 3SUM in O(n1.9992) and APSP in O(n2.9995). Those are the first polynomial improvements over the textbook bounds.

Known reductions carry the damage further. The paper also refutes real-valued 3SUM and APSP, Exact Triangle, Zero-Weight k-Clique, and three rectangular online matrix-vector conjectures from van den Brand and coauthors.

SETH survives. The paper does not derive SETH’s failure from 3SUM or APSP. It also does not give a truly super-linear lower bound from SETH to 3SUM. The reduction skeleton of fine-grained complexity still stands. But several nodes that people treated as stable are gone. You cannot call a problem “3SUM-hard” or “APSP-hard” and stop there. The reduction chain itself now needs another look.

SSETH is a different case. If decades of work from IKW, Williams, and others suggest that exponential-time algorithms may have methodological blind spots, then strong hypotheses like SSETH are more fragile than they look. This release does not touch SSETH. It does make that fragility easier to believe.

k-server Closes

The k-server conjecture dates to Manasse, McGeoch, and Sleator in 1988. It says a deterministic online algorithm can reach competitive ratio kk on every metric space. Koutsoupias and Papadimitriou proved the work function algorithm hits that bound on restricted metrics in 1995. The general case stayed open.

Coester, Koutsoupias, and Zbysiński prove it. Their proof encodes the work function as a matrix. The determinant gives the work function value. Requests update the matrix through basis changes and row replacements. The competitive analysis becomes a statement in matrix algebra.

The paper says the first three-server proof came without AI help. Human insight still found the key turn. AI helped with search and verification.

The randomized version is less dramatic. Bubeck, Coester, Rabani, and others disproved the random k-server conjecture in 2022 and 2023. They gave an Ω(log⁡2k) lower bound. The new work removes a log⁡nlogn factor from algorithms on HST. That is strong. It is also the kind of improvement people expected once the deterministic proof was understood.

The Math Community Pushes Back

AGMAI, the Advisory Group on Mathematics and Artificial Intelligence, is hosted at the Institute for Advanced Study. Its members include Timothy Gowers, Martin Hairer, and Edward Witten. In late September, the group released recommendations. One line is blunt: “We do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.”

AGMAI also wrote: “The future of mathematical research cannot be confined to understanding the results of AI labs. Mathematicians must be able to pose their own questions, develop their own research methods, and explore research directions that have not yet been selected as examples of AI system capabilities.” The concern is not only release style. It is who controls the questions.

The complaints are specific. One mathematician called the release “mobster behavior.” In August, OpenAI gathered about 40 mathematicians. The company implied the model had solved hundreds of long-standing problems and said it would not release everything at once. An OpenAI spokesperson later said they did not know about that guarantee. Bryna Kra of Northwestern recalled that attendees asked for papers, not blog posts or tweets, so mathematicians could digest and use the work. Those requests were ignored.

The Navier-Stokes fight is sharper. OpenAI heard that others were close to a solution. It deployed thousands of agents to solve the Millennium Prize problem first. NYU professor Tristan Buckmaster accused OpenAI of front-running his work with Anthropic employee Levent Alpöge. Buckmaster announced three proofs on Tuesday. Less than 24 hours later, OpenAI published a full proof of Navier-Stokes existence and smoothness. Scooping is a serious violation in academia. In an AI lab’s release schedule, it can look like competition.

AGMAI’s recommendations include another line: “refrain from treating the release of mathematical results as marketing vehicles to promote their models.” That names a structural problem. AI labs both produce results and publish them. Peer review, revision cycles, and citation norms do more than check quality. They also spread power. A GitHub dump does not.

OpenAI promised to “continue exploring other community-hosted alternatives for this release” and to “further improve the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding.” Those promises point in a useful direction. They still leave the process in the lab’s hands. AGMAI’s warning asks a harder question: when knowledge production shifts, can the academic community still repair itself?

What AI Can and Cannot Chew

The release shows a clear divide. AI punches through problems with clear statements, mechanizable verification, and structured search spaces. Derandomization, 3SUM, APSP, and k-server fit that pattern. The model shows systematic search, not sudden inspiration.

AI still cannot chew general circuit lower bounds. P=BPP needs them. The release does not contain them. It gives arithmetic circuit bounds but not Boolean ones. It does not imply P≠NP. The boundary is visible: AI does well when a reduction path and a verifier exist. It struggles when a problem needs a new method.

This puts numerical-improvement work at risk. Integer multiplication at n(log⁡n)1−ϵn(logn)1−ϵ, FFT at similar bounds, matrix multiplication at n9/4, space-time simulation at T2/5: these papers aim at asymptotic improvements. Verification is mechanical. Search can be parameterized. When a model can try thousands of parameter sets in three hours, a human moving an exponent from 1.999 to 1.998 has less room.

TCS is not dead. The Simons Institute’s September report, “AI and TCS: The Next Six Months,” lists 12 actions. They include a standard “AI Methodology” section in papers, tracking AI-generated ideas, requiring authors to attest that they understand and have verified the work, clear exposition in the first 10 to 12 pages, a submission limit of five per author, and allowing reviewers to use AI while keeping human judgment on novelty and importance. The report says the community needs to adjust evaluation, not abandon the problems. TCS still has PP vs NP, Unique Games, and ETH.

The real question is which problems deserve a PhD student’s five years. If AI keeps accelerating on problems with verifiers, humans may do better by posing new problems and building new reductions. Those are the areas where current AI is weakest.

Where This Leaves Researchers

The full impact of the 722 manuscripts will take time to settle. Some signals are already clear.

Derandomization and fine-grained complexity are shaking. SETH and general circuit lower bounds still stand. TCS reduction methodology remains a human strength. Algorithm optimization is being absorbed by AI.

The ethics fight is about who controls knowledge production. When AI labs produce and publish results, academic self-regulation takes pressure. The Simons Institute’s 12 actions are a bottom-up response. Whether they can keep pace with lab release schedules is open.

For TCS researchers, the useful move is to treat AI as a powerful collaborator. Coester and coauthors did this with Claude’s algorithm. They understood it, simplified it, strengthened it, and extended it. AI searched. Humans understood and generalized. That division requires evaluation systems to change. If discovering an algorithm is worth less, understanding it must be worth more.

Welcome to the perilous new world. The nine mathematicians of AGMAI would not use the word “welcome.”

tech newsfact or fictionthought leaders

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin