One Wrong Sign: The Tiny Error That Broke OpenAI’s 722-Paper Math Drop
OpenAI published hundreds of AI-written math proofs in a day. A flipped +1 brought three down, and exposed a trust problem no benchmark can fix.

On October 6, 2026, OpenAI published 722 mathematical manuscripts on GitHub, all generated by its internal unreleased frontier model, covering 372 result families. Within 24 hours, three of them were retracted.
The reason for the retraction was a sign error. In a geometry manuscript discussing the algebraicity of Weil classes, the model wrote the sign of a geometric operation as +1, whereas according to the paper’s own convention it should have been -1. This error invalidated the paper’s central argument. The stabilization-trace argument required the sum of positive and negative counts to cancel to zero. With the sign reversed, a count that should have canceled to zero became nonzero. Two derivative results that depended on its construction, “Algebraicity of the Kuga–Satake correspondence for K3 surfaces” and “The rational Hodge conjecture for products of K3 surfaces,” were also retracted.
Alex Townsend, associate professor of mathematics at Cornell University, was not surprised. “It is not surprising that one error cascades and causes problems in three manuscripts,” he said. “I suspect more errors will be found.” His suspicion has a structural basis: of the 719 manuscripts, only about 42% of the main results had completed Lean formal verification. The retracted and substantively revised content almost all came from the parts that had not been formalized in Lean.
Lean is an open-source interactive theorem prover that can turn mathematical proofs into formalized code that a machine can check line by line. Manuscripts that have not passed Lean verification are essentially still at the stage of “natural-language claims.” And with 719 manuscripts published at once, totaling tens of thousands of pages of mathematics, human review is physically impossible to complete in the short term. Townsend pointed out that, given the mathematical community’s skepticism toward AI, OpenAI “should have first published the results that had passed Lean verification,” and then separately requested help verifying the rest.
The response from an OpenAI spokesperson revealed a more specific decision logic. The company admitted that its advisory body, the Mathematics and AI Advisory Group (AGMAI), had advised publishing results without waiting for “complete formalization,” and that “about 50% of the results were published without confirmation.” AGMAI’s recommendation text explicitly opposed “using internal models inaccessible to outsiders to test difficult mathematical problems” and required labs to disclose model names, prompts, compute costs, and failure records. OpenAI claimed to have consulted AGMAI’s recommendations, but the Association for Human Mathematics (AHM) believes the company “ignored the core premise of the recommendations.” Of the 719 manuscripts, only 10 disclosed the model’s reasoning process for arriving at its results.
To understand why a frontier model would make a mistake on something as basic as a “plus or minus sign,” one needs to return to the computational nature of large language models. The core mechanism of an LLM is autoregressive next-token prediction. It learns probability distributions over massive amounts of text and generates the statistically most likely next symbol. In the context of mathematical proof, this means the model is at every step “guessing” the most likely next operation. When a paper’s own sign convention conflicts with the more common convention in the training data, the model is pulled by the “probability gravity” of the training distribution and tends to choose the high-frequency symbol, even if that symbol is wrong in the current context.
A 2025 paper diagnosed this phenomenon as “comprehension without competence.” Researchers found that LLMs could “explain” the rules of decimal comparison with textbook precision, yet when actually executing them would conclude “9.90 > 9.11.” They called this pattern “computational split-brain syndrome.” The model “can perfectly explain principles it cannot reliably execute.”
This diagnosis precisely describes the error mechanism in OpenAI’s retraction. The model “understood” the high-level structure of the stabilization-trace argument. It knew a cancellation needed to be constructed, and that the result after cancellation could support the algebraicity theorem. But when executing the individual sign operation within that structure, it relied on statistical patterns rather than symbolic semantics. The paper’s own convention (“the sign here is -1”) is a local constraint, whose information is concentrated at a specific point in the argument rather than widely distributed in the training data. When generating the content at this point, the model was overridden by the prior that “+1 is more common” in the training distribution.
If unformalized parts are the high-incidence area for errors, is the part that completed Lean formalization foolproof?
Researchers at Cambridge University and King’s College London gave a negative answer in a paper published in October 2026. The paper, titled “Navier–Stokes lost in translation,” studied OpenAI’s previously announced proof of a blow-up solution to the Navier–Stokes equations and found that “the formalized Lean proof does not correspond to the natural-language proof,” although this does not affect the possibility that the proof is mathematically correct.
The core issue is autoformalization, the process of translating natural-language mathematical text into Lean code itself. The researchers proved that providing a “semantically faithful translation” requires resolving ambiguity in natural-language mathematical text, and this problem’s computational complexity is “arbitrarily high in the Solvability Complexity Index (SCI) hierarchy,” harder than any computational problem including the halting problem. In other words, no general algorithm can guarantee error-free translation of any natural-language mathematical proof into formal code.
This means a widely accepted assumption, that passing Lean verification equals a correct proof, has a fundamental hole. If the translation process itself is erroneous, then Lean verifies only the internal consistency of the “translated code,” not the correctness of the “original natural-language argument.” The preliminary step of formalization “is itself a mapping problem from informal to formal. If the mapping is wrong, then ‘verification’ is meaningless.”
Terence Tao named the current situation “proof indigestion”: AI can rapidly generate propositions, proofs, and counterexamples, but humans cannot verify, understand, write about, teach, and absorb them in time, resulting in large numbers of “proofs no one can digest.” The metaphor is precise because it identifies a different problem: even if the proofs are right, they may not be useful.
The pace of traditional mathematical research stands in sharp contrast to this “proof indigestion.” The publication of a mathematical paper is not only the presentation of logical deduction but also a social process: through seminars, informal discussions, and peer review, the intuition of the proof is gradually explained, the formulation of methods is refined, and connections with the existing body of knowledge are established. When a commercial laboratory publishes hundreds of manuscripts at once, this process is completely bypassed. Bryna Kra, a mathematician at Northwestern University, pointed out that “peer review is gone,” and the subsequent costs of verification, revision, and knowledge integration are “shifted to the academic community.”
The Association for Human Mathematics statement put the conflict more sharply: “Mathematicians did not ask for this work. Publishing more than 700 documents at once is not an academic presentation but a display of power.” This statement was reposted by Terence Tao and co-signed by 25 Fields Medalists including Deng Yu and Peter Scholze.
The substance of this conflict is the gap between the speed at which AI generates results and the human capacity to understand them. Mathematical research is not merely the determination of the truth value of propositions; it is also the pursuit of “why.” A proof that has passed formal verification may be logically flawless, but if no one can explain “why it holds,” what its relationship is to existing theoretical frameworks, and what new directions it opens, then its contribution to mathematical knowledge is limited. Although Tao has been an active user of AI-assisted research, he explicitly pointed out that if the person prompting the model to solve a problem “cannot explain the result, participate in reporting it, or respond to peer questions, the results will be difficult to truly integrate into mathematical research.”
After the retraction, an OpenAI spokesperson said, “We welcome scrutiny and feedback from the mathematical community,” and added, “Where errors are found, we will work to correct them promptly.” From a procedural standpoint, OpenAI’s errata process was transparent and swift: within 24 hours it retracted the problematic manuscripts, revised 14, updated 13 citations, and attached a retraction note in the GitHub repository.
The problem is the publication strategy itself, which externalizes the cost of verification. AGMAI’s recommendations explicitly opposed “using internal models inaccessible to outsiders to test difficult mathematical problems” and required labs to disclose model names, prompts, compute costs, and failure records. OpenAI claimed to have consulted AGMAI’s recommendations, but AHM believes the company “ignored the core premise of the recommendations.” Of the 719 manuscripts, only 10 disclosed the model’s reasoning process for arriving at its results, and the “machine-readable associated information” required by AGMAI, used to correspond natural-language proofs with formalized results, was also not provided.
The consequences of this information asymmetry are concrete. When a company claims that its internal model has proved hundreds of unsolved mathematical problems, the mathematical community faces a choice: either invest substantial manpower to verify these claims, or choose not to trust them and refuse to participate. Townsend’s words may represent the feelings of a considerable number of mathematicians: after 15 years of mathematical research, he said he felt “threatened.”
More fundamentally, this kind of publication is changing the incentive structure of mathematical research. If the speed of “solving difficult problems” becomes the core metric for measuring AI capability, and if commercial laboratories can unilaterally decide which mathematical problems are worth attacking and which results are worth making public, then the direction and pace of mathematical research will increasingly be influenced by a small number of technology companies. Terence Tao, in a statement he previously co-signed with 24 other Fields Medalists, warned that treating difficult mathematical problems as benchmarks for model capability “may divert technological competition from the original purpose of mathematical research: the pursuit of understanding and insight.”
Returning to the sign error itself. The error is a symptom of a structural defect: in mathematical reasoning, current large language models have a systematic rupture between capturing high-level argument patterns and executing low-level sign conventions. The model can at an abstract level “know” that the stabilization trace should cancel to zero, but when generating the specific sign, it relies on statistical priors in the training data rather than the local semantic constraints of the current text.
Among these manuscripts, there are indeed results that have passed Lean formal verification; about 42% of the main results completed machine checking. Lean verification itself remains one of the most reliable proof-checking tools available. The problem is that there is a fundamental contradiction between the incompleteness of verification and the completeness of publication. When fewer than half of the results have been formally verified, publishing all of them means placing unverified content and verified content within the same narrative framework of “AI mathematical breakthrough,” making it difficult for readers to distinguish the difference in credibility between the two.
A single sign error points to the boundary of AI’s mathematical capability. The model can generate large numbers of “plausible-looking” mathematical arguments. These arguments structurally conform to the patterns of mathematical writing, yet may contain errors in local operations that human mathematicians would not make. These errors are difficult to detect precisely because the macro-structure of the argument is “correct”: it looks like a mathematical paper, uses correct terminology, and follows a reasonable argumentative procedure. Only by going deep into the specific level of sign conventions does the error become visible.
OpenAI wrote at the end of its retraction note: “We thank the mathematical community for its patience.” The commit history of the GitHub repository shows that the retraction operation occurred at 3:17 a.m. on October 7. The links to the three manuscripts were replaced with a brief errata note, whose last line reads: “The correct versions will be published subsequently.”
As of press time, the correct versions have not yet been published.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.