Ai Engineering 4 min read

OpenAI Withdraws Three AI Math Papers After a Sign Error Surfaces

OpenAI's math repository recorded its first withdrawals on October 7: three K3-surface-related papers fell to a sign error that invalidated a stabilization-trace cancellation argument, while 14 manuscripts were revised and formalizations reached 42% of top-line results.

The first withdrawals have landed in OpenAI’s AI-mathematics repository. An update to the project’s history file recorded on October 7 that a “sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers,” leading to the withdrawal of three manuscripts: the algebraicity of Weil classes on split abelian eightfolds, the algebraicity of Kuga-Satake correspondences for K3 surfaces, and the rational Hodge conjecture for products of K3 surfaces. The withdrawn papers now carry notices explaining the flaw and linking to archived versions, so nothing has been memory-holed. This is the verification grind beginning, exactly one day after the 722-manuscript release flagged that unformalized results “could have issues.”

The Withdrawals Are the System Working

It is worth being precise about what happened, because the temptation will be to read this as the AI-math bubble puncturing. A sign error in one argument invalidated a construction, which took down the two papers that depended on it, and all three were withdrawn with public notices and archived originals. Compare the alternative history: three flawed papers quietly sitting in a repository labeled as results, discovered by some graduate student a year later, with no version notes. The withdrawal is the release’s versioning promise, “revisions will be recorded as new versions with old releases preserved,” executing exactly as written. It is also the AGMAI guidance working as designed: the checklist demanded provenance, disclosure, and accountability, and a public withdrawal notice with an archived original is what accountable correction looks like in a post-AI-math world.

The Rest of the Update Is Genuinely Encouraging

The same update revised 14 other manuscripts with “proof repairs, corrected statements, clearer hypotheses and dependencies,” spanning five affected clusters from Lipschitz heights to the exact Birch-Swinnerton-Dyer formula, and updated 13 more papers to reference the revised editions. Formalization progressed too: six new Lean formalizations bring the top-line results to 300 of 719, about 42%, up from the release-day count. The formalization number is the one that matters most, because Wolfram’s warning was that autoformalization can verify the wrong theorem, and a growing Lean-verified fraction is the only cure for both wrong theorems and squirrely interpretations. One day into a months-long verification program, a 42% formalization rate moving upward and three bad papers already caught is, frankly, a better error rate than most human subfields manage at this stage of a large release.

What Three Withdrawals Teach About the 722

The patterns in the update are information about the model’s failure modes. A sign error propagating through a cancellation argument and its dependents is classic algebraic-geometry failure: locally plausible steps, globally wrong, exactly the kind of error that survives casual reading and falls to formalization or deep review. That suggests where the remaining 400-plus unformalized manuscripts are most likely to be wrong: not in the statements, but in long chains of technical argument, which is also where human mathematicians err. The practical calibration for anyone using these results: the 42% formalized set is the trustworthy core, the rest is draft-quality pending verification, and the two categories are clearly labeled. No human journal has ever offered that granularity.

What to Watch

Three things. First, whether the withdrawn cluster returns: sign errors are repairable, and a corrected resubmission of the K3-surface papers would complete the demonstration that the release-response cycle works end to end. Second, the formalization curve: 300 of 719 moving toward some equilibrium tells the community how much of the corpus can ever be kernel-verified. Third, the meta-question this update answers for every lab watching: OpenAI published flawed results, withdrew them publicly within days, and the sky did not fall. That precedent, fast public correction as the standard response to AI-generated errors, may end up mattering more than any single theorem in the repository.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How Function Calling Works in LLMs

Function calling lets LLMs interact with external systems by requesting structured tool executions. Here's how the loop works, how to define tools, and what to watch for across providers.

Ai Engineering

OpenAI Publishes 722 AI-Generated Mathematics Manuscripts, From π to the Hodge Conjecture

OpenAI released 722 mathematical manuscripts in 372 families produced by an internal model, hosted on GitHub with Lean formalizations and reasoning traces, alongside an investigation clearing accusations that the model stole a researcher's results.

Ai Engineering

Stephen Wolfram on AI Mathematics: Results Without Human Framing Are 'Born Alien'

Stephen Wolfram's September 28 essay argues AI can mine and connect existing mathematics but that the essence of pure math is human choice of what to ask, warning that AI-generated results absent human framing are 'born alien.'

Ai Engineering

AGMAI Publishes Responsible-Release Rules for AI-Generated Mathematics

The mathematicians' advisory group published its first guidance on September 29, asking labs to stop testing advanced math problems on proprietary models, funding human understanding of ununderstood AI proofs, and warning of a two-tier system if model access stays closed.

Ai Engineering

Top Mathematicians Form an Independent AI Advisory Group, With OpenAI as Its First Client

Nine leading mathematicians announced the Advisory Group on Mathematics and AI on September 21, hosted at the Institute for Advanced Study. It formed after OpenAI approached members about an external advisory board, and its first task is advising OpenAI on releasing a large batch of AI-produced mathematical results.