Ai Engineering 3 min read

OpenAI Publishes 722 AI-Generated Mathematics Manuscripts, From π to the Hodge Conjecture

OpenAI released 722 mathematical manuscripts in 372 families produced by an internal model, hosted on GitHub with Lean formalizations and reasoning traces, alongside an investigation clearing accusations that the model stole a researcher's results.

OpenAI has released 722 mathematical manuscripts, organized into 372 families, produced by a single internal unreleased model, hosted on GitHub under Apache 2.0 with Lean formalizations and abridged reasoning traces for ten highlighted families. The announcement page, “Sharing AI progress in mathematics”, hit 803 points on Hacker News with 734 comments, and the scope is hard to overstate: results touching the irrationality exponent of π, the Mahler conjectures, Kaplansky’s direct-finiteness conjecture, the Mézard-Parisi spin glass formula, free group factor isomorphism, the 3D relativistic Vlasov-Maxwell system, and the Hodge Conjecture for CM abelian varieties, among others. Most results came from one procedure at roughly three hours of ChatGPT Pro thinking compute per result, across about 4,000 posed problems.

The Buckmaster Investigation Is the Other Half of the Story

The release ships alongside a finding that resolves the ugliest thread of the September controversy: accusations that the model had stolen the work of a human researcher. Per the Navier-Stokes solution page, an investigation “confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.” That is a direct answer to the insinuation, circulated after the September 8 announcement, that the Navier-Stokes result had absorbed a human mathematician’s unpublished progress through prompts. Clearing it before shipping 722 more manuscripts was reputationally necessary; shipping the provenance trail alongside is what makes the answer checkable.

This Is AGMAI’s Checklist, Mostly Implemented

Measured against the responsible-release guidance AGMAI published on October 1, the release implements a striking amount of it. Deposit in independent scholarly repositories: GitHub with versioned releases. Provenance disclosure: the model is identified as an internal unreleased system, with compute described per result. Formalization where possible: a Lean directory with a formalization catalog, though not all 722 are formalized. Reasoning traces: published for ten families, abridged. What remains unimplemented is AGMAI’s human-understanding requirement: the README states plainly that results are at mixed verification stages, that unformalized ones “could have issues,” and that fixes will come as new versions. Stephen Wolfram’s warning that results without human framing are “born alien” is the exact risk here, acknowledged rather than solved.

The Reception Is Split Down the Same Old Line

The Hacker News thread captures the fork. One camp read the release as vindication-by-pressure: OpenAI engaging with the mathematical community only “even if they had to be publicly shamed into doing so.” Another camp argued the opposite, that gatekeeping had made permission-seeking necessary at all, with several commenters noting the community asked for something different from what was delivered. The substantive dispute underneath is about what 722 unverified manuscripts are: a gift of candidate results the community can now triage, or 722 obligations of verification dropped on a field with finite labor. AGMAI’s funding demand, labs paying for the human understanding that follows, is the unanswered half of exactly that question.

What to Watch

Three things. First, the verification grind: 722 results at mixed verification stages will be triaged for months, and the first formalized-and-confirmed results (or the first refuted ones) will calibrate what three hours of thinking compute per result actually buys. Second, AGMAI’s response, since the group exists to advise on exactly this kind of release and its judgment of whether OpenAI met the guidance is the first real test of the standard. Third, the provenance precedent: the Buckmaster investigation sets a template, investigation findings published alongside releases, that every lab shipping AI-generated results will now be measured against.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How Function Calling Works in LLMs

Function calling lets LLMs interact with external systems by requesting structured tool executions. Here's how the loop works, how to define tools, and what to watch for across providers.

Ai Engineering

Stephen Wolfram on AI Mathematics: Results Without Human Framing Are 'Born Alien'

Stephen Wolfram's September 28 essay argues AI can mine and connect existing mathematics but that the essence of pure math is human choice of what to ask, warning that AI-generated results absent human framing are 'born alien.'

Ai Engineering

AGMAI Publishes Responsible-Release Rules for AI-Generated Mathematics

The mathematicians' advisory group published its first guidance on September 29, asking labs to stop testing advanced math problems on proprietary models, funding human understanding of ununderstood AI proofs, and warning of a two-tier system if model access stays closed.

Ai Engineering

Top Mathematicians Form an Independent AI Advisory Group, With OpenAI as Its First Client

Nine leading mathematicians announced the Advisory Group on Mathematics and AI on September 21, hosted at the Institute for Advanced Study. It formed after OpenAI approached members about an external advisory board, and its first task is advising OpenAI on releasing a large batch of AI-produced mathematical results.

Ai Engineering

Claude Formalized Fermat's Last Theorem in Lean After 11 Days of Autonomous Work

Anthropic announced Claude produced a complete machine-checked Lean formalization of Fermat's Last Theorem after working largely autonomously for 11 days, generating verified proofs of 30,300 intermediate theorems.