AGMAI Publishes Responsible-Release Rules for AI-Generated Mathematics
The mathematicians' advisory group published its first guidance on September 29, asking labs to stop testing advanced math problems on proprietary models, funding human understanding of ununderstood AI proofs, and warning of a two-tier system if model access stays closed.
AGMAI, the independent advisory group formed by nine leading mathematicians on September 21, published its first substantive document on September 29: guidance on the responsible release of AI-generated mathematics, built on more than 600 replies to a community survey. The headline ask is blunt. Labs should stop using the field as an unpaid benchmark: the group writes “we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models,” a direct reference to the OpenAI situation the group was formed around, in which OpenAI announced the existence of many results without giving details, including the Navier-Stokes blow-up claim from September 8 that remains unverified.
The Core Rule: Release Comes With a Duty to Explain
The guidance’s central principle converts a public-relations announcement into a binding obligation: labs that release AI-generated results without immediate human understanding “must take responsibility for ensuring that human understanding will follow,” including paying for it. Understanding itself must stay “organic and community led,” not directed by the lab that produced the result. The document then splits mathematical output into two regimes with different rules. Human-understood results follow traditional norms: preprints, peer review, talks. Ununderstood results, the AI proof nobody has checked, carry a checklist: proper citation of prior literature; clean, conventional proof write-ups; deposit in independent scholarly repositories with persistent identifiers; disclosure of the model name, prompts, a summarized chain of thought, compute time, and cost; formalization where possible; and documentation of how the problems were chosen and what failed along the way.
That checklist is quietly the most technical artifact the AI-governance debate has produced. It specifies provenance metadata for machine-generated claims, which is exactly the kind of standard that can be enforced by repositories and journals rather than trusted to lab goodwill.
Step II Makes Verification a Funded Discipline
The guidance’s second step addresses the economics that have governed this whole saga. Verifying an 88-hour multi-agent proof is months of skilled labor that nobody is currently paid to do, which is why the Navier-Stokes claim has sat unverified since September 8. AGMAI’s answer: labs that release ununderstood results should fund conferences, workshops, working groups, postdocs, and expository writing, so that human understanding becomes a funded workstream rather than a volunteer effort squeezed between other duties. The group’s founding statement insisted members take no payment and hold no decision power at any AI company, and the funding flows in this document respect that design: money goes to the community’s understanding of results, never to the group itself, and the community input behind the document came from more than 600 mathematicians rather than a closed committee.
The Access Warning: A Two-Tier System
The third section widens the lens from release practice to market structure. Proprietary internal models that produce mathematics the public cannot access risk “a two-tier system where labs outrun the rest of the field,” and the group advises labs to “grant the global mathematical community broad, equitable access to their publicly available models.” This is a research-access demand aimed at exactly the moment when frontier labs are proving results internally on unreleased models (OpenAI’s 100-plus problems claim, the Navier-Stokes result) while selling older generations publicly. It also intersects with this week’s model news: as GPT-6.1 Sol arrives at $2/$10, raw access to a strong model is cheap, but the models doing the landmark mathematics remain internal, and AGMAI is on record that this gap is the problem.
What to Watch
Three threads. First, whether OpenAI responds by opening its claimed results through AGMAI’s process, which would convert the group from petitioners to gatekeepers and test the “no decision power” clause in both directions. Second, whether repositories and journals start requiring the provenance checklist for AI-assisted submissions; if arXiv or the Clay Institute adopts even part of it, the guidance becomes infrastructure. Third, whether other fields copy it: the structure (declare results, fund verification, disclose provenance, keep understanding community-led) is field-agnostic, and the month’s pattern suggests mathematics is simply first. The group has now done in nine days what most AI governance efforts manage in years: formation, consultation, and a concrete, enforceable-seeming standard.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How Function Calling Works in LLMs
Function calling lets LLMs interact with external systems by requesting structured tool executions. Here's how the loop works, how to define tools, and what to watch for across providers.
Top Mathematicians Form an Independent AI Advisory Group, With OpenAI as Its First Client
Nine leading mathematicians announced the Advisory Group on Mathematics and AI on September 21, hosted at the Institute for Advanced Study. It formed after OpenAI approached members about an external advisory board, and its first task is advising OpenAI on releasing a large batch of AI-produced mathematical results.
Researchers Welcome Embedded Safety Evaluators, Then Ask the Obvious Question
Safety researchers welcomed Anthropic and OpenAI's commitment to embed independent evaluators inside their labs as unprecedented access, while TechCrunch's coverage asks whether the evaluators will really be independent.
Po-Shen Loh Makes the Case That We Will Always Need Human Mathematicians
In a guest post on Terence Tao's blog, Carnegie Mellon mathematician Po-Shen Loh argues that AI progress multiplies the number of control points requiring skilled human oversight, so demand for human experts will grow faster than AI can replace them.
OpenAI's Chief Scientist Says No Lab Has Solved Alignment, Calls for Slowdowns
In an essay titled An Alien Mind, OpenAI chief scientist Jakub Pachocki describes AI as grown rather than designed, admits no lab has solved alignment, and calls for voluntary slowdowns and third-party audited safety frameworks.