25 Fields Medalists Warn of a Severe Misalignment Between AI and Mathematics
Terence Tao joined 24 other Fields Medalists in a September 11 declaration arguing that AI companies using mathematical problem-solving as benchmarks is severely misaligned with the needs of mathematics itself.
Twenty-five Fields Medalists, including Terence Tao, published a joint declaration on September 11 titled A Severe Misalignment of AI in Mathematics, and the argument deserves more attention than the headline. The signatories are not warning that AI will destroy mathematics. They are warning that AI companies are using mathematical problem-solving as a benchmark for progress in ways that misalign incentives with what mathematics actually needs, a critique published days after Anthropic’s Claude formalized Fermat’s Last Theorem in Lean.
The Complaint Underneath the Benchmark Fight
The declaration’s core concern is what optimization does to a discipline. When labs race to solve famous open problems as demonstration exercises, the problems become benchmark fuel rather than research objects: solutions get burned as evaluation targets, progress gets measured by spectacle instead of depth, and the mathematical community’s own priorities (building verified libraries, formalizing foundational work, training the next generation) get displaced by whatever makes a good launch demo. It is the same structural critique now familiar from other fields: whenever a domain becomes a measuring stick for AI capability, the measuring stick gets consumed.
The Irony, and the Signal
The irony is pointed. This month also brought Claude’s 11-day machine-checked formalization of Fermat’s Last Theorem, an achievement the mathematical community largely verified and admired. The medalists are not saying the capability is bad; they are saying the incentive structure around demonstrating it is. The signal for AI labs is that the prestige-benchmark era is closing: the people whose problems are being benchmarked now have the standing to publicly object, and their objections will shape which evaluations are considered legitimate. For evaluation teams, the takeaway is to build benchmarks with domain communities rather than on top of them, or watch 25 Nobel-equivalent signatures land on a blog post about your roadmap.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
gr.Workflow Turns AI Pipelines Into Deployable APIs
Learn how to model, debug, expose, and deploy multi-step AI pipelines with Gradio Workflow and daggr.
Claude Formalized Fermat's Last Theorem in Lean After 11 Days of Autonomous Work
Anthropic announced Claude produced a complete machine-checked Lean formalization of Fermat's Last Theorem after working largely autonomously for 11 days, generating verified proofs of 30,300 intermediate theorems.
DeepMind's AlphaGenome Atlas Predicts Effects of All 9 Billion DNA Variants
Google DeepMind released the AlphaGenome Atlas, a petabyte-scale catalog predicting the regulatory effects of every possible single-letter DNA change, built on the AlphaGenome model published in Nature.
Google Research: AI Benchmarks Need 10+ Human Raters for Reliable Results
New Google Research shows that standard AI benchmarks require more than 10 raters per item to capture human nuance and ensure scientific reproducibility.
DeepMind Embeds AI Researchers at A24 in $75M Studio Deal
Google has invested $75 million in independent film studio A24 to build artist-driven pre-production tools without granting access to existing training data.