Ai Engineering 2 min read

25 Fields Medalists Warn of a Severe Misalignment Between AI and Mathematics

Terence Tao joined 24 other Fields Medalists in a September 11 declaration arguing that AI companies using mathematical problem-solving as benchmarks is severely misaligned with the needs of mathematics itself.

Twenty-five Fields Medalists, including Terence Tao, published a joint declaration on September 11 titled A Severe Misalignment of AI in Mathematics, and the argument deserves more attention than the headline. The signatories are not warning that AI will destroy mathematics. They are warning that AI companies are using mathematical problem-solving as a benchmark for progress in ways that misalign incentives with what mathematics actually needs, a critique published days after Anthropic’s Claude formalized Fermat’s Last Theorem in Lean.

The Complaint Underneath the Benchmark Fight

The declaration’s core concern is what optimization does to a discipline. When labs race to solve famous open problems as demonstration exercises, the problems become benchmark fuel rather than research objects: solutions get burned as evaluation targets, progress gets measured by spectacle instead of depth, and the mathematical community’s own priorities (building verified libraries, formalizing foundational work, training the next generation) get displaced by whatever makes a good launch demo. It is the same structural critique now familiar from other fields: whenever a domain becomes a measuring stick for AI capability, the measuring stick gets consumed.

The Irony, and the Signal

The irony is pointed. This month also brought Claude’s 11-day machine-checked formalization of Fermat’s Last Theorem, an achievement the mathematical community largely verified and admired. The medalists are not saying the capability is bad; they are saying the incentive structure around demonstrating it is. The signal for AI labs is that the prestige-benchmark era is closing: the people whose problems are being benchmarked now have the standing to publicly object, and their objections will shape which evaluations are considered legitimate. For evaluation teams, the takeaway is to build benchmarks with domain communities rather than on top of them, or watch 25 Nobel-equivalent signatures land on a blog post about your roadmap.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading