Overview
- Anthropic's Claude AI has achieved a significant milestone by producing the first fully computer-verified proof of Fermat's Last Theorem in just 11 days, generating what is now the longest mathematical proof ever created.
- In contrast, a human-led initiative at Imperial College London, which aims to accomplish the same task, has been ongoing since 2024 and is far from completion. Claude's accomplishment surpasses this effort.
- Kevin Buzzard, the mathematician overseeing the human project, has evaluated Claude's proof and affirmed its validity based solely on fundamental mathematical principles.
Anthropic has announced that its Claude AI successfully generated the longest mathematical proof in history, formally establishing Fermat's Last Theorem, a conundrum that has perplexed mathematicians for 358 years.
In a remarkable feat, Claude completed this task in 11 days, primarily autonomously, producing 13 million lines of code that can be verified by a computer, rather than relying on a mathematician's confirmation.
Fermat’s Last Theorem posits that no three positive integers can be found such that when each is raised to a power greater than 2, the first two will sum to the third. This assertion was initially noted in the margins of a mathematics book by Pierre de Fermat in 1637, who claimed to have a "truly marvelous proof" that was too lengthy to fit in the margin.
After Fermat's death, mathematicians dedicated the subsequent 358 years to reconstructing his presumed proof.
Understanding Proofs and Verification
A mathematical proof consists of a sequence of logical deductions, and if any single link is flawed, the entire proof fails. Identifying such flaws, buried within extensive arguments, can consume years of a mathematician's time.
Formalizing a proof involves converting it into a language that a computer can independently verify without subjective interpretation.
Historically, mathematicians have struggled with this verification. A 1908 German prize, worth approximately $1 million to $2 million today, for the first valid proof of Fermat's theorem attracted 621 incorrect submissions in its inaugural year alone.
Verifying a major mathematical proof can take years. Formalization—converting the mathematical reasoning into a format that computer proof assistants like Lean can validate—can facilitate this process.
Last month, Claude achieved the first formalized proof of Fermat’s Last Theorem, one of… pic.twitter.com/pdT8zwlV4A
— Anthropic (@AnthropicAI) September 4, 2026
The definitive proof of Fermat's Last Theorem was not established until 1995 by British mathematician Andrew Wiles, who faced challenges while presenting his solution across three lectures in June 1993, only to have a reviewer later identify a flaw.
Wiles, along with his former student Richard Taylor, spent almost a year rectifying the proof, nearly giving up before finally publishing a corrected 129-page proof in May 1995. This proof utilized mathematical theories that did not exist during Fermat's lifetime, resulting in skepticism about Fermat's original "marvelous proof."
In 2024, Kevin Buzzard launched a project at Imperial College London to replicate what Claude has accomplished: translating Wiles's proof into Lean, a computer-verifiable language. This effort requires a substantial volunteer mathematician base and is projected to continue through 2029.
Claude completed this entire task in just 11 days.
The Mechanism Behind Claude's Success
According to Anthropic, Tianyi Peng, who develops AI formalization tools at Columbia, sought to determine how far Claude could progress independently. A team of Claude agents worked collaboratively, generating definitions, proving smaller results, and integrating them into larger proofs, with minimal human guidance such as prioritizing certain theorems.
Initially, the process faced obstacles, as the agents struggled to track their progress and collaborate effectively, contributing to approximately 7% of the final proof consisting of false starts.
The situation improved with the introduction of a tool called Prove2Me, developed by Peng's team, which provided all agents with a shared live to-do list of pending smaller proofs, preventing duplication of efforts and ensuring efficient organization for faster verification by Lean. It also maintained clear notes on each result to facilitate reuse among agents.
In total, Claude proved over 30,000 supporting theorems and processed billions of tokens while operating on a research model comparable to Claude Fable 5.1, which was subsequently released to the public. The completed proof encompasses 13 million lines, exceeding five times the size of Mathlib, the existing shared library for similar mathematical work.
To put this in perspective, a typical novel contains around 80,000 words; Claude's proof equates to 160 novels filled with rigorous logical reasoning.
Is This Significant?
Buzzard, whose own project continues to receive funding until 2029, has reviewed Claude's proof and endorsed it, asserting that it proves the theorem "with no assumptions other than the axioms of mathematics."
However, this does not imply that Claude has unearthed new mathematical concepts. Wiles had already proven Fermat's theorem three decades prior—Claude merely created a machine-checkable record of it. This is crucial as mathematicians are increasingly inundated with unverified proofs, including those generated by AI, faster than they can conduct manual verifications.
Moreover, these formal proofs are deterministic and less susceptible to human error, a vital aspect of mathematical rigor.
This issue is not new. A computer-assisted proof of the Kepler conjecture took four years before a review panel could only assert "99% certainty," while Grigori Perelman's proof of the Poincaré conjecture required a similar duration to be fully accepted.
For those skeptical of Anthropic's claims, the entire 13-million-line proof is available on GitHub, accessible for any mathematician willing to dissect it thoroughly.
