Illustration for: Claude Just Formalized Fermat's Last Theorem in 11 Days

Claude Just Formalized Fermat's Last Theorem in 11 Days

Anthropic's Claude produced the first complete, computer-checked proof of Fermat's Last Theorem in Lean, working largely autonomously over 11 days across 13 million lines of code.

By the Numbers

11 days
Time to complete
13M lines
Lean code generated
29,500
New theorems proved
~6B
Output tokens used
TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

It's a rare capability demo with a verifiable pass/fail outcome, not a benchmark score a lab chose to highlight.

2

The breakthrough came from open-source scaffolding (Prove2Me), not the base model alone -- tooling is doing real capability work right now.

3

It resets the moat question for any 'AI for science' startup pitching fine-tuned models as their edge.

TC

The VC Read · Trace's Take

Trace Cohen

The diligence question I'm now asking every 'AI for science' founder: what does your model do that pointing Claude or GPT-6 at the right open-source scaffolding wouldn't also do? Prove2Me existing as a free, open tool is what actually unlocked this, not a proprietary Anthropic breakthrough -- which means the moat for math- and science-focused AI startups just got thinner, not thicker.

Analysis

I don't think most of the AI industry has clocked what Anthropic just did, because it doesn't look like a product launch. Claude spent 11 days largely autonomously producing the first complete, computer-checked proof of Fermat's Last Theorem in the Lean 4 programming language -- 13 million lines of Lean code, 29,500 new formal theorems, roughly six billion output tokens -- using an open collaborative platform called Prove2Me built by Tianyi Peng and collaborators at Columbia. Mathematician Kevin Buzzard, who has spent years working on Lean formalization himself, called it an achievement that arrived far faster than experts predicted.

Here's why I'd put this above most of this week's funding headlines. Formalizing an existing proof in Lean isn't the same as discovering new mathematics, but it's a much better test of sustained, multi-step autonomous reasoning than any benchmark I've seen used to justify a funding round: 11 days of largely unsupervised work, self-correcting across millions of lines, with a verifiable pass/fail outcome (Lean either accepts the proof or it doesn't). Anthropic's own account notes its first attempt at formalizing Wiles' proof failed outright -- the breakthrough only came once Claude had access to Prove2Me's scaffolding. That's the real signal: capability gains right now are coming as much from tooling and infrastructure around the model as from the model itself.

The gap between checkable translation and genuine mathematical discovery is exactly where I'd expect the next 18 months of hype to outrun the actual capability.

For anyone diligencing an "AI for science" or "AI for math" startup, this changes the bar. If a frontier lab's general-purpose model, pointed at the right open-source scaffolding, can do this without months of task-specific fine-tuning, a startup's pitch needs to explain what moat exists beyond access to Claude or GPT-6 and a well-built tool. Formal verification, drug discovery, and materials science all share this structure -- an expensive, checkable ground truth plus enough compute to brute-force the search -- and enough of that stack is now assembled from off-the-shelf pieces that "we fine-tuned a model for X" is a much weaker answer than it was a year ago.

Room for disagreement: this is one demonstration on one famous, well-studied theorem, not evidence of general mathematical creativity. Fermat's Last Theorem already has a validated human proof (Wiles, 1994) for Claude to formalize against -- it's a translation task with a known correct answer, not open research. Skeptics are right that "formalized an existing proof" and "discovered a new one" are very different capabilities, and Anthropic hasn't claimed the latter. The gap between checkable translation and genuine mathematical discovery is exactly where I'd expect the next 18 months of hype to outrun the actual capability.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by Anthropic · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.