Analysis
The detail that should stick with anyone evaluating AI research investments isn't that a model solved hard math problems -- it's the price tag. An internal version of OpenAI's Astra generated formal, machine-checkable proofs for ten previously open problems in mathematics and theoretical computer science, including the first explicit construction of a non-sofic group, a question open since Mikhail Gromov introduced the concept in 1999. Total compute cost: roughly $2,000.
Rigorous, Not Just Impressive
The validation is unusually rigorous for an AI research claim. Fields Medalist Tim Gowers said he would recommend the proof for publication in a top mathematics journal without hesitation, and a team of nine mathematicians including Gowers and Noga Alon later published a companion paper explaining the result for human readers. The formal Lean 4 proof certificates, released publicly on GitHub, carry a zero "sorry" count -- meaning every logical step across all ten proofs is machine-verified, not just plausible-looking.
“For anyone pricing AI research capability, that $2,000 figure reframes what "research output" costs relative to human capital.”
For anyone pricing AI research capability, that $2,000 figure reframes what "research output" costs relative to human capital. Work that would previously have required years of a specialized mathematician's time -- and carries real career and citation value once published -- came from a few thousand dollars of inference. That doesn't make human mathematicians obsolete, but it does compress the cost curve on a category of intellectual labor that felt immune to automation longer than most.
The implication for AI lab valuations and research funding is direct: if verified, publishable-grade research output scales this cheaply, the return on frontier model R&D spend looks a lot better than pure product-revenue multiples alone would suggest, and it strengthens the case for labs continuing to burn enormous sums on model scale rather than optimizing purely for near-term commercial applications.
What to watch: whether Astra's full public release reproduces this result reliably across a wider set of open problems, or whether this specific result reflects favorable problem selection rather than a repeatable research capability.