Analysis
OpenAI announced on Saturday that an internal version of its next major model, code-named Astra, had solved ten previously-open problems spanning mathematics and theoretical computer science, publishing machine-checkable Lean proofs on GitHub alongside a roughly 249-page report detailing the methodology. The results include a construction proving the existence of non-sofic groups -- a question that had stood open in group theory for years -- and new bounds on sphere-packing, a classical problem with a long history of only incremental progress.
What separates this from prior AI math-solving demonstrations is the endorsement, not just the output. Fields Medal winner Timothy Gowers, one of the most respected working mathematicians alive, said he would recommend one of the Astra-generated proofs for publication in the Annals of Mathematics, widely regarded as one of the field's top journals, without hesitation. That's a meaningfully different bar than a model performing well on a benchmark of known-answer competition problems -- these were genuinely open questions, and the model's proofs are being evaluated on the same terms a human mathematician's work would be.
“The compute cost is the detail most likely to reshape how labs and investors think about AI-driven research economics: roughly $2,000 total across all ten results.”
The compute cost is the detail most likely to reshape how labs and investors think about AI-driven research economics: roughly $2,000 total across all ten results. For comparison, DeepMind's AlphaProof and AlphaGeometry efforts, and OpenAI's own earlier competition-math systems, have typically required far more extensive compute and human-curated training pipelines to reach far narrower results. If $2,000 genuinely bought ten previously-open results at this level, the marginal cost of AI-assisted mathematical discovery has fallen by an order of magnitude or more within a single model generation.
For AI investors, the signal isn't that a model can do math -- reasoning models have chased competition-math benchmarks for two years -- it's that verified, novel research contributions are now arriving at a cost structure closer to a cloud compute bill than a research grant. That reframes what 'AI for science' companies need to prove to justify venture-scale valuations: the bottleneck shifts from raw model capability toward domain-specific problem framing and verification infrastructure, since Lean's machine-checkable proof format is what let Gowers evaluate the results quickly and confidently in the first place.
What to watch: whether Astra's full public release reproduces these results at similar cost once it's generally available rather than an internal build, whether other frontier labs claim comparable open-problem results in adjacent fields, and whether Gowers' proof actually clears peer review at the Annals -- a genuine publication would be a categorically different milestone than an informal endorsement, however credible the endorser.