Analysis
Mathematicians working with Meta's Muse Spark model co-authored six research papers, five of which answer previously unsolved open problems, according to a summary of the results. The team used Muse Spark 1.1 and 1.2 in Thinking Mode through the standard meta.ai chat interface, with no custom research scaffolding built around the model.
The six papers span probability, differential equations, group theory, optimization, arithmetic physics and non-associative algebra, according to a list posted by Scale AI's Alexandr Wang -- including a sharp threshold for fitting random Gaussian points to ellipsoids and a resolution to finite-time blow-up for the mass-critical biharmonic nonlinear Schrodinger equation, a problem that had been open since 2015, plus a counterexample showing semiabelian groups need not be monomial.
The result adds Meta to a research-benchmark conversation OpenAI and Google DeepMind have both leaned on to demonstrate frontier reasoning progress, intensifying competition over whose model is advancing genuine mathematical reasoning versus pattern-matching familiar techniques from training data. That competition for research-credibility benchmarks is increasingly a marketing and recruiting tool for labs trying to attract both enterprise customers and top research talent, alongside more traditional product benchmarks like coding and computer-use scores.
Mathematical open-problem-solving is a notably harder benchmark to game than standardized test scores, since a proof either holds up to peer review or it doesn't -- which is part of why labs increasingly reach for this kind of result over benchmark leaderboards that critics argue can be optimized for directly during training. Whether Muse Spark's contribution constitutes genuine novel reasoning or sophisticated synthesis of techniques already scattered across the mathematical literature is itself a live debate among researchers reacting to the papers.
What the results don't settle: the papers were produced in Thinking Mode with no disclosed time or compute budget per problem, so it's unclear how reproducible or scalable the approach is beyond these six specific problems. A general-purpose chat model contributing to real open-problem solutions, rather than a specialized research system built for the task, is the more interesting signal for AI-for-science investors than the papers themselves -- it suggests the capability may already be more broadly accessible than a dedicated research tool's narrower user base would imply.

