OpenAI says an internal version of Astra resolved or substantially advanced ten long-standing problems across geometry, coding theory, complexity, group theory, cryptography, and combinatorics.
That is an extraordinary claim, and extraordinary claims in mathematics have a useful property: they do not become true because a model sounds confident.
They have to survive proof.
OpenAI says humans prepared the arguments into manuscripts with the model, after which Astra formalized each argument into a Lean certificate. The company also published reasoning walkthroughs and made an unusually clear attribution choice, stating that presenting an AI-generated argument as human work would misrepresent how the result was produced.
The headlines will focus on ten problems. The more durable story is the verification stack around them.
Discovery is becoming cheaper than trust
Generative AI can create candidate ideas much faster than a mathematical community can inspect them. If research models continue improving, the scarce resource will not be the first plausible proof. It will be expert attention, formal verification, reproducibility, attribution, and the slow work of placing a result inside an existing field.
Lean matters because it changes part of that trust problem from persuasion into machine-checkable structure. A formal certificate does not decide whether a theorem is important or whether an argument teaches mathematicians something useful, but it can verify that the encoded proof follows from its stated assumptions.
That distinction is critical. The model can produce an elegant-looking mistake, while a formal system demands exact steps. Human experts still have to confirm that the formal statement matches the intended problem, that the assumptions are appropriate, and that the result belongs in the literature as claimed.
OpenAI estimates that the model work behind the ten results would have cost roughly $2,000 at GPT-5.6 Sol API rates. Even if that estimate excludes substantial human review and infrastructure, it points toward a dramatic shift in research economics.
When generating a candidate proof becomes cheap, verification becomes the workbench.
The next serious scientific AI systems will not be judged only by what they discover. They will be judged by how clearly they expose provenance, uncertainty, formal evidence, expert review, and responsibility when the answer is wrong.
Source: OpenAI