Dashboard

OpenAI's Navier-Stokes Claim: What Is Settled

Three separate claims are being reported as one story, and only the first of them has an answer you can check mechanically.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
9 September 20261 min read

OpenAI's Navier-Stokes Claim: What Is Actually Settled

Three separate claims are being reported as one story, and only the first of them has an answer you can check. On 8 September 2026 OpenAI published a result on the Navier-Stokes equations, produced by roughly 10,000 agents running for 88 hours. Coverage since has merged three questions: whether a proof exists, whether it meets the Clay Mathematics Institute's prize criteria, and who got there first. Those questions have very different evidentiary status, and keeping them apart is the whole exercise.

Question one: is there a proof, and can anyone check it

This is the part with a real answer. According to Quanta Magazine's report, the agents produced a proof of finite-time blowup, meaning a solution that stops being smooth after a finite amount of time, and a second model then formalised it in Lean over a further 17 hours.

The Lean formalisation is the load-bearing detail. Lean is a proof assistant: it checks each inferential step mechanically, and it does not care who or what wrote the argument. A proof that compiles in Lean is correct in the narrow sense that its conclusion follows from its stated assumptions. That is a genuinely different category of claim from a benchmark score, which you have to take a vendor's word for. If you want the longer version of why that distinction matters, we covered it in what formal verification means in AI.

So the useful reading of question one is not "AI solved a Millennium Prize problem". It is closer to: a machine-checked artifact exists, and machine-checked artifacts are the rare kind of AI output that does not require trusting the lab that produced it.

Question two: does it meet the Clay criteria

Here the coverage disagrees with itself, and you should treat the disagreement as the current state of knowledge rather than picking the version you like.

Quanta's account describes a singularity found in the three-dimensional equations with a smooth forcing function and treats the result as satisfying the prize requirements. Forkast's analysis takes the opposite position, arguing that the Clay Mathematics Institute defines its criteria on the unforced Navier-Stokes equations while the OpenAI result addresses the forced version, which adds an external forcing term to the system.

That is not a trivial disagreement about wording. Adding a forcing term changes what you are allowed to assume, and mathematicians have long known that forced variants can be made to blow up in ways the unforced problem may not. Whether OpenAI's specific forcing function falls inside or outside the prize statement is exactly the question a referee would spend months on.

The practical implication: as of today, "OpenAI solved a Millennium Prize problem" is a contested claim, not a settled one, and the contest is technical rather than rhetorical.

Question three: who got there first

The third thread is a provenance dispute, and it is the one with no mechanical arbiter at all.

Forkast lays out a timeline in which OpenAI contacted NYU mathematician Tristan Buckmaster on 6 September, describing a proof produced by an internal model using an approach similar to work he had developed with Levent Alpöge, and proposing co-authorship that excluded Alpöge on the grounds of his Anthropic affiliation. Buckmaster and Alpöge released preprints on 7 September. OpenAI announced publicly on 8 September. Semafor reported that Buckmaster accused OpenAI of muscling in on their work and that the company denied the claim.

Note what a Lean certificate does and does not resolve. It establishes that a proof is valid. It says nothing about whether the ideas in it were arrived at independently or absorbed from work in circulation. Provenance is a social question, and the tooling that makes AI mathematics newly credible does not touch it.

Why this matters if you are building with agents, not proving theorems

The interesting transferable detail is the shape of the run: about 10,000 agents, 88 hours, then a separate verification pass by a different model. That is not one model being clever. It is a search process with a checker attached, and the checker is what makes the output worth anything.

The same structure is available at much smaller scale. An agent that writes code and then runs a test suite is doing a weaker version of the same thing: generate widely, verify mechanically, keep only what survives. The lesson from this week is not that agent swarms are magic. It is that swarms are only as trustworthy as the verifier you point them at, and most tasks people give agents have no verifier at all.

That is also the reason claims like this one should raise your standards rather than lower them. Our guide on spotting an inflated AI benchmark claim applies here in reverse: this result comes with unusually strong evidence for one of its three claims, and unusually weak evidence for the other two. Earlier OpenAI mathematics announcements had the same shape, as we covered in OpenAI Astra and the ten open math problems.

Where to keep an eye

The things that would move any of the three questions:

  • Public release of the full proof for independent review. Forkast notes OpenAI has declined to release it publicly so far.

  • A statement from the Clay Mathematics Institute on whether the result addresses the prize problem as stated.

  • Independent mathematicians working through the Lean development and confirming what its assumptions actually are.

Until at least one of those lands, the accurate summary is short: a machine-checked proof of something exists, whether it is the prize problem is disputed, and who originated the approach is unresolved. If you follow this category closely, our routine for keeping up with AI news covers how to hold a story like this open without refreshing your feed all day.

FAQ

Did OpenAI solve the Navier-Stokes Millennium Prize problem?

That is disputed. Quanta Magazine's coverage treats the result as meeting the prize requirements, while Forkast argues the proof addresses the forced equations, whereas the Clay Mathematics Institute's criteria are defined on the unforced version. No adjudication from the Clay Institute has been reported.

What does the Lean formalisation prove?

That the argument is logically valid given its assumptions. Lean mechanically checks each step. It does not verify that the problem statement matches the prize criteria, and it says nothing about who originated the ideas.

How many agents were used, and for how long?

Roughly 10,000 agents over 88 hours for the proof, with a further 17 hours for the Lean formalisation by a separate model, per Quanta's reporting.

What is the dispute with Tristan Buckmaster about?

Buckmaster, of NYU, alleges OpenAI approached him on 6 September about a proof using an approach similar to his work with Levent Alpöge, and proposed co-authorship that excluded Alpöge. He and Alpöge posted preprints on 7 September, ahead of OpenAI's 8 September announcement. OpenAI has denied the accusation.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.