26 years. 88 hours. $1 million. 🧮 For 26 years, only one of the seven Millennium Prize problems had ever been cracked. This week, OpenAI says an AI may have taken down a second.
- 🌊 The target: Navier-Stokes — the 90-year-old equations that describe how every fluid moves. Weather forecasts, airplane wings, blood-flow simulations… all of it runs on this math. ✈️
- 🤖 They didn't use anything public. They pointed an unreleased model — described as significantly more capable than GPT-6 Astra — at the problem, then unleashed 10,000 AI agents on it at once.
- ⏱ 88 hours. 2.7 million messages exchanged. One proof.
- ⚖️ The catch: it still has to survive two years of review by the math community before it qualifies. And OpenAI says it won't even claim the million dollars.
- 🤯 But if it holds? AI just solved the problem generations of the world's best mathematicians spent entire careers chasing — and never cracked. Wild time to be alive.
Context & limitations
The numbers all hold up. OpenAI's own write-up says roughly 10,000 concurrent agents, a result 88 hours after the run started, 2.7 million messages and about 130 billion output tokens on this problem alone — 4.9 million messages and 300 billion tokens across everything they threw at it. GPT-6 Astra then spent another 17 hours formalising the argument in Lean, which is the part that matters: Lean doesn't have opinions, so a machine can check a machine. Mark Chen put the compute bill in the millions of dollars.
The word doing the quiet work is "forced". The 166-page proof builds a fluid that blows up in finite time while something external keeps pushing it. Half the internet says that means it doesn't count — and that's wrong on the paperwork. Charles Fefferman's official problem statement asks for "a proof of one of the following four statements", and (C) and (D) are precisely the forced-breakdown cases, as long as the force is smooth and decays properly. So it's inside the rules as written. But the version mathematicians actually lose sleep over is (A) and (B), where nobody is pushing and the fluid tears itself apart on its own. That one is still open. Both of those are true at the same time, and almost every take you'll read picks one and drops the other.
The one number in the caption I'd change is the 90 years. The equations are about two centuries old — Navier in 1822, Stokes in 1845. What's roughly 92 years old is Leray's 1934 paper, the one that framed the existence question everybody's been stuck on since. And the Millennium Prize problem itself only turns 26 this year. "90-year-old equations" makes the maths younger than it is.
The two-year wait is real and it's a rule, not a vibe: Clay requires publication in a qualifying outlet, at least two years elapsed, and general acceptance in the global mathematics community. Clay president Martin Bridson called the announcement exciting and said the evaluation would be "deliberately unhurried". The institute still lists Navier-Stokes as unsolved. The Lean files are public, but the review status on them is self-assessed — as of now no outside group has re-run the check.
And the part nobody posts is who was already standing there. Diego Córdoba and Luis Martínez-Zoroa spent about a year building the forcing technique both proofs rest on — Buckmaster says Martínez-Zoroa "deserves a Fields Medal". Levent Alpöge (Anthropic) and Tristan Buckmaster (NYU) proved forced-Euler blowup on 15 August and had it Lean-verified by 22 August. OpenAI's run started on 1 September, after hearing a rumour of exactly that. So the 88 hours is honest, and it's 88 hours standing on a year of someone else's work. Buckmaster has since alleged that OpenAI offered him co-authorship while excluding Alpöge over his Anthropic affiliation, and that he was asked "Why would you ruin your career?"; OpenAI's Sébastien Bubeck calls the allegations "false and inflammatory". OpenAI's own post contains this sentence: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." Terence Tao's warning is the one worth keeping — that efforts on this scale, triggered by a hint of what another group is doing, could stop researchers sharing what they're working on at all.
Do this yourself
Verification makes iteration useful only within the checks' scope; attempts still consume time and resources.















































