ISA After Hours Augsburg
An empty lecture hall at dawn, its blackboard wiped almost clean, a closed laptop on the lecturer's desk.

AI at work · Research

AI solved a 100-year-old mystery

· About 14 min read
In this article
  1. What the problem actually is
  2. A double-edged sword
  3. First edge: the tools we'll never invent
  4. Second edge: trusting what we can't follow
  5. Whose idea was it?
  6. The same sword, at work
  7. Sources

For about a century, mathematicians could not answer a basic question about the equations that describe everything that flows (the modern study of it began in 1934). On 8 September, OpenAI announced that one of its AI models had answered it in about 88 hours.

The model has not been released to the public. OpenAI says it ran as "on the order of 10,000" cooperating copies of itself, known as AI "agents", each working on part of the problem, and that it reached the answer "about 88 hours" after they were launched (OpenAI). OpenAI's Sébastien Bubeck has spoken for the company throughout the story (Bubeck on X).

The machine did not start from nothing. It built on a method developed over several years by two mathematicians in Madrid, Diego Córdoba and Luis Martínez-Zoroa, using "analytic techniques that don't rely on computers at all" (Quanta). In an interview with the Spanish daily El País, the two said plainly that without their idea, the AI would not have solved it (El País).

And there was a rival team. About twelve hours before OpenAI posted, Tristan Buckmaster, a professor at New York University, released related results he had built with Levent Alpöge, a mathematician who works at the AI company Anthropic and says this was a personal project (Buckmaster's statement). OpenAI admits that rumours of their work are what set its own run going: it says its effort began on 1 September after "we heard rumors that two Millennium Prize problems had been resolved" (OpenAI).

Which leaves the question the whole world is now asking. If a machine can do in a long weekend what people could not do in a century, what is left for us?

What the problem actually is

The Navier–Stokes equations are the rules physics uses for anything that flows: air around a wing, blood in an artery, a storm crossing the Atlantic. The French engineer Claude-Louis Navier wrote them down about two hundred years ago, in 1821 and 1822 (MacTutor). Engineers use them every day, on computers, and they work.

The open question was about something deeper. Can a smoothly flowing fluid, following these rules, suddenly reach infinite speed at a single point, something no real fluid could do? If it can, the equations stop describing reality at that point. Mathematicians have attacked that question since at least 1934, when Jean Leray published the first major results on it (Wikipedia), and nobody could settle it. In 2000 the Clay Mathematics Institute made it one of seven "Millennium Prize Problems", with a million dollars for a solution. OpenAI's answer is yes, it can happen. It published a 166-page paper and a version a computer can check line by line, and said it does "not intend to claim the Millennium Prize" (OpenAI).

Two caveats belong in the same breath. The proof answers a narrower version of the question than most mathematicians had in mind: it lets a gentle outside push act on the fluid, which the official problem allows but the version people dream about does not. And nobody has finished reviewing it. The Clay Institute said on 11 September that the problem "has apparently been settled" and that its evaluation is "deliberately unhurried"; its rules require "at least two years" after publication and "general acceptance" before any prize is considered, and the problem's page still lists it as active. But if it holds, Quanta Magazine wrote, it is "by a significant margin, the most important mathematical proof to have been arrived at by an artificial-intelligence model to date".

Nor was it cheap. OpenAI published no cost. Its own researchers said "several million dollars" (Quanta) and "in the millions of dollars" (Wired). Outside estimates, based on OpenAI's own figures for how much computing it used, range from $15 million (Simon Willison) to $22.5 million (TechCrunch), and Córdoba estimated about €15 million (El País).

A double-edged sword

There is no point pretending this is anything but a leap. A question that resisted the best minds for generations gave way in days. That is what AI does at its best: it throws us forward, and in science, medicine and engineering it will keep doing so.

But progress has a price, and this story shows both edges of it. The first edge is the tools we will never invent. Hard problems have always forced people to build new ideas on the way, and those ideas went on to solve other things; an instant answer skips the struggle that produced them. The second edge is blind trust. When a machine hands us a result that nobody around us fully understands, we can only take it on faith, and nobody is in a position to catch it when it is wrong.

First edge: the tools we'll never invent

The clearest warning about this edge was written five days before the announcement, when no one knew it was coming.

On 3 September Terence Tao, probably the best-known mathematician alive, posted a six-part argument on Mathstodon, a social network for mathematicians. He started by saying the Navier–Stokes question matters surprisingly little for practical engineering: an answer "would not radically transform the way we would, for instance, model weather prediction or climate change". Its value lay in what the chase produced. The attempts to solve it, he wrote, "have historically led to fundamental insights and influential theorems in fluid mechanics, analysis, and partial differential equations", the branches of mathematics these equations belong to, and he listed half a dozen of those results by name (2/6).

Then came the core of the argument (6/6):

"In most cases in pure mathematics, the problems are posed not because we desperately want the solution to these problems in and of themselves, but because we have seen from past experience that human-directed efforts to solve these problems tend to spur further development of the field through the efforts to solve such problems, and then to digest any partial or complete solutions that emerge for further insights. Prematurely solving the problem by purely AI-powered methods - particularly without full transparency into the solution process - can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole."

He was specific about where the value comes from. Progress comes from trying an approach, finding out exactly why it fails, adjusting and trying again, and the dead ends, he wrote, can "end up being highly instructive in the nature of their failure". He warned of a scenario in which an AI system does all of that internally "while the AI company running the harness keeps the process to arrive at that ansatz almost completely out of public view" (5/6). The harness is the software that runs the AI's many copies; an ansatz is the educated guess at the shape of a solution. In that case, "one of the most prominent open problems in mathematics would now be solved; but there would be almost no value added to mathematics as a consequence." Two days later he put a name on it: a "substantial opportunity cost" in turning a great problem "into a mere viral social media post advertising some benchmark progress", a benchmark being a standard test used to rank AI systems (5 Sep).

History has a well-documented case of what those by-products can be worth. For more than two thousand years, mathematicians tried to prove Euclid's parallel postulate, the assumption that through a point beside a line there is exactly one parallel line, from his other rules. Many of those attempts were accepted as proofs for long periods until the mistake was found, and in the nineteenth century the failures finally revealed that consistent geometries without that assumption exist (MacTutor, Non-Euclidean geometry). Bernhard Riemann developed that idea of curved space into a general theory in 1854, and it was in Riemann's mathematics, his biographers write, that "Einstein found the frame to fit his physical ideas" for general relativity (MacTutor, Riemann). Nobody set out to prepare the ground for relativity. It was a by-product of a long, failed attempt at something else.

What this means for you: the answer is often not what an organisation is really paying for. The understanding and the tools gained on the way to it are, and they do not arrive when the answer arrives from outside.

Second edge: trusting what we can't follow

On 11 September, Tao and 24 other winners of the Fields Medal, the field's highest honour, published a declaration titled "A Severe Misalignment of AI in Mathematics". It never names OpenAI. Its central line: "solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal." It warns that "the mass production at faster and faster pace of 'true/false' statements could destroy fertile ground instead of breathing life into new ideas", and that without people to absorb and pass on new ideas, "the crucial human transmission chain between mathematicians would be lost". A true statement that nobody understands is something you can only take on trust.

In this story, the experts' job is already shifting from finding answers to checking a machine's, and the checking is still thin.

OpenAI's result came with a version written in Lean, a programming language in which a computer checks a proof step by step, the way a spreadsheet recalculates every cell. That check tells you there is no gap in the reasoning; it cannot tell you whether the computer checked the question people actually cared about. As Quanta put it: "The crucial bit of verification that must still be done by humans is to guarantee that the statement being shown to be true in Lean is logically equivalent to what mathematicians set out to prove."

Here there is some good news, and some that is not news yet. The question the computer checked the proof against was not written by OpenAI. It comes from Formal Conjectures, an open project at Google DeepMind that writes famous unsolved problems in a form a computer can check. A contributor to that project wrote this one in May 2026, months before OpenAI's run, so the target was not tailored to the answer; OpenAI says it adapted that version. On 9 September one of the project's reviewers suggested "waiting a little bit just so experts can audit both the Lean statements and the proof". On 11 September the same reviewer reported that "an independent expert had looked at the formalisation", without naming the expert, and the people who run the project agreed to link to OpenAI's proof, one of them writing "I think we have sufficient consensus" (the project's discussion page). OpenAI's own project file still lists its review status as "self-assessed". Luis Martínez-Zoroa, whose idea the proof builds on, told El País he had so far only skimmed the 166 pages.

So the proof has a machine's approval, a question written by outsiders before the race, one project's sign-off, a look from one unnamed expert, and a proper review by mathematicians that has barely begun. Tao, for his part, said on 8 September that pushing these methods further with "an enormous amount of compute" is an exercise that "does not particularly hold my interest"; he is "far more interested in digesting the proof methods and extracting out the key new insights" (Mathstodon). That is the job that remains, and the only defence against blind trust: understanding what the machine found, not producing it.

What this means for you: the checker's job is only as good as the question it checks against, and only as good as the checker's own understanding. Somebody has to own both, and in most organisations nobody has been asked to.

Whose idea was it?

The struggle in this story did happen. It was just done by people, before the machine arrived. Córdoba and Martínez-Zoroa spent several years on a way to make fluid equations break down: stacking an endless cascade of ever-smaller swirls, each feeding the next, with techniques that need no computer at all (Quanta). "I don't use AI: I have Luis," Córdoba told Quanta. Speaking to El País, he described what the AI added as brute force, and noted, laughing, that he and Martínez-Zoroa had done the same work in a year on their ministry salaries, and enjoyed it.

The rival team's story is harder. Buckmaster and Alpöge had used AI heavily too, including OpenAI's own tools. Buckmaster says OpenAI offered to let him write up its result alone, without Alpöge, and that when he threatened to go public he was asked, "Why would you ruin your career?" (statement). OpenAI's Sébastien Bubeck replied: "I never ever asked for Levent to be removed from authorship of his own work"; in his and OpenAI's chief executive Sam Altman's telling, the offer was a rewrite of OpenAI's proof with Buckmaster as lead author, and Altman wrote that the team "acted with integrity and generosity throughout". OpenAI first said it "cannot rule out" that the pair's use of its products had helped improve its models (OpenAI on X). Two days later it said an investigation had confirmed their prompts "could not have influenced the system in any way" (OpenAI). Both accounts are on the record; which is right cannot be settled from here.

Buckmaster's own conclusion was not about OpenAI. "This is a Deep Blue-Kasparov moment," he wrote, referring to the chess computer that beat the world champion in 1997. What matters is "the way we train students, assign credit, referee, and decide what is worth one human life's attention" (statement). A few days later he told Australia's ABC that racing to publish had become "pointless": "What do we do about credit? What do we do about hiring? What do we do about PhDs?"

Credit fights are as old as the equations. As Geektime, republishing a Davidson Institute piece, recounts, Navier wrote the equations down around 1821–22 without fully understanding the physics behind them. About twenty years later the Irish-born George Stokes derived them again on a sounder footing, only to discover that Navier had got there first, and that Siméon Denis Poisson had reached similar results in between. Stokes published anyway, in 1845, and history eventually joined the Frenchman's name to the Briton's (Geektime). The difference in 2026 is that one of the parties in the race is a machine run by a company. The old etiquette, credit your sources and credit your collaborators, has no settled answer for that.

What this means for you: when an AI system produces the result, the contribution that made it possible is easy to lose. Somebody's idea, somebody's prior work, somebody's notes typed into a tool. Deciding who gets credit is now a decision someone has to make on purpose.

The same sword, at work

You do not have to care about fluids to feel both edges. Replace "open problem" with the analysis, the strategy paper, the audit or the design review your team produces, and they arrive in the same form.

Can anyone here catch it when it is wrong? If an AI produces the answer, somebody still has to understand it well enough to spot the flaw, defend it and change it. Checking is not the same as understanding: a machine can confirm that an argument has no gaps, but not that it answers the question you meant to ask. Before an AI-assisted deliverable goes out, name the person who could explain it without the tool, and who owns the question it was meant to answer. If there is nobody, you are trusting it blind.

Is the team still building what the struggle used to build? The hard version of a job is where people learn the skills that solve the next job, and the one after that. If the tool now does the hard version, that learning has to happen somewhere else on purpose: a share of work done the slow way, reviews where people have to rebuild the reasoning, juniors who work a problem before they see the machine's answer. Otherwise, in two years, the team can operate the tool and no longer do the work.

Who gets credit? This story began with a human idea from Madrid, months of human work in New York, and a rumour. Credit went first to whoever published fastest. Decide in advance how AI-assisted work is credited, and make sure the person whose idea it was is named.

One practical question for whoever looks after your data protection: when your people type unpublished work into an AI tool, what may the vendor do with it, including in "de-identified" form, with names stripped out? In this story the vendor gave two different answers in two days. That is a question for the person who owns your AI contracts, not legal advice.

About a century of human effort against about 88 hours of machine time (OpenAI): the speed is real, and it is not going away. What it costs is the path, the years of failed attempts that taught people how to think about the problem and handed them tools for the next one. Take the speed. Just decide, on purpose, which parts of the path you are still willing to walk.

Sources