On September 1, OpenAI launched a batch of agents. On the order of ten thousand of them, running at once, split into groups, each group working a different variant of the Navier–Stokes problem.
Eighty-eight hours later, on September 5, they returned a result: a fluid starting at rest, pushed by a smooth external force, winds into a vortex that spirals inward while stretching along its axis, and its speed goes to infinity in finite time. The equations blow up. Seventeen hours after that, the proof had been translated into Lean and checked by machine, line by line. According to press accounts, this one problem consumed about 130 billion output tokens.
OpenAI announced it at noon on September 8. Navier–Stokes is one of the seven Millennium Prize Problems the Clay Mathematics Institute posted in 2000, each carrying a million-dollar prize. If the result survives peer review, it will be the second of the seven ever solved. The first was the Poincaré conjecture.
About twelve hours before the announcement, NYU mathematician Tristan Buckmaster put out a statement of his own. All summer, Buckmaster and Levent Alpöge — an Anthropic employee, working in a personal capacity — had been using AI on a simpler cousin of the problem, the Euler equations with an external force, following the "forcing" approach developed by Diego Córdoba and collaborators. They had a result in mid-August and had it verified in Lean before the end of the month. Buckmaster's account is that OpenAI learned of their progress days earlier and used the same route.
OpenAI's own write-up tells a different story. On September 1, it heard rumors that two Millennium Problems had been solved, and launched a project to throw agents at all the unsolved ones. Along the way, roughly a hundred agents working for about fifty hours solved an "easier" problem: the Euler equations blow up even without an external force. Seeing that, OpenAI judged Navier–Stokes the most promising target, moved agents over from the other problems, and used the Euler solution as the prompt. It says it never saw Buckmaster and Alpöge's work before they made it public, credits them with priority on the forced Euler result, and says the two proofs differ substantially.
A few days before all this, on September 4, Anthropic had published a result of its own. In eleven days, Claude produced a complete formalization of Fermat's Last Theorem in Lean: 13 million lines of code, more than five times the size of Mathlib, the community's entire library of formalized mathematics, with over thirty thousand intermediate theorems proved along the way. This was a community project Kevin Buzzard of Imperial College London kicked off in 2024, expected to take years. The mathematical input from humans was limited to occasional high-level instructions from the project lead.
On September 11, twenty-five Fields Medalists published a joint declaration titled "A Severe Misalignment of AI in Mathematics." One sentence in it: "solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight."
Four lanes over five weeks. Machines did the search; people decided where to aim. Sources: Anthropic's two research posts; OpenAI's announcement (September 8, updated September 10); Buckmaster and Alpöge's August dates and the "about twelve hours earlier" via the Wikipedia article on the priority controversy; the declaration text on Terence Tao's blog. The hour the OpenAI run started is not public.
For the rest of September, the question in that world was some version of: is math over?
My answer: math isn't over. What's over is the part of math that has a fixed finish line and a cheap judge. And this re-sorting isn't only happening in math. It's happening in every trade at once, including the shop floor.
1. Who did what in those 88 hours
Take the two stories apart and ask what the machines did and what the people did.
The machines did the heavy lifting. Ten thousand agents, over a hundred billion tokens, eighty-eight hours of trying without a break. Thirteen million lines of Lean in eleven days. Search at that scale is more than the entire community of human mathematicians could do in decades.
The people did two other things.
First, they set the finish line. The statement of the Navier–Stokes problem was fixed by the Clay Institute in 2000, every word precise. The proof of Fermat's Last Theorem was given by Andrew Wiles three decades ago; formalizing it means translating something already known. Before either job started, what "done" looked like was perfectly clear.
Second, they picked the direction. Where the technical stepping stone came from is disputed: Buckmaster says the route was theirs first; OpenAI says its own agents found the unforced Euler blowup. Peer review will have to settle that. But one thing isn't in dispute on either side: people decided where to aim. The project started because of a rumor. Putting ten thousand agents on Navier–Stokes was OpenAI's call after looking at the Euler result — its write-up says it concluded Navier–Stokes was the most promising place for a breakthrough. And the stepping stone itself — do the Euler equations blow up? — was also a problem with a pre-written statement and a Lean check.
There's one more role in both stories, and in neither is it played by a person: the judge. Lean is a proof checker. It goes line by line; right is right and wrong is wrong. Each ruling costs almost nothing.
A fixed finish line and a cheap judge. Once both are in place, what remains is heavy lifting — and heavy lifting is exactly what machines are best at right now.
In that sense — and only in that sense — the Millennium Problem was the easy part. The hard parts had been done before the first agent launched: someone wrote the statement, someone built the checker, someone decided where to aim.
2. Two kinds of work: capped and open
Generalize that, and nearly all work sorts roughly into two kinds.
Capped work: you know in advance what "done" looks like, and there's a cheap judge that can tell you whether you got there. A school essay has a rubric. A support ticket is resolved or it isn't. 3D part localization is off by some number of millimeters. Formalizing Fermat has Lean. Even Navier–Stokes, once its statement is fixed and Lean can check the proof, falls on this side.
Open work: you have to define the finish line yourself. Which question to ask; which of dozens of routes is worth walking; what counts as "understanding"; what the tolerance on a part should be, and what "clean" actually means. None of these come with a judge. Put differently: whoever does the work is the judge.
AI acts on these two kinds of work differently.
On capped work, AI is addition. It pushes everyone some distance toward the ceiling, and because the ceiling is fixed, the people farthest from it move the most. Economics has a clean piece of evidence here. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,179 customer-support agents given a generative-AI assistant. Issues resolved per hour rose 14% on average and 34% for novice and low-skilled workers, with minimal effect on the most experienced. The gap compressed.
Support work has a finish line and a judge, so AI acts as addition: the people farthest from the ceiling gain the most. Data: Brynjolfsson, Li & Raymond, "Generative AI at Work," Quarterly Journal of Economics, 2025. "Minimal" is the paper's word; the bar is drawn near zero.
On open work, AI is multiplication. There's no ceiling, so what AI amplifies is the judgment you already have. The larger the base, the larger the absolute gain.
I find it easiest to think about this with basketball. This is an analogy, not a statistic. Over the past decade, NBA role players have been pulled close to the stars: everyone shoots threes, everyone switches on defense, the floor of basic skill has risen across the league. Yet the stars have pulled further away. Part of the game is capped — shooting form, defensive positioning; once you've drilled it, you've drilled it. Part of it is open — reading the game, deciding who gets the ball in the last two minutes. The first part gets flattened. The second gets stretched.
The economist Sherwin Rosen described the same mechanism in 1981 in "The Economics of Superstars": when technology widens one person's reach, small differences in talent turn into large differences in reward.
Mathematics supplied a direct example this year. Mehtaab Sawhney, an assistant professor at Columbia, took leave to join OpenAI, and with colleagues has used internal models to settle a string of Erdős problems. When Sawhney joined, OpenAI's chief research officer Mark Chen wrote that "accelerating Mehtaab will also meaningfully accelerate mathematics."
Nobody says that about a support agent.
Read the Fields Medalists' declaration again and you'll see they drew the same line I'm drawing. Solving problems is the proxy; understanding is the point. Solving is capped. Understanding is open. AI is eating the first.
3. Where the stars are scarce
If the top gets stretched, what exactly is scarce about the people at the top?
Not the ability to prove things. September showed that as long as the statement is fixed and the judge is present, a proof can be ground out by ten thousand agents.
The scarcity is in two things: picking the direction and being the judge.
Picking the direction is research taste. In September, direction got a very concrete price tag: a rumor that two Millennium Problems had fallen — no names, no methods — was enough for a company to launch a project that ended with ten thousand agents on a single problem. The bare signal that a direction can be walked was worth that much.
Being the judge is subtler. In capped work, the judge comes ready-made — Lean, a rubric, a caliper. In open work there is no ready-made judge, so the person who can say "this result is right, this one matters, this one is worth pursuing" becomes the scarcest part of the whole system.
It's the same structure as my August piece, "The Builder Cannot Be the Verifier." As generation gets cheaper, verification gets more valuable. That piece was about robot programs. This one is about people: once AI makes generation nearly free, human value concentrates at the judgment end.
There's one more layer. In a guest post on Terence Tao's blog, the philosophers Silvia De Toffoli and Eamon Duede put it plainly: the machine delivered an answer, a conclusion checked to be true; mathematicians want a solution, which includes why it's true. They quote the Clay Institute's own explanation of why a proof matters for Navier–Stokes: "Because a proof gives not only certitude, but also understanding." A hundred and thirty billion tokens bought certainty. Understanding is a separate bill. NPR's headline on September 22 was blunter: "AI solved one of math's hardest problems. Humanity learned nothing (so far)."
4. Where the basketball analogy breaks: the ball
At this point I have to admit the analogy has a big hole in it.
In the NBA, the stars and the bench play with the same ball. At the frontier of mathematics, they don't.
Press estimates put the compute for OpenAI's push in the millions of dollars. An independent mathematician without access to that kind of compute, however good their taste, cannot run a route to the end in eighty-eight hours. Meanwhile a researcher with ordinary taste who happens to sit inside a lab might get there first.
That's precisely MIT Technology Review's worry: that mathematics becomes the private preserve of well-capitalized AI companies. The declaration's complaints — results announced in a rush, no time for a proper write-up, prior work left uncited, "severe attribution and plagiarism questions" — connect to the same thing. Access to compute is concentrated in a few companies, and the scarcity may be captured by institutions instead of staying with people.
So the more accurate statement is: scarcity at the top is person × access — the person, times a channel to the compute.
Countries are building that channel differently. In the US, in September, it looked company-driven: OpenAI launched its push on a rumor that someone else had gotten there first, and the whole thing was a race.
China is on a different road.
Two things are on the public record. The State Council's "AI+" action plan, issued in August 2025, lists six priority actions, and the first is "AI + science and technology." In late August this year, Beijing's Zhongguancun hosted a national AI-for-Science conference chaired by fifteen members of the Chinese Academies, with the stated goal of moving research from "people using AI to compute" to "AI conducting research autonomously." And from where I sit, I see more AI-for-Science conferences being convened, and, in manufacturing, companies closing technical gaps quickly.
On one side, a handful of companies build the channel with capital. On the other, the state lays it down by plan. Which road grows more person × access combinations is not visible yet. But the question is worth writing down. It will be one of the most important questions on the US–China line for the next several years.
5. The frontier's other direction: down
Everything so far has been about moving up — essay, code, research, top research. Each step up, the finish line gets blurrier and AI finds it harder to flatten.
But the frontier doesn't only run up. It also runs down.
Down is the physical world: where the part is sitting, where the burr is, whether there's porosity inside the weld. The difficulty there isn't abstraction. It's that you have to touch the thing to know.
The same re-sorting is happening along this line too.
In my piece on Cognex and RealSense, I wrote about Inbolt, a Paris startup I kept running into at IMTS. A standard 3D camera costing a few hundred dollars, plus Inbolt's software, does bin picking and part localization on a robot arm. By the company's own numbers, it has more than 200 systems installed in more than 75 factories, with 0.5–0.7 mm repeatability on a standard camera. Three days after the show closed, Cognex announced it would buy RealSense, the company behind the camera in Inbolt's public demos, for about $500 million in cash.
3D localization got good because it's capped. The part's true pose is right there; millimeters of error make a cheap judge. A finish line plus a judge means it can be flattened. And the result of flattening is that it stops being a craft and becomes a module you can buy.
So where did the value go? To the next place that doesn't have a cheap judge yet.
Deburring is one. After the robot finishes, in most shops "is it clean?" is still answered by a veteran running a gloved hand over the part. Welds are another: you can photograph the bead, but porosity inside takes NDT, which is slow and expensive. Sub-millimeter precision assembly is worse still: what you can measure is often not the quantity you actually care about.
What these processes share is that the finish line hasn't been written down and the judge hasn't been built.
In "Engineering the Map" I argued that where a process sits on the automation map isn't fixed: hang an independent re-inspection behind it, and the cost of verification drops from a veteran's glove to a single image. In today's framing, what that means becomes clearer. Building a judge drags a process from the open side to the capped side. Once it's across, it gets flattened, and the value moves on to the next place without a judge.
That road doesn't end. Every judge you build moves the frontier down one more notch.
Two axes. Horizontal: is "done" defined in advance? Vertical: what does one judgment cost? Bottom-left is capped; top-right is open. Even a Millennium Problem, once the statement is fixed and Lean can check it, sits bottom-left. The arrow is Section 5: hang an independent camera re-check on deburring, and it gets dragged from the middle to the capped side. Positions are qualitative — my judgment, no units.
6. Three things for people who make things
First, ask two questions about the work in front of you. Do you know in advance what "done" looks like? Who judges it, and what does one judgment cost? If both have clear answers, that work is on the side being flattened. Hand it to AI early and move your people off it; don't staff up on it.
Second, the most valuable seat is the person who builds the judge. The one who sets the tolerance. The one who designs the independent re-check. The one who turns "a veteran runs a glove over it" into a judgment you can record and replay. In mathematics this is called research taste; on the shop floor it's called process knowledge. It's open work, so it gets stretched.
Third, access doesn't only live in data centers. In mathematics, access means ten thousand agents, and most people can't get it. In manufacturing, much of the access is physical: cameras, probes, NDT equipment, the line itself. The people standing on the line already hold part of the channel. It's one of the few structural advantages people who make things have in this round of re-sorting. Don't give it away.
One falsifiable marker, for the record: if, in the next twelve months, an AI poses a new problem or a new theory — with no human having supplied the statement or the route — and the mathematical community accepts it as important, then this piece's claim that the open side belongs to people needs revising.
If you know a grad student who spent September wondering whether their field is "over" — send them this.






The “build the judge” point stuck with me. An AI-generated answer might look good in one run, but would it still hold up when materials change, a line starts up, or something goes wrong? That’s the gap between a promising result and something a plant can rely on.