Ten thousand agents and one PhD thesis
OpenAI cracked Navier-Stokes with 10,000 agents in 88 hours. The direction came from a 2021 thesis written without a computer. Compute scales the search. Taste still picks it. Here is how to extract taste to steer an autoresearch loop, and three UX to try.
OpenAI put roughly 10,000 agents on the Navier-Stokes problem for 88 hours, burned something like 130 billion output tokens, and came out with a singularity. [1] The strategy those agents followed came from a 2021 doctoral thesis that, by its author’s own account, avoided computers entirely. [2]
That is the whole post, really. The rest is me arguing that this is not a fun coincidence but the shape of every autoresearch loop you will run this year.
What the compute did and did not do
The story as it stands on September 12. Diego Córdoba and Luis Martínez-Zoroa spent years building an “infinite cascade” of non-singular layers that stack into a singular solution. By 2023 they had it working for Euler variants. The remaining gap was that stacking the layers wrecked the smoothness of the forcing term, which is exactly the thing the Millennium Prize demands. [2]
Tristan Buckmaster and Levent Alpöge picked up that approach and spent most of a year grinding on it with public models. They got a Lean-verified Euler result by August 22. Then word got out. [3]
OpenAI, on a call with Buckmaster on September 6, said it had only tried the problem in the past week, after the rumors started. [3] It used the same approach. First about 100 agents for 50 hours on an Euler variant, then the 10,000-agent run on full Navier-Stokes. [2] Estimates of the bill run from one million dollars at internal rates to over twenty at list price. [4]
Charles Fefferman’s read: “The heroes of the story are Córdoba and Martínez-Zoroa.” [2] Córdoba’s read: “I don’t use AI: I have Luis.”
So the machine did the thing machines do. It searched a space that a human had already pointed at, faster and wider than any human could. What it did not do was choose the space. Nobody at OpenAI woke up on September 1 with a new idea about fluid singularities. They woke up with a rumor about which old idea was about to work.
Taste is the part that did not scale
Karthik Duraisamy, after costing out the run, added a line I keep coming back to: “there must have had a good deal of (human) intuition that was involved. We’d never know.” [4]
We would never know because nobody logs it. The compute has a receipt. The taste that selected the direction was a chain of judgment calls made in offices in Madrid and New York over five years, and none of it was written down as an instruction the agents could read.
MIT Technology Review put it plainly: research taste, meaning the ability to pick the promising question, has long been the named obstacle for AI in math, and if the agents took the Córdoba and Martínez-Zoroa route because humans had shown it was promising, then human taste “played an essential role in OpenAI’s success.” [5]
The DeepMind and Gómez-Serrano fluid work has the same shape with a smaller budget. The neural networks did not discover self-similarity. The humans chose the self-similar ansatz, chose the equations, chose the boundary conditions, and steered the networks “toward solutions with features they knew the singularities should have.” [6] The networks got a billion times more precise. The pointing was still manual.
Tao’s worry is the dark version of the same observation: if labs strip-mine open problems for answers the moment a promising direction leaks, mathematicians stop sharing promising directions, and the pipeline that produces the next direction dries up. [7] You can read that as an ethics complaint. I read it as a supply chain complaint. Taste is the scarce input and the current setup burns it without replacing it.
The autoresearch version of this problem
Now scale it down to something you can run on one GPU.
Karpathy’s autoresearch loop is beautifully dumb: an agent edits train.py, runs for five minutes, reads val_bpb, keeps the change if the number went down. [8] The human writes a program.md up front and prunes the log in the morning.
That is the Navier-Stokes setup in miniature. One oracle, the GPU, answers “did it work.” Everything the oracle cannot see is invisible to the loop. Ugly wins that game the eval split. Beautiful losses that needed twenty minutes instead of five. A hundred optimizer tweaks when the data pipeline is where the gain is. And the biggest one: the moment you realize the metric itself is wrong.
The human is supposed to fix all that by writing better instructions. Which is asking a person to verbalize their taste before they have seen any evidence. Polanyi’s line covers it: we know more than we can tell. Ask a researcher why they would try idea A before idea B and you get a shrug and “it just smells better.” Some people call that vibe. I call it the most valuable signal in the building and the one we have no interface for.
Compare, don’t describe
Here is the fact the whole design turns on. People who cannot describe their taste can still exercise it instantly on a comparison.
Show a researcher two diffs and ask which they would rather run. Ten seconds, a confident answer, zero prose. That is Bradley-Terry preference data, the exact thing RLHF was built on, and it has been sitting there unused in every research harness I have seen.
So the harness I want treats the human as a second oracle with the same properties as the GPU: slow, expensive, budgeted, and authoritative on a different question. The GPU says whether it worked. The human says whether it was worth trying. The harness’s job is to schedule queries to both so every unit of cost buys the most information.
Three rules fall out of that.
Taste steers exploration, never acceptance. The taste scorer ranks which of the agent’s K proposals get GPU time. It does not decide what counts as a win. The moment it can, the agent learns to write diffs the scorer likes instead of diffs you like, and you have Goodharted your own judgment.
The model does the verbalizing. After a batch of comparisons the agent writes its own guess at your taste as a document. You edit only where it is wrong. Recognizing wrongness is cheap for humans and producing prose is cheap for models, so the split matches each side’s strength. The comparison log stays the source of truth. The document is a cache.
Spend human minutes on surprises. A case where taste and metric agree teaches nothing. An ugly win or a beautiful loss carries all the signal. Open every human session with the disagreements.
None of this needs a trained reward model on day one. In-context taste, a few exemplars plus the last thirty comparisons, is enough to test whether the loop changes what the agent proposes. If it does, then you build the scorer.
Three UX to try
The question is whether taste can actually be extracted, not whether it exists. That is an interface problem, so here are three interfaces. Each one is a bet on a different place the taste is hiding.
1. Tinder for diffs
A ten-minute daily session. Two proposals side by side, keyboard left or right, and a third key for “run both at a longer budget.” The pairs are not random. The harness serves the surprise queue first (taste rank and outcome disagreed), then the pairs where the scorer is uncertain and the experiment is expensive, then nothing. It stops asking when it has learned enough, which is the part that keeps you coming back.
The bet: taste lives in fast comparative judgment, and the friction of writing was the only thing hiding it.
2. Redline my taste
The agent maintains a taste.md it wrote itself, regenerated after every session, shown to you as a diff against last week’s version. You do not write it. You strike lines, add “no, the opposite,” and move on. Every strike is a labeled example that goes back into the comparison log.
The bet: humans are far better at spotting a wrong description of their taste than at producing a right one, and a model is far better at producing a plausible description than a human is at producing any. Let each side do the half it is good at.
3. The scoreboard
Reserve a slice of every batch for ideas your taste ranked last. Run them anyway. Show a running tally: how often the metric agreed with your taste, how often it embarrassed you, and which kinds of ideas you keep vetoing that keep winning. Same view for the agent’s own proposals, so you can see whether it is drifting toward what the scorer rewards rather than what works.
The bet: taste is not fixed, and the loop should sharpen the human as much as the search. A researcher who sees their own veto lose ten times in a row updates. That is the one part of this Córdoba and Martínez-Zoroa never got: five years of judgment calls, and no scoreboard telling them which ones were right until the end.
I want to try the first two in Caliper first. If the agent’s proposals look different after a week and I am spending under fifteen minutes a day on it, the third one gets built. If not, I will write up the failure, which is honestly the more likely post.
References
[1] OpenAI, "On the Navier-Stokes Millennium Prize Problem," September 8, 2026. openai.com
[2] Quanta Magazine, "AI Has Solved One of Math's $1 Million Millennium Prize Problems," September 8, 2026. quantamagazine.org
[3] Fortune, "OpenAI says it cracked Navier-Stokes, one of math's grand challenges," September 8, 2026. fortune.com
[4] Karthik Duraisamy, "Navier-Stokes Regularity, what does the computation cost," September 2026. karthik-duraisamy.blogspot.com
[5] MIT Technology Review, "What OpenAI's latest controversy tells us about the future of math," September 8, 2026. technologyreview.com
[6] Quanta Magazine, "Using AI, Mathematicians Find Hidden Glitches in Fluid Equations," January 9, 2026. quantamagazine.org
[7] Terence Tao on Mathstodon, on AI companies and the strip-mining of open problems, September 2026. mathstodon.xyz
[8] Andrej Karpathy, "autoresearch," an agent loop that edits a training script against a single metric. github.com