phluent weekly — the mathematicians push back
25 Fields Medallists declare AI companies 'severely misaligned' with math — the same month OpenAI claims an unreleased model cracked a Millennium Prize problem.
Twenty-five Fields Medallists — math's equivalent of Nobel laureates — signed a public declaration this week saying the AI companies solving famous open problems are “severely misaligned” with mathematics itself. They did it in a hurry, skipping the usual consultation, because they judged it urgent. And they did it the same month OpenAI announced an internal model had cracked the Navier–Stokes problem, one of the seven Millennium Prize problems that have stood for roughly 90 years. The machines are winning the trophies. The people who built the game are asking them to stop treating it as a game. That tension is this week's whole story.
The rundown
Fields Medallists call AI-in-math “severely misaligned”
Terence Tao — arguably the most famous living mathematician — revealed he's one of 25 initial signatories, all Fields Medallists, to a new declaration on math and AI. Their argument is subtle and worth getting right: it isn't that AI can't do math, it's that treating famous open problems as benchmarks to be “solved first” is corrosive to the field. Those problems are landmarks whose value is the new ideas, methods, and students that grow up around the attempt — not trophies to be claimed in a press release. When labs race to announce “we proved X” on social media, they skip the community's actual stewardship (sharing methods, mentoring, not scooping people at the finish line) and quietly damage the careers of humans working the same ground. The Economist ran it under “Top mathematicians are outraged by OpenAI's methods.” When the field's most decorated researchers move this fast and this collectively, that's the signal.
OpenAI says an unreleased model solved Navier–Stokes
Here's what they're reacting to. OpenAI published a claimed solution to the Navier–Stokes existence-and-smoothness problem — one of the Millennium Prize Problems, open for ~90 years — proving that smooth fluid motion can break down into a singularity in finite time. They shared both a writeup and a Lean formalization (a proof written in software that mechanically checks every logical step, so it can't hand-wave). The kicker: it was produced by “an internal model significantly more capable than GPT-6 Astra,” explicitly released, in OpenAI's words, “to inform the world about the pace of AI progress.” Read that next to the declaration above and the misalignment snaps into focus — one side treats a 90-year problem as a landmark for human understanding; the other treats it as a pace-of-progress advertisement for a model no one outside the building has seen. Both things are true at once, and that's exactly why the mathematicians are alarmed. We covered the first tremor of this back when Claude formalized Fermat's Last Theorem — at the time it read as hopeful collaboration. The mood has shifted.
“We must pace the frontier” — and the rebuttal that landed the same day
Dario Amodei · the open-weights reply
Anthropic's CEO published a direct policy essay committing his company to embedded third-party evaluators — independent auditors sitting inside the lab — and asking governments to require every frontier company to match. Sam Altman reportedly agreed within hours. It's the highest-profile “let's tap the brakes” argument yet from someone actually building the thing. Then, the same day, came a sharp rebuttal: embedded evaluators are theater, the argument goes, because the real slowdown lever is a law requiring any model offered to the public be released as open weights — meaning the trained parameters are published for anyone to download and run, rather than rented through an API. The logic has teeth: frontier funding depends on sky-high valuations that assume the weights stay proprietary. Mandate openness and you deflate the capital driving the race. Whether or not you buy it, it reframes “pacing” from a governance question into an economic one — and that's the more honest fight.
The “coding is solved” narrative meets a private codebase
While the labs claim math trophies, here's a bracing reality check on the day job. Specific Labs licensed real production codebases from actual companies — billing systems, tax logic, customer migrations with genuine business consequences and company-specific conventions — and benchmarked frontier coding agents on tasks whose code and answers exist nowhere on the public internet. The result: the best agent (Fable 5.1 / Claude Code) resolved only 38.8%; GPT-6 Astra 33.8%; others down to 16%. Compare that to the near-solved-looking scores on public benchmarks like SWE-bench, which the models have effectively memorized. The gap is the whole point. Agents are dramatically better than a year ago and genuinely useful — but “coding is solved” is a benchmark artifact, not a description of the messy, context-heavy work most of us actually do. If you've felt that gap in your own repo, you're not imagining it.
Mistral raises €3B on “sovereign, open-weight AI”
The largest equity round ever completed by a European tech company — €3B at a €21B+ valuation, led by Samsung — and the pitch is pointed: “sovereign, open-weight AI” for governments and enterprises who don't want to be locked into an American API they can't inspect or control. Set it beside the open-weights rebuttal above and a real thesis emerges: open weights aren't just an ideology, they're becoming the business model that big money is willing to back. “Sovereignty” — the idea that a country or company should be able to run frontier AI on its own hardware, under its own rules — is turning into the organizing principle for everyone who isn't OpenAI or Anthropic. The open-vs-closed split is hardening into two economies.
Quick hits
- DeepSeek v4.1-Flash — DeepSeek ships a Flash-tier model it says is both cheaper and more capable than its own v4 Pro. The lab that keeps resetting the price floor did it again; the cheap-fast-frontier squeeze continues.
- Rust is now a Tier-1 language at Microsoft — Microsoft officially elevated Rust alongside C++, C#, and TypeScript, meaning first-class tooling and investment across one of the largest codebases on earth. The memory-safe language's march from insurgent to establishment is basically complete.
- Stockfish 19 — the venerable open-source chess engine gains up to +44 Elo, still trained on donated volunteer CPU time through its distributed Fishtest framework. A decades-running community project quietly getting stronger the patient, human, un-hyped way.
Closing
The through-line, if you want one: capability keeps arriving faster than anyone's figured out what it's for or who it's for. The mathematicians aren't afraid the machines will fail — they're afraid of what winning does to the game. That's a more interesting fear, and a more human one. And if you needed a reminder of why the doing still matters, a developer named Joel Otter wrote a raw little essay this week titled “Fuck it, make it anyway” — a case for building things by hand for the joy of it, output be damned. Sometimes that's the whole answer. See you next Sunday.
Was this issue useful?