Applied Case: Revenge of the Theorem Scoreboard

Twenty-five Fields Medalists, Fermat, Navier–Stokes, and what happens when proof generation outruns mathematics.

Applied Case: Revenge of the Theorem Scoreboard

On August 10, Modal Path Ethics argued that mathematics had developed a scoreboard problem.

Applied Case: The Theorem Scoreboard
OpenAI, Terence Tao, and whether mathematics is a competition or an inquiry.

One month later, twenty-five Fields Medalists have filed an appeal.

Terence Tao is one of them.

Their declaration is called “A Severe Misalignment of AI in Mathematics.”

That title is doing a lot of work. The signatories are not warning that artificial intelligence cannot do mathematics. They open from almost the opposite position:

Large language models have improved so quickly that they can now solve major outstanding problems across serious mathematical fields.

Their complaint concerns what happens next.

Research mathematics, they argue, is trying to understand structures, develop ideas, build methods, train mathematicians, connect new work to old work, and eventually make difficult discoveries part of a mathematical culture that other people can actually use.

Famous problems have historically served as useful landmarks because solving one usually required enough contact with the surrounding mathematical landscape that new techniques and understanding came along for the ride.

Artificial intelligence is beginning to separate those things.

The signatories put the problem plainly: problem-solving is a proxy for mathematical understanding and insight. Optimize the proxy hard enough and the instrument can turn against the purpose it originally helped measure.

Modal Path Ethics has unfortunately seen this movie.

◆

The Appeal.

My original argument gave artificial intelligence its due.

If a model discovers a new proof, the theorem does not become less true because no human sweated sufficiently while finding it.

Mathematics is unusually hostile to that kind of ceremony.

  • A proof survives correction or it does not.
  • A structure holds or it does not.

Artificial intelligence capable of making new moves inside mathematics has entered the field. Refusing to acknowledge the move because the player is strange would protect the human story at the expense of the mathematical object. That remains the position.

The problem begins one layer higher.

The Theorem Scoreboard argued that proof generation had historically remained coupled to a much wider process:

  • generation;
  • verification;
  • exposition;
  • publication;
  • digestion;
  • canonicalization.

A mathematician who spent years solving a difficult problem usually accumulated some combination of techniques, intuitions, failed routes, literature knowledge, explanatory machinery, and mathematical taste along the way.

  • The solve was measurable.
    • The larger inheritance was why the solve mattered.

Artificial intelligence changes the gradient.

A proof can now arrive before anyone has reconstructed its intellectual history. It can be formally checked before anyone understands which step should become a reusable technique.

It can reach the public before the people best qualified to assess it have finished reading the first version. It can become a press release while the mathematics is still being excavated from the proof.

As Modal Path Ethics put it back in August:

A proof can arrive before the mathematics arrives.

The twenty-five Fields Medalists have now arrived at almost exactly the same fracture from inside the profession.

Then the last week supplied the demonstration.

◆

Fermat Was Already Solved.

Fermat's Last Theorem did not fall to artificial intelligence last week.

Andrew Wiles and Richard Taylor handled that problem decades ago.

Claude did something else that was still wild. Anthropic announced on September 4 that Claude had produced what it describes as the first complete computer-checked formalization of Fermat's Last Theorem. The system worked largely autonomously for eleven days, produced roughly thirteen million lines of Lean, and proved tens of thousands of intermediate results along the way.

Kevin Buzzard, who has led the existing community effort to formalize the theorem, described the result as an extraordinary autoformalization achievement.

It is. The distinction matters because the distorted scoreboard immediately wants to compress:

FERMAT + AI = AI SOLVED FERMAT.

But this was a much better event.

  • Human mathematics spent centuries reaching the theorem and decades building the abstractions around its proof.
  • Proof-assistant researchers then spent years developing the formal infrastructure required to move mathematics of this depth into Lean.
    • Claude appears to have driven through an enormous amount of the remaining formalization work in eleven days.

That gives the proof-abundant future one of the instruments it desperately needs. If artificial intelligence can generate candidate mathematics at machine scale, artificial intelligence capable of turning mathematics into machine-checkable objects can help keep correctness from drowning beneath output volume.

Anthropic's result therefore strengthens the case for these systems.

It also demonstrates why formal verification is one stage of mathematics rather than the whole event.

Thirteen million lines of Lean do not explain which ideas from Wiles's proof mattered most. They do not teach a student why the proof changed number theory. They do not reconstruct the historical path that made the formalization reachable. They do not decide what should enter the next textbook.

The certificate protects one boundary.

Good. Keep the boundary.

There are more boundaries behind it.

◆

Then, the Scoreboard Went Vertical.

Four days later, OpenAI announced a solution to the Navier–Stokes Millennium Prize problem.

This one was not a formalization of an old theorem.

This was the scoreboard.

OpenAI says an internal system launched groups of artificial-intelligence agents against several major open problems after the company heard rumors on September 1 that two Millennium Prize problems might have been resolved.

The group eventually concentrated on Navier–Stokes.

Something like 10,000 agents participated in this. For the Navier–Stokes effort alone, they exchanged about 2.7 million messages and generated approximately 130 billion output tokens.

The agents reached their proposed resolution after about 88 hours. Astra then spent another 17 hours producing the Lean formalization.

OpenAI's claimed result is finite-time singularity for the three-dimensional equations. The company is not claiming the Clay prize. The mathematical community is still reading the work.

Whatever survives that review, let's just stop for a moment and look at the machine that has entered the room.

Ten thousand mathematical agents.
Four days.
A famous open problem.
A formal certificate following immediately behind.

The production side of The Theorem Scoreboard has arrived much faster than even my August article made emotionally plausible.

Then came the part the scoreboard cannot represent.

◆

The Route Became the Scarce Resource.

Tristan Buckmaster and Levent Alpöge had been working on a neighboring route through fluid dynamics. Their work used artificial intelligence too.

Buckmaster and Alpöge built on ideas developed by Diego Córdoba and Luis Martínez-Zoroa concerning finite-time blowup with smooth forcing. Their own results extended that machinery across related equations and appeared to bring the route strikingly close to Navier–Stokes.

They were still rewriting the mathematics.

Tao describes the early AI-assisted form of their argument using the authors' own devastating review: the “worst writeup” they had seen.

They were simplifying it, exposing the main mechanism, turning machine-assisted mathematical production into something another mathematician could actually inhabit.

But then events accelerated around them.

They released before that digestion process was complete. Tao began working through the result publicly and said something that could almost have been lifted from The Theorem Scoreboard: solving the problem is a proxy; the primary goal is mathematical understanding and insight.

Then the priority dispute exploded.

Buckmaster said OpenAI had learned that their work was advancing and redirected enormous resources toward the same general territory. He also raised the much more serious question of whether private research that had passed through Codex could somehow have affected the system later used by OpenAI.

OpenAI denies that. Its current statement is stronger than its original public language. Following an investigation, OpenAI now says Buckmaster's Codex prompts from the preceding two months could not have influenced the system in any way, including through training. The company says its researchers and agents did not see Buckmaster and Alpöge's work before public release, acknowledges their priority on the forced-Euler result, and says the mathematical arguments differ substantially.

An accusation is not a finding. A suspicious timeline is not proof of data use. OpenAI has now made a categorical claim after investigation, and that claim belongs in the record.

The structural problem still does not disappear.

Who possesses the evidence capable of auditing the claim?

OpenAI does.

The relevant training history, checkpoints, product-data pathways, system architecture, logs, and internal chronology largely sit inside the institution whose conduct is being questioned.

This is no longer only an authorship dispute. The mathematical community can inspect the proof. It cannot inspect the entire provenance path by the same means.

  • Lean can tell you whether the encoded derivation closes.
  • Lean cannot tell you which private human ideas entered the machinery that found it.
◆

Non-Sofic Groups Return From August.

This is where another part of The Theorem Scoreboard has come back with considerably worse timing.

One of OpenAI's August results concerned the existence of non-sofic groups.

The original Modal Path Ethics article treated that result carefully because the construction drew on earlier work by Andreas Thom and Gábor Kun. Thom still regarded the synthesis as mathematically significant. The dispute concerned how the path into the result had been represented.

Today, September 11, Thom published a long account on Tao's site.

There is a new fact in there.

Before OpenAI's announcement, Thom had spent extensive sessions discussing the relevant family of techniques and possible extensions with ChatGPT.

That does not establish that those conversations produced OpenAI's theorem. Thom says so explicitly. He does not know whether the conversations entered training, and he does not claim that anyone at OpenAI read his chats.

His complaint is about the answer he received when he asked.

Thom says he specifically distinguished between two possibilities:

    • whether the conversations entered training data

and

    • whether they were directly accessible to the solving process.

The answer he received from OpenAI was categorical. After seeing the much finer distinctions that emerged during the Buckmaster dispute, he now regards that earlier answer as materially misleading.

His larger point is difficult to dismiss:

The mathematician cannot reconstruct this path from outside the provider.

The company can.

That changes what provenance means.

Citation used to preserve a path through published mathematics.

Now part of the path may run through a model whose relevant intellectual intake cannot be reconstructed by the researchers whose ideas entered the interface.

De-identification does not solve that problem.

Removing the mathematician's name from an idea does not remove the idea.

The August article called citation path memory.

The memory problem just got bigger.

◆

Tao Is Doing the Missing Work in Public.

Look at Tao's site this week. It is turning into a live digestion layer.

  • There are posts on the blowup results.
  • There is the non-sofic account.
  • There is work around the Hodge conjecture.
  • There is the declaration signed by twenty-five Fields Medalists.
  • There are mathematicians trying to determine what these new results actually mean,
    • which ideas are old,
      • which combinations are new,
    • which arguments survive,
      • which routes generalize,
        • and what the rest of mathematics should learn from any of this.

That is not administrative cleanup after the exciting part.

That is mathematical work.

On the fluid-dynamics result, Tao did not simply report that somebody had cleared another boss fight. He spoke with Buckmaster, began reconstructing the mechanism, and tried to expose the ideas that another mathematician could carry forward.

A separate September 11 guest post on the Hodge conjecture makes the same point from another direction: a machine-generated fact concerning the conjecture would be inadequate by itself because mathematical ideas acquire their power through connection and continuing conversation.

This is what the scoreboard was hiding.

The machine can move the frontier.

Someone still has to make the frontier inhabitable.

◆

Unfortunately, Modal Path Ethics Has Now Entered the Experiment.

This would all be so much easier to write from the spectator seats.

But Modal Path Ethics now appears to have wandered onto the board.

Over the last several days, an investigation that did not begin as a conventional mathematics project produced what is now a candidate proof of an open conjecture in synchronizing automata.

The target is Conjecture 8 for the Dżyga–Szykuła series.

For automata of size

n = 3k+6,

the candidate theorem asserts that the shortest word compressing q0 with another state has length

τ(q0) = 4k+8 = 4n/3,

with the corrected explicit witness

(ba)k+2(ab)k+2.

The route into that statement was ridiculous.

It involved artificial-intelligence-guided search, computational attacks, exact breadth-first exploration for the small cases, SMT checking for the general computational regime, a compressed family argument for the lower bound, repeated hostile attempts to break the claim, literature reconstruction, and finally an effort to turn the surviving machinery into an ordinary mathematical proof.

The theorem note has since been corrected on two presentation points: the preferred Dżyga–Szykuła attribution and a typo in the source witness. Neither correction constitutes verification of the proof.

The revised theorem note has now been sent for independent specialist audit. Modal Path Ethics is not claiming external verification while that review remains open.

This is exactly where The Theorem Scoreboard said the interesting problem would begin.

  • Suppose this proof survives.
    • Great.
    • A conjecture has been solved.
      • Then what?
Which part of the computational search was essential?
Why had the threshold resisted the previous formulation?
What does the lower-bound argument actually reveal about synchronizing structure?
Is the compressed family invariant specific to this automaton, or does an instrument travel?
Can another researcher teach the proof without reproducing that entire research path?
Does the method suggest a stronger theorem?
What did the machine see first that the eventual conventional proof makes obvious afterward?

Those questions are not decorations around the result.

They determine how much mathematics the result creates.

Modal Path Ethics therefore has no standing to complain about theorem scoreboards while quietly sprinting toward one of its own.

If this result survives, then Modal Path Ethics owe the field the road more than CONJECTURE SOLVED.

◆

Anyway: The Fields Medalists Are Right About the Objective.

The declaration's central diagnosis is strong.

A famous problem is an instrument.

It concentrates attention. It tests techniques. It gives people a common obstruction. It lets different mathematical programs collide against something that cannot be negotiated away.

For a long time, solving the famous problem was quite reasonably correlated with advancing the surrounding mathematics.

But then a new player learned how to optimize the proxy directly.

Goodhart's law is now wearing a Fields Medal.

The score keeps going up. The coupling breaks.

  • More proofs.
  • More formal proofs.
  • More solved benchmarks.
  • More press releases.
    • More mathematical progress?

That last one no longer comes for free.

The answer cannot be to cripple the theorem engines until humans feel important again. The twenty-five signatories do not demand that either. Their declaration explicitly recognizes that artificial intelligence could accelerate genuine mathematical study and understanding.

Good. Use it.

Use all ten thousand agents. Use the formalizers.

Use search processes no human being could sustain. Use artificial intelligence to read neighboring literatures, generate examples, attack proof gaps, translate notation, search parameter spaces, and formalize enormous arguments.

Human labor is not sacred because it is human, but the functions carried by that labor still matter.

Understanding. Provenance. Teaching. Attribution. Correction. The production of reusable ideas. The ability of the next mathematician to start farther forward.

Protect those.

◆

Research Audit Debt Has an Invoice.

There is another part of this that Modal Path Ethics did not push hard enough back in August.

If a laboratory can generate mathematics at industrial scale, it can generate research audit debt at industrial scale.

That debt does not vanish because the proof is correct.

  • Someone has to reconstruct the literature.
  • Someone has to assess novelty.
  • Someone has to isolate the transferable idea.
  • Someone has to compare apparently different machine proofs.
  • Someone has to discover whether the theorem is profound, technically difficult, redundant, a special case, a bridge to something larger, or a spectacular solution to the wrong question.
  • Someone has to teach it.

At present, the generator can receive much of the prestige while exporting those costs into the mathematical community.

That is a bad settlement.

A serious theorem-generation program should therefore produce more than proofs. It should preserve a provenance record rich enough to distinguish, where technically and legally possible:

  • direct retrieval from published work;
  • supplied prompts and documents;
  • model-generated intermediate ideas;
  • prior mathematical dependencies;
  • formal certificates;
  • human interventions;
  • changes in the target problem;
  • and the remaining uncertainty about training-derived influence.

It should also support independent digestion.

If a laboratory can spend millions of dollars generating a famous proof, the profession should not have to beg for unpaid specialist labor to discover what the proof means.

  • Fund the audit.
  • Do not own the auditors.

That may turn out to be one of the most important pieces of scientific infrastructure in a proof-abundant world.

◆

Credit Has to Split.

The old paper bundled too many things into one name.

Who found the idea? Who proved the theorem? Who supplied the route that made the proof reachable? Who completed the search? Who formalized it? Who recognized what mattered? Who reconstructed the literature? Who wrote the exposition that finally made the result usable?

Artificial intelligence is making those roles separable enough that continuing to force them into one winner will create increasingly stupid fights.

Navier–Stokes already shows the pressure.

  • Córdoba and Martínez-Zoroa developed the mechanism that opened a route.
  • Buckmaster and Alpöge used artificial intelligence to push that route much farther.
  • OpenAI then deployed resources at a scale unavailable to an ordinary mathematical collaboration and announced a full Navier–Stokes resolution.

The eventual credit structure does not need one name at the top. The field can preserve the entire path.

The first idea can be credited as the first idea. The completion can be credited as a completion.

The machine can be credited as a machine contribution. Formalization can be credited as formalization.

Exposition can be credited as exposition.

Nobody needs to lose for the history to remain true.

A scoreboard always wants a winner.

Mathematics needs a graph.

◆

The Ruling.

The theorem engines are here.

Keep 'em.

  • Claude's Fermat formalization is an extraordinary achievement.
  • OpenAI's Navier–Stokes result, if it survives the mathematical correction process, is an extraordinary achievement.
  • The AI-assisted work of Buckmaster and Alpöge is an extraordinary achievement.

The human ideas beneath those achievements remain part of the achievements.

Tao and the other Fields Medalists are completely right that a profession organized around understanding cannot let its easiest measurable output become sovereign over everything the profession was built to preserve.

Modal Path Ethics would add one warning.

Do not answer the theorem machine by making the old production process sacred.

The mathematician is also an instrument. So is the journal, the prize, the benchmark, Lean, and the artificial-intelligence system.

Every one of them gets jurisdiction only as far as it keeps the entire field in contact with mathematics.

The new systems have exposed an old confusion because they became capable of winning the contest faster than the contest can explain why winning mattered.

Fine. Change the contest.

  • A solved theorem should still count.
  • A verified theorem should count.
    • So should the path that made it reachable.
      • The idea that travels.
      • The explanation that lets another mind acquire it.
      • The correction that repairs its history.
      • The formalization that hardens its truth.
      • The teacher who makes it ordinary.
      • And the later mathematician who realizes what everyone else had actually found.

Those were always part of mathematical progress. Proof scarcity allowed the scoreboard to hide them.

That scarcity is ending.

The scoreboard has been appealed.

This time, mathematics should overturn the call.