> ## Content Index
> Fetch the complete content index at: https://modalpathethics.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Applied Case: The Theorem Scoreboard
- URL: https://modalpathethics.com/applied-case-the-theorem-scoreboard/
- Published: 2026-08-10T10:30:54.000Z
- Updated: 2026-08-10T10:30:54.000Z
- Description: OpenAI, Terence Tao, and whether mathematics is a competition or an inquiry.
- Author: Aidan Lawson
- Tags: Applied Case, Instrument Jurisdiction, Modal Path Ethics

Mathematics has developed a scoreboard problem.

On August 1, OpenAI announced that an internal version of its forthcoming Astra model had produced ten new results across high-dimensional geometry, coding theory, group theory, operator algebras, complexity theory, quantum information, lattice problems, and extremal combinatorics. OpenAI says the model generated the mathematical arguments, humans prepared the manuscripts with it, and the model then formalized the arguments as Lean certificates. The token cost required to find the ten solutions, priced at current Sol API rates, was roughly two thousand dollars.

That is an extraordinary technical event. It also arrived in a field that already knows how to turn extraordinary technical events into races.

- Who solved the problem first?
- Which lab has the strongest model?
- Which theorem still belongs to humans?
- Which benchmark has fallen?
- How many open problems can the next system clear before lunch?

Then mathematicians began reading the paper.

[Scientific American](https://www.scientificamerican.com/article/openais-latest-math-breakthroughs-commit-research-misconduct-experts-say/?ref=modalpathethics.com) reported complaints that some of the most impressive results depended on recent human work that the initial public framing did not adequately acknowledge. One mathematician accused OpenAI of research misconduct over a sphere-packing argument. In the group-theory case, researchers found a genuine new result assembled from recent ingredients that made the field less stagnant than the original announcement had suggested.

OpenAI defended the work and said it was holding itself to ordinary mathematical standards. Its current release explicitly says that questions about artificial intelligence in mathematics cannot be answered by a technology company alone.

Two days later, [New Scientist](https://www.newscientist.com/article/2583307-why-mathematician-terence-tao-thinks-ai-must-spark-a-rapid-revolution/?ref=modalpathethics.com) put Terence Tao into the middle of the same argument.

Tao's intervention is useful because he turns past the obvious question.

Artificial intelligence is getting better at mathematics. 

- **What is mathematics trying to do?**

That is suddenly the harder problem.

> ◆

# **OpenAI Actually Did Something.**

The weakest response to the Astra results is also the easiest one:

- *The model only combined things humans already knew.*

Welcome to mathematics.

Mathematicians inherit definitions, lemmas, methods, failed approaches, standard constructions, tricks, examples, counterexamples, notation, conjectures, and entire theories from other mathematicians. A new proof does not have to arrive from a spotless dimension containing no previous work. Mathematical creativity often consists in seeing that two structures already present in the field can be made to answer one another in a way nobody had completed before.

The group-theory result is a clean example. Astra's construction drew on ideas from earlier papers. That does not by itself make the resulting theorem fake. Andreas Thom, one of the mathematicians whose earlier work enters that history, described the result as creative and elementary. The interesting act may have been the synthesis.

That is still an act.

[Field Instruments: The Mathematics](https://modalpathethics.com/applied-case-the-mathematics-problem/) already gave mathematics its strongest possible status inside Modal Path Ethics. Mathematics is one of humanity's most powerful instruments for contacting stable relation. Once the field has been cut into variables, units, operations, and admissible relations, mathematical rigor can carry an argument much farther than intuition alone ever could.

A machine that can operate seriously inside that closure has entered a serious field. So give Astra the point.

If the proof is new, correct, and independently checkable, then something new has entered mathematical reachability. The theorem does not become *less true* because a model found it. A proof does not acquire a human essence through the sweat traditionally required to produce it. 

The structure either survives the correction regime or it does not.

This is where some anti-artificial-intelligence rhetoric breaks contact with mathematics itself. If the model genuinely found a proof, insisting that it did not really count because the path looked unlike human creativity places psychological ceremony above the result.

Modal Path Ethics has no reason to protect that ceremony.

It **does** have a reason to protect the field around the proof.

That is where the argument changes.

> ◆

# **Citation Is Path Memory.**

- A theorem can be correct while its history is represented badly.

This is the part of the controversy that survives every argument about machine creativity.

[Scientific American](https://www.scientificamerican.com/article/openais-latest-math-breakthroughs-commit-research-misconduct-experts-say/?ref=modalpathethics.com) reports that Steven Miller objected to the sphere-packing result because a key argument had appeared in his earlier work with a collaborator. The same report says the non-sofic group construction combined ideas from papers in 2016 and 2019, while OpenAI's first public framing made the area sound more dormant than specialists believed it was. OpenAI has since updated its language and says it plans ordinary revisions to the paper.

The accusation should remain an accusation while the relevant mathematicians and publication processes do their work. The structural point does not depend on deciding the misconduct charge here.

**Citation is path memory.**

A **citation** tells the field that this route did not begin at the current paper. It preserves the bridges that made the new move reachable. It tells later mathematicians where an instrument entered, which prior theorem carries load, which problem was already moving, and which neighboring path may still contain useful structure.

Credit matters. Careers, jobs, grants, reputation, and intellectual ownership all run through attribution.

There is also a deeper epistemic function.

A field that receives results without their path becomes easier to misunderstand.

[The Invisible Board](https://modalpathethics.com/applied-case-the-invisible-board/) made this point through Go on the same day this article was written. A visible position is historical. The stones in front of the player are compressed evidence of earlier moves. The position cannot be understood completely as a fresh inventory detached from the route that produced it.

Mathematics has the same problem at a different scale.

A proof sits inside a literature. The definitions have ancestors. The lemmas have routes. The obstruction somebody finally bypasses may be visible only because ten earlier people spent years discovering where the wall actually was.

Delete the path and the theorem remains true.

But the mathematical field becomes thinner around it.

This is why attribution is larger than etiquette. Proper attribution keeps the discovery connected to the structure that made it possible.

The machine may discover the theorem and still fail to discover where the theorem came from.

That failure becomes more dangerous as generation accelerates.

> ◆

# **The Opponent Is Part of the Instrument.**

There is a good case for competition in mathematics.

Mathematical history contains prizes, priority races, rival schools, Olympiads, public challenges, departmental competition, grant competition, journal competition, and the quieter competition inside a room where two people have different ideas about how a proof should work and each would enjoy being correct.

Competition can apply pressure.

Pressure can expose weak structure.

[The Invisible Board](https://modalpathethics.com/applied-case-the-invisible-board/) makes the same argument for games: the opponent is part of the instrument. A private theory can protect itself. Another player has an active interest in locating exactly where your reading of the board fails. The shared rules let disagreement become consequence.

Mathematics has its own adversarial correction machinery.

A conjecture is public enough to attack. A proof can be checked line by line. Another mathematician can search for the hidden assumption, the missing case, the false equivalence, the easier route, the stronger theorem, or the older paper everyone somehow forgot.

Priority can help this field. Prizes can help it. Difficult benchmarks can help it. Even the deeply silly human desire to be the person who finally solved the thing can keep somebody inside a problem long enough for reality to answer.

Competition belongs inside inquiry as one of its instruments.

The instrument works while winning remains coupled to the larger field.

- You solve the problem first, and the result becomes available to others.
- You win the prize, and the proof enters the literature.
- You defeat a rival approach, and the failure teaches the field something.
- You train for the Olympiad, and the difficult problems train capacities that travel far beyond the medal.

The competition earns its place because the contest produces more mathematical contact than the scoreboard alone contains.

Then artificial intelligence changes the coupling.

> ◆

# **The Scoreboard.**

Terence Tao's July 24 public lecture at the International Congress of Mathematicians begins by asking the audience to provisionally grant a strong assumption:

- Suppose artificial intelligence soon performs a meaningful fraction of research-level mathematical tasks at usable levels of cost and quality.

He does not spend the rest of the talk celebrating or mourning that possibility.

He asks **what** the mathematical community is actually optimizing.

His list is wider than problem solving:

- solve open problems;
- develop theories and techniques;
- understand the world;
- build a mathematical community;
- train future mathematicians;
- contribute to shared mathematical knowledge;
- create enduring work with aesthetic value.

Historically, these goals often moved together. Someone who solved an important problem usually had to understand enough mathematics to explain something, teach something, build something, or leave a technique other people could use. The solve-count could stand in as a rough proxy because other goods frequently arrived with it.

Tao's warning is Goodhart's law.

- Once the proxy becomes an optimization target, the old correlation can break.

A system can increase the number of solved problems without increasing mathematical understanding at the same rate. It can generate correct proofs faster than the community can read them. It can optimize a class of benchmark-friendly problems while theory-building, pedagogy, taste, exposition, and the selection of genuinely important questions lag behind.

This is [the degenerate metagame](https://modalpathethics.com/applied-case-the-solved-game-the-degenerate-meta/) arriving in mathematical research.

A degenerate meta appears when an incentive system finds an exploit and begins compressing the field around it. Everyone can be playing correctly. The score can keep improving. The activity can still lose the structure that made the score worth pursuing.

Artificial intelligence does not create this problem.

Publication counts were already gameable. Citation counts were already gameable. Priority already rewarded speed. Universities already converted research into rankings. Grant systems already learned to demand legible output. Mathematics already possessed competitions whose local incentives could distort the wider inquiry.

Artificial intelligence changes the gradient.

Suddenly one of the easiest outputs to measure may become one of the cheapest outputs to generate.

- **Problems solved** can go vertical.
  - The rest of mathematics does not automatically follow.

That is the fracture.

The competition was tolerable while winning remained coupled to inquiry.

The new systems make it possible to win faster than the field can understand what winning produced.

> ◆

# **Proof Before Mathematics.**

Tao's most useful contribution is a pipeline.

A mathematical result does not enter the field at the moment a candidate proof first appears.

It develops.

- **Generation ->**  
  - **verification ->**  
    - **exposition ->**  
      - **publication ->**  
        - **digestion ->**  
          - **Canonicalization.**
- The first stage finds an argument.
  - **Verification** establishes that the argument actually proves what it claims to prove.
    - **Exposition** makes the crucial structure intelligible to other people.
      - **Publication** places the result into a community process of review and responsibility.
        - **Digestion** lets other mathematicians understand what was found, compare it to the literature, identify the reusable techniques, decide how deep the result is, and connect it to other work.
          - **Canonicalization** occurs when the result has been absorbed into the stable theory of the subject strongly enough to teach the next generation and support later work.

Tao argues that artificial intelligence accelerates the early stages much more aggressively than the later ones. 

Proof generation moves first. Formal verification follows. Exposition can be assisted. Community acceptance remains slow. Canonicalization is slower still.

In his ICM lecture, Tao calls canonicalization the stage least amenable to artificial-intelligence optimization and the most valuable part of the whole process.

That sentence should end the fantasy that a theorem count is the same thing as mathematical progress.

A proof can be **true** before anybody understands its significance.

A proof can be **verified** before the literature around it has been reconstructed.

A proof can be beautifully typeset before anyone has identified the one idea worth carrying forward.

A proof can **exist** while the field still does not know **what** it has.

A **proof** can arrive before the **mathematics** arrives.

This is the new bottleneck.

Tao calls the emerging condition **proof abundance** and **proof indigestion**. The field has been organized for a world where producing a serious proof was scarce enough that verification, explanation, and digestion could usually gather around it. If generation becomes cheap, those downstream capacities become the scarce resource.

Modal Path Ethics can name the exported burden more generally:

- **research audit debt.**

A system produces a result quickly.

Then somebody else has to determine:

- whether the proof is actually correct;
- whether the formalization encodes the intended theorem;
- what is genuinely new;
- which prior work made it reachable;
- whether the citations are complete;
- whether the result is deep, useful, trivial, redundant, or strangely beautiful;
- which idea inside the proof deserves to become a reusable instrument;
- what the field should do next.

The producer can externalize that work onto the mathematical community.

At small scale, this is ordinary scholarship. Everyone depends on reviewers, readers, editors, seminar audiences, librarians, teachers, and future researchers.

At machine scale, the burden can change phase.

One lab can generate hundreds of pages at a tempo that requires dozens of specialists to inspect. Formal certificates lower one part of the burden while leaving several others untouched. The proof assistant can tell you that the encoded derivation closes under its rules. It cannot, by that fact alone, tell you whether the result was framed honestly, whether the right theorem was encoded, whether the literature was represented well, whether the result deserves attention, or what it teaches the field.

The [Leiden Declaration on Artificial Intelligence and Mathematics](https://leidendeclaration.ai/?ref=modalpathethics.com) is basically a constitutional attempt to protect those downstream functions before they disappear beneath output volume. It emphasizes attribution, independent verification, transparent arguments, evaluation of depth and significance, community understanding, and human responsibility for the result and its citations.

The declaration is defending the field around proof generation against being compressed into proof generation alone.

> ◆

# **Lean Does Not Digest the Theorem For You.**

OpenAI did something unusually responsible with the Astra work: it formalized each argument in Lean.

Formal verification is exactly the kind of instrument a proof-abundant world will need. If machines can generate candidate mathematics faster than humans can check every line, machine-checkable proofs can preserve a hard boundary against a flood of plausible nonsense.

The boundary is valuable precisely because it is narrow.

- Lean can certify a formal relation under a specified environment.
  - It cannot certify the entire social and intellectual event surrounding that relation.

This is a direct continuation of the [Mathematics Problem](https://modalpathethics.com/applied-case-the-mathematics-problem/). Formal rigor begins after the cut. A formal proof can be immaculate while the larger selection remains confused.

- What theorem did we choose to encode?
- What definitions were selected?
- Which earlier result was imported?
- What does the theorem illuminate?
- Which questions does it make reachable?
- Which mathematical community can now use it?
- What should be taught from it?
- What was hard?
- What was surprising?
- What should be forgotten as scaffolding once the deeper structure is visible?

Lean's jurisdiction ends there.

[Field Instruments: The Scientific Method](https://modalpathethics.com/applied-case-the-scientific-method/) makes the parallel point for science. A method becomes trustworthy partly through correction paths: replication, criticism, open records, adversarial review, better instruments, and public repair. The discipline is powerful because it can answer its own earlier claims.

Proof assistants belong to that family. They strengthen correction.

The mistake arrives when the presence of a very strong correction instrument lets the producer skip the rest of the inquiry.

The computer can check the bridge.

The community still has to decide where the bridge goes.

> ◆

# **The Constitutional Crisis.**

Tao now describes this as mathematics' first major foundational crisis in more than a century.

The early twentieth-century crisis concerned foundations in the familiar technical sense: sets, infinity, axioms, paradoxes, proof, consistency.

The new crisis concerns values and practices.

- What counts as progress?
- What counts as authorship?
- What earns priority?
- What work deserves prestige?
- What should be automated?
- What has to remain a human responsibility?
- Who decides which problems deserve attention?
- What does it mean to have understood a theorem?
- What is a mathematician for after theorem generation becomes abundant?

Tao's current public summary, updated August 6, sharpens this into an unusually urgent warning. He says mathematicians have only months to begin reimagining the profession, calls for organizing and activism rather than passive acceptance, and says the field has temporarily lost the narrative about what a mathematician is.

That language is revealing.

A technology company does not need to pass a law defining mathematics.

It can define the profession indirectly through demonstrated capability.

If every major public announcement says:

- *Our model solved ten hard problems.*
  - then **problem solving becomes the legible unit of mathematical value**.

If the benchmark says:

- *Our model beats experts.*
  - then **expertise becomes whatever the benchmark measures**.

If the press cycle says:

- *Another human domain has fallen.*
  - then **mathematics becomes territory in an artificial-intelligence competition**.

The field begins inheriting an external constitution through repetition.

This is an instrument-jurisdiction problem.

Technology companies have standing to say what their systems can do. They have standing to publish results, show evidence, build tools, and participate in mathematics once their systems genuinely participate in mathematics.

They do not acquire sovereignty over the definition of mathematics by becoming exceptionally good at one part of it, however.

The theorem generator does not get to define the theorem field.

The scoring instrument does not get to define the purpose of play.

The proof benchmark does not get to define mathematical life.

Tao's intervention is valuable here because he is not trying to defend a pre-AI profession as sacred territory. His positive vision includes human-machine collaboration, large-scale distributed mathematics, automation of routine work, population studies across huge families of problems, and new infrastructure built around proof abundance.

He expects the instrument to enter. He is arguing over the **constitution** it enters.

That is the correct fight.

> ◆

# **Mathematics Has Already Seen the Invisible Board.**

There is a funny timing problem here.

[The Invisible Board](https://modalpathethics.com/applied-case-the-invisible-board/) was published on the same day Modal Path Ethics encountered this controversy.

That article moved from Go into mathematics through exactly the distinction now under pressure.

The visible symbols do not exhaust the active relation.

A proof is public correction of mathematical intuition. It makes a suspected structure inspectable outside the mind that found it.

That is powerful. Yet a proof still lives on a wider board.

It has history.

It has neighboring work.

It has a community capable of understanding it badly or well.

It has students who may inherit a technique without inheriting the path that made the technique intelligible.

It has institutional incentives capable of rewarding volume, priority, prestige, difficulty, novelty, beauty, usefulness, fashion, or whatever else the local scoring system can see.

Artificial intelligence changes the visible pieces on that board dramatically.

It may become a stronger problem solver than any individual human.

That does not erase the invisible board.

It makes the invisible board more important.

The question moves from:

- *Can the model solve this problem?*

into:

- *What kind of mathematical field does this new capability generate?*

That is a [Field Intelligence Gap](https://modalpathethics.com/applied-case-the-field-intelligence-gap/) question.

- Strategic intelligence asks how to win under the existing rules.
- Field intelligence asks what game everyone inherits after the move.

Artificial intelligence can become extraordinarily strategically intelligent inside mathematics. It can search, combine, calculate, formalize, test, and perhaps increasingly prove.

The field-intelligent question arrives one level up.

- If proof generation becomes cheap,  
  - what happens to review?
- If priority becomes a prompting race,  
  - what happens to literature search?
- If theorem count becomes an artificial-intelligence benchmark,  
  - what happens to problem selection?
- If humans stop traversing difficult proofs,  
  - what capacities fail to develop?
- If machines can search the long tail,  
  - which neglected regions suddenly become visible?
- If machines can connect literatures humans kept separate,  
  - which new theories become reachable?
- If the field refuses the tools out of status fear,  
  - which discoveries does it close?
- If the field surrenders its values to the tools,  
  - which parts of mathematics disappear while the proof count rises?

That is the board.

> ◆

# **The Better Game.**

There is no need to choose between **artificial intelligence** and **mathematics**.

There **is** a need to choose the relation.

The Better configuration is almost embarrassingly available already.

Let artificial intelligence search aggressively.

Let it formalize.

Let it test conjectures, generate counterexamples, explore long-tail problems, combine literatures, perform routine derivations, build candidate proofs, and expose structures no human would have found on the same schedule.

Then, preserve the rest of the field **around** that capability.

- Require serious provenance work before priority claims become settled.
- Keep formal verification as a strong correction layer without pretending it exhausts interpretation.
- Reward exposition that identifies the actual new idea.
- Treat peer review and digestion as mathematical labor rather than volunteer janitorial work performed after the glamorous part is over.
- Give canonicalization prestige because stable shared understanding is what turns isolated discoveries into a mathematical civilization.
- Build public and academic infrastructure so the field is not structurally dependent on whichever company currently owns the strongest theorem engine.
- Train students through generative resistance where the struggle is itself part of acquiring mathematical perception.
- Use competition where competition produces contact.
- Stop the competition from becoming sovereign over the inquiry.

Tao has even begun pushing priority in this direction. In his current account, the race should move away from who can make an artificial-intelligence system produce the first raw proof and toward who can provide a genuinely public scientific exposition of the result.

That is a profound change in the win condition.

The first person to reach the destination is less important than the first person who can build a road other mathematicians can actually use.

Competition survives, now answerable to inquiry again.

The same move works at larger scale.

Artificial intelligence companies can compete to build stronger systems.

Mathematicians can compete to solve hard problems.

Journals can compete for excellent work.

Students can compete at Olympiads.

None of those contests needs to become the constitution of mathematics.

The larger game is cooperative.

Everyone is trying to gain better contact with structure that no participant owns.

A theorem cannot be conquered.

It can be found.

Then understood.

Then connected.

Then taught.

Then used to find something else.

> ◆

# **The Ruling.**

OpenAI's Astra results are impressive.

The strongest criticism does not require pretending otherwise.

If the model generated genuinely new, correct mathematical arguments, then artificial intelligence has crossed another important threshold in research mathematics. Formal verification makes the event stronger. The ability to produce ten substantial results at low marginal inference cost changes what kinds of mathematical search are reachable.

The complaints about attribution are also serious.

A correct proof does not erase the path that made it possible. Citation preserves mathematical memory. Literature review keeps the discovery connected to the field that generated it. Priority without provenance turns inquiry into a race whose finish line moves faster than its mapmakers.

Terence Tao sees the larger fracture.

For a long time, mathematical competition and mathematical inquiry were coupled tightly enough that winning one contest often advanced several other goods at once. Solve the hard problem. Build the theory. Train the mathematician. Explain the technique. Add to the literature. Give the next generation a stronger starting point.

Artificial intelligence can separate those goods.

A lab can win the theorem race while the community inherits verification load, citation reconstruction, exposition work, review pressure, and canonicalization debt.

A model can generate the answer while other people still have to discover what the answer means.

The model can enlarge mathematics while exposing the scoreboard's insufficiency.

The theorem does not care who won.

The theorem does not know OpenAI from UCLA. It does not know whether a human spent twelve years finding the bridge or a model searched the relevant space before dinner. It does not award moral points for suffering. It does not protect anybody's professional identity.

The relation either holds or it does not.

**Inquiry** cares about the path.

It cares because later understanding depends on knowing where the result came from, what it changed, which instruments survived correction, who can explain it, what it connects, what remains open, and whether the field became more capable of continuing after the discovery.

Competition can serve that process.

Competition cannot own it.

If artificial intelligence becomes capable of winning the mathematical contest, mathematicians have been handed an unusually sharp opportunity to ask whether the contest was ever the purpose.

That answer will not be found on the leaderboard.

The leaderboard is one of the things now under review.