AI News

The End of Mathematics? OpenAI Says Its AI Solved 100+ Problems

What happens to a mathematics PhD when a machine can produce the result that would have earned it?

OpenAI has given that question a concrete news hook. In its September 21 announcement, the company says a new internal model has resolved more than 100 long-standing open problems across most areas of mathematics. The announcement does not supply an itemized inventory with a proof and independent review status for each result. The headline figure remains an OpenAI claim. Read the announcement.

Even with that qualification, the prospect is disruptive. A profession organized partly around finding difficult proofs has to reconsider what those proofs demonstrate about the people presenting them. A correct result could become easier to obtain while the ability to understand, teach, and use it remains expensive to develop.

“Math is cooked” captures the anxiety. The serious version is more specific: can universities preserve a community of people who understand mathematics if producing impressive mathematical work no longer requires developing that understanding?

Analysis published September 21, 2026. Claims, public proof artifacts, and mathematical acceptance are distinguished below. This article assesses the implications of the public record; it does not certify the proofs.

What the 100-problem claim actually establishes

There are several different levels of evidence in this story. They should survive the journey from a research announcement to a headline.

Development Public evidence What readers should conclude
May 20: Erdős unit-distance conjecture An AI-generated construction and a companion account by external mathematicians A substantive disproof with a human-verified exposition; it does not determine the exact maximum for every number of points.
August 1: ten mathematical and theoretical-computer-science advances OpenAI released arguments, reasoning walkthroughs, and Lean certificates A collection of inspectable research claims with different questions of scope and novelty.
September 8: Navier–Stokes OpenAI released a proposed proof and formalization A major claim with public artifacts. Clay’s page still labels the problem “Active.”
September 21: more than 100 open problems An aggregate claim in the advisory-group announcement A reported expansion of capability, without a result-by-result evidence ledger on that page.

Sources: the unit-distance announcement, external companion paper, ten-result release, Navier–Stokes announcement, and Clay’s current problem page.

These are not interchangeable milestones. A public manuscript lets other people inspect an argument. Formal verification adds a particular kind of assurance. A specialist’s assessment adds context about whether the statement is important, correctly framed, and new. A prize institution has its own process.

Clay’s rules, for example, require publication in a qualifying outlet, at least two years after that publication, and general acceptance by the mathematical community before it will consider a proposed solution. That timetable is an award procedure, not a reason to disregard a promising argument in the meantime. Clay’s prize rules.

Kingy.ai’s Millennium Prize Problems guide and Mathematics & Science Breakthrough Tracker provide the broader evidence trail. Counting every announcement as another settled theorem would make that trail less useful.

The unsettling change is what a proof says about its author

Imagine two students presenting the same correct argument. One can explain why its assumptions matter, repair a broken step, and adapt the method to an unfamiliar example. The other can reproduce the finished document but cannot answer those questions.

The theorem has the same truth value in both cases. The students have demonstrated very different levels of mathematical ability.

AI makes this distinction harder to ignore. A polished paper can be evidence that a useful result exists without being sufficient evidence that its nominal author possesses the expertise a department wants to hire. That creates a practical problem for anyone using publications as a shortcut for evaluating people.

Daniel Litt explored the institutional danger in his August essay, The End of Mathematics. He framed it as a hypothetical adverse scenario, not a forecast: increasingly capable AI could coexist with weakening human engagement and a profession that continues rewarding the wrong outputs. His September follow-up, A beginning for mathematics, argues for adaptation and stronger attention to human understanding. Reading only the first title would misrepresent his position.

The distinction matters beyond universities. If a company hires someone to improve an optimization system, it needs an employee who can recognize when the assumptions behind a proposed method fail. Being able to obtain a plausible proof is useful. Taking responsibility for the consequences of applying it is a different demand.

Will AI replace mathematicians?

The evidence does not justify a timetable for the disappearance of mathematicians. It does justify asking which activities institutions will continue paying humans to perform, and how they will assess competence.

“Mathematician” bundles together several jobs: investigating problems, constructing arguments, developing concepts, teaching, evaluating other people’s work, and helping collaborators formulate useful questions. Automation can affect those activities at different speeds.

Here is an illustrative way to separate the pressures. It is a framework for thinking about work, not a measured forecast of job losses.

Activity Pressure from stronger AI A useful test of human competence
Finding an argument for a specified problem More candidate solutions could reduce the premium on being first Explain the central idea and its limitations without relying on the generated text.
Applying a theorem A model may identify relevant tools quickly Identify whether the real case satisfies every necessary assumption.
Reviewing a result More output could expand the volume requiring assessment Locate the decisive claim, check its dependencies, and distinguish novelty from reformulation.
Teaching mathematics Explanations and worked examples could become easier to generate Diagnose a particular learner’s misunderstanding and establish that it has been resolved.
Choosing research directions Models may propose many plausible questions Make a convincing case for which questions deserve sustained attention.

None of the last column should be advertised as permanently beyond AI. A career strategy built on the latest model’s weakest skill can age quickly. Nor does the possibility of automating an activity tell us how much of it a university, laboratory, or company will choose to automate.

The harder economic question is whether demand expands enough to support more human work. If much more mathematics becomes available, there could be more to explain, apply, and teach. But additional need does not automatically create a funded position. Departments and funders would have to assign value to that work and allocate resources accordingly.

Claims that every mathematician is doomed and assurances that every mathematician will become more valuable both skip that institutional step.

Why “the proof checks” does not answer every question

Formal proof assistants are a major part of this transition. Lean represents mathematical statements and proofs in a language whose rules can be checked. Its documentation describes the underlying logical system and the role of additional axioms. The assumptions used by a development therefore matter. Lean’s account of axioms and computation.

A checked formal proof establishes a relationship between a formal statement, its proof, and its assumptions. Readers also need to know how that statement corresponds to the claim made in ordinary mathematical language.

Consider a simple hypothetical example. An announcement says a method works for every network. Its formal theorem assumes the network is connected. The proof could be correct while the announcement is too broad. The missing work is identifying the mismatch.

That example makes no allegation about any OpenAI proof. It shows why a release should make its definitions and scope easy to inspect.

There is also a distinction between having reason to trust a result and understanding what to do with it. A researcher might accept a theorem while remaining unsure which part of the construction generalizes, which hypothesis is essential, or whether a shorter explanation exists. Those questions can produce valuable mathematics after the original proof is complete.

The strongest counterargument to “math is cooked”

The unit-distance result offers a concrete reason to resist fatalism.

The problem concerns how many pairs among a collection of points in the plane can be exactly one unit apart. The AI-generated construction overturned a conjectured bound. External mathematicians then produced a shorter, human-verified account, connecting the argument to existing mathematics and discussing its significance. Their paper explicitly distinguishes the original AI argument from the version they simplified and generalized. The companion remarks.

That is a visible model for productive work after an AI discovery. The initial result can become something other researchers can inspect, understand, modify, and build on. Mathematical activity can continue around the new knowledge.

The optimistic case deserves more than a sentence at the end of an otherwise apocalyptic story. A student with access to reliable mathematical assistance could explore more examples before meeting a supervisor. A researcher crossing into another field could test whether a promising analogy survives basic scrutiny. A small department could gain useful support for subjects it cannot cover with its own faculty.

These are opportunities, not benefits established by the 100-problem announcement. They depend on access, reliability, and the way people use the tools. A system that gives a learner the answer while concealing every useful intermediate step could be less educational than a weaker system designed to expose the structure of a problem.

What a serious mathematics department could change

Waiting for a precise date when AI becomes “better than mathematicians” is an unhelpful planning strategy. Institutions can respond to the evidence already available without predicting a final ceiling on capability.

First, they could separate assessment of a result from assessment of a researcher. The correctness and importance of a theorem deserve review. A candidate’s ability to explain assumptions, work through a variation, and recover from an error deserves its own evaluation. One assessment should not silently stand in for the other.

Second, they could ask for a contribution record with each substantial project. What did the researcher choose? What did a model propose? Which suggestions were rejected, and why? Who checked the final argument? Such a record would not make every contribution perfectly measurable, but it would give evaluators more useful information than an undifferentiated authorship claim.

The Leiden Declaration on AI and Mathematics already calls for disclosure of automated tools, careful attribution, and continued human responsibility for published work. Implementing those principles consistently would be a practical start.

Third, departments could make work on existing results count explicitly. A rigorous explanation, a useful counterexample to an overbroad interpretation, or a new application can demonstrate substantial expertise even when the first proof came from elsewhere. Hiring and promotion criteria should say when such contributions count, instead of leaving researchers to guess.

These changes involve real tradeoffs. Detailed assessment takes faculty time. Contribution records can be incomplete. Prestige will not become fair merely because an institution adds another form. The aim should be to improve the information behind decisions, with requirements proportionate to the claim being made.

The future of mathematics will depend on what we fund

OpenAI’s announcement is a reason to take the disruption seriously. It is not evidence that mathematical inquiry has reached its end, or that everyone using a public chatbot can reproduce the capabilities of an internal research system.

The risk is that institutions continue rewarding easy-to-count outputs while underinvesting in the people who can interpret them. The opportunity is to make difficult ideas more accessible and give more researchers a chance to investigate them.

Both futures require decisions about money, access, assessment, and teaching. They are not automatic consequences of a benchmark score.

For a department reviewing its next hiring round or PhD assessment, the useful question is concrete: what evidence will convince us that this person understands the mathematics well enough to extend it, explain it, and take responsibility for using it?

Read the companion analysis: OpenAI’s Math Advisory Group: Who Controls the Future of Mathematics?