Opinion · Mathematics and AI · October 8, 2026
AI will accelerate mathematics. A profession that values understanding should demand better proofs, better tools and wider access. Treating the right to attempt a problem as something an established community can withhold is a mistake.
Terence Tao’s blog has become a focal point for a fight over the future of mathematics. After OpenAI released hundreds of mathematical manuscripts on October 6, the Association for Human Mathematics responded with a statement that Tao republished the following day. It challenged the release’s value and urged mathematicians to stop working with OpenAI.
My view is blunt: the boycott is the wrong response. The public case against AI mathematics too often mixes legitimate concerns about correctness and attribution with an assumption that established mathematicians should decide which discoveries may happen, when they may happen, and who may make them. That assumption deserves pushback of its own.
Mathematics belongs to everyone capable of contributing a sound argument. A company does not acquire mathematical authority by publishing a large repository. A prestigious group of mathematicians does not acquire ownership of the frontier by objecting to it.
There is a professional ego problem in the rhetoric of this dispute. That is an editorial judgment about a public claim to authority, not a finding about Tao’s private psychology. The strongest criticism concerns the demand for deference: discovery is being judged partly by whether it respects an incumbent community’s preferred process. We should judge a proposed theorem by its statement, evidence and proof, then judge its broader contribution by what people can learn and build from it.
Scope and disclosure: This is an argued opinion based on linked public sources, checked on October 8. It does not certify OpenAI’s collection or claim to know anyone’s personal motives. AI tools assisted research, drafting and publication preparation. Kingy.ai’s earlier local proof-checking pilot is cited with its limitations; no new mathematical experiment was run for this article.
Tao hosted the statement. AHM wrote it.
The distinction is essential. The October 7 post on Tao’s blog explicitly identifies itself as an AHM guest post, reproduced from the association’s website. The page’s byline says Terence Tao, but the statement is signed by AHM’s Communications Working Group. Hosting it gives the argument visibility. It does not establish that Tao wrote every sentence or personally adopted every demand.
The AHM statement says, “Mathematicians did not ask for this work to be done.” It criticizes OpenAI for proceeding despite advice against testing advanced problems on internal models, rejects the company’s characterization of the release as progress, describes the mass release as a “demonstration of power,” and calls for mathematicians to discontinue their work with OpenAI.
That is a substantial escalation from asking for more transparent benchmarks or cleaner citations. A boycott is a proposal to withdraw cooperation from an organization. It needs a case about consequences, proportionality and alternatives. An expression of professional indignation cannot supply that case by itself.
There are also several groups and documents here, with different positions. They should not be collapsed into a single anti-AI faction.
The September declaration on misalignment in AI mathematics, signed by Tao and other prominent mathematicians including Peter Scholze, Maryna Viazovska and Martin Hairer, argues that solving problems as benchmarks can undermine the development of ideas, students and understanding. It also recognizes AI’s potential to accelerate mathematical study. Its objection concerns the organization and purpose of research, rather than a simple assertion that machines cannot do mathematics.
The Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, includes Timothy Gowers, Hairer, Ravi Vakil, Edward Witten and others. Tao is not on its published membership list. It describes itself as independent, says members accept no payment for this work, and makes clear that it has no decision-making power at AI companies. Its October 6 response withholds endorsement while committing to continued engagement with frontier labs. AHM’s call to stop collaborating is therefore a distinct position.
Disagreement inside the profession matters. No association, advisory committee, celebrated blog or collection of signatures can stand in for every mathematician, let alone everyone who benefits from mathematical advances.
Mathematical discovery does not require a profession’s invitation
The sentence about mathematicians not asking for the work is the weakest part of AHM’s argument. Research does not normally require an invitation from the people already working in a field. A new proof can be unwelcome, inconvenient, inelegant or badly presented and still deserve examination.
Someone who attempts a public mathematical problem owes the world honesty about the result. They owe previous researchers accurate credit. If they announce a breakthrough, they should make the argument available and correct it when it fails. None of that creates a right for an incumbent group to prevent the attempt.
There is an important boundary here. Nobody is entitled to commandeer another person’s unpublished work, private correspondence or labor. A researcher can decline to referee a manuscript. An academic department can choose what it studies. An association can recommend a boycott. Those freedoms coexist with the freedom to investigate publicly stated problems using new methods.
The invitation argument confuses a community’s legitimate authority over its own participation with authority over the existence of discoveries. The first is indispensable. The second is gatekeeping.
AGMAI’s September 29 guidelines have a more specific target: advanced problem testing on proprietary models unavailable to the broader scientific community. They ask labs to stop that practice, while also recommending responsible, prompt release of significant results, funding for human understanding, better attribution and exposition, and broad access.
The access concern is serious. If a company has a powerful mathematical tool and everyone else has only its selected outputs, the company gains an advantage in choosing research directions and presenting its capabilities. That concern should lead to stronger access demands and better reproducibility.
I disagree with the proposed remedy of stopping advanced problem testing on those models. Developing a research tool requires discovering what it can and cannot do. A ban on testing the most consequential tasks can leave the public less informed about an emerging capability. The appropriate answer is disclosure, independent assessment and a path to wider access.
A temporary advantage deserves scrutiny. It does not turn a mathematical question into reserved territory. Nor does institutional prestige supply a veto: a proof does not become incorrect because its author failed to consult an advisory group, and consultation cannot make an incorrect proof true.
The critics are right that mathematics needs more than answers
The best argument on Tao’s side of this debate is about what makes a result useful. A theorem can be correct and still be difficult to understand, teach, generalize or connect to other work. A field can accumulate more claims while making disappointingly little progress in organizing them.
In Mathematics in the age of AI, Tao considers a future in which AI can perform substantial research tasks. He describes bottlenecks between generating proofs, verifying them, making them readable and absorbing them into mathematical knowledge. He also emphasizes that teaching, theory-building and sustaining a community are goals alongside problem-solving. This is a conditional analysis of how the profession should respond, not a blanket denial of AI capability.
That argument has a long intellectual history. William Thurston’s On proof and progress in mathematics, published in 1994, discusses mathematical progress that formal theorem proofs alone do not capture. Understanding involves methods, examples, communication and the ability to see why an argument works.
Consider a proposed proof that relies on a hundred unfamiliar intermediate claims. A checker might establish that the formal steps fit together. A researcher still needs to know which assumptions are essential, where the new idea lies, whether the method applies elsewhere, and how to explain it to students. Those are substantial mathematical tasks.
The mistake comes when this good description of intellectual work becomes a reason to resent the arrival of more material to understand. An incomplete process can have a valuable first step. A difficult theorem can become useful before it has a textbook treatment. The need for explanation supplies a reason to invest in explanation.
Human understanding also develops through more than one sequence. Sometimes a person grasps the idea and then proves it. Sometimes an example, computation or difficult proof arrives first, and people reconstruct the organizing idea afterward. Declaring the second sequence illegitimate would discard a useful way of learning.
The proof and its explanation can come from different contributors. That division of labor should receive accurate credit. The mathematician who extracts a reusable technique from an unwieldy argument may produce something more valuable than the original statement. A profession confident in the importance of understanding should reward that contribution directly.
Unsolved problems are not a fixed supply of jobs
One anxiety in the September declaration concerns how students and ideas develop through the process of solving problems. The concern is understandable. A student can spend months learning through an attempt, while a fast automated answer may change the project overnight.
But an unsolved problem is doing at least two jobs. It is a question about mathematics and a training opportunity. Those jobs can separate. We already teach students through problems whose answers are known. A theorem’s resolution does not erase the intellectual work required to learn its proof.
Research training differs from coursework because originality matters. Universities will have to adapt project selection, assessment and credit. That adaptation can be painful, especially for people halfway through a thesis. Departments should help them rather than pretend that a sudden solution has no personal cost.
Still, preserving uncertainty so a particular career structure can continue would be a poor research objective. A conjecture is not a job reservation. The public value of mathematics cannot depend on keeping answers unavailable until the preferred candidate finds them.
Newly verified results can also suggest new projects: weaken a hypothesis, classify the equality cases, simplify an argument, compare approaches, find a constructive version, or identify an application. I expect AI to make many of those tasks easier as well. That expectation creates a moving frontier, not a guarantee that every displaced project will be replaced conveniently.
A sensible training system should prepare students to formulate good questions and assess evidence under changing tools. Protecting them from a faster method would leave them less prepared for the work they will encounter.
The evidence for acceleration extends beyond OpenAI’s latest release
The forecast that AI will accelerate mathematics does not need all of OpenAI’s October manuscripts to survive review. There are earlier, more bounded demonstrations of useful capability. They involve different tools and tasks, so they should not be added together as if they were one benchmark.
AI can help generate mathematical insight
The 2021 Nature paper Advancing mathematics by guiding human intuition with AI reported work involving machine learning and mathematicians in knot theory and representation theory. The researchers used models to identify relationships and guide conjectures, including a connection between algebraic and geometric knot properties.
This matters because the benefit concerns the development of ideas. Pattern recognition can give a human a promising direction to investigate. A computer does not have to write an entire autonomous paper to improve the rate at which researchers learn something useful.
Search becomes productive when candidates can be checked
The 2023 Nature paper Mathematical discoveries from program search with large language models introduced FunSearch. It combined language-model-generated programs with evaluation and search, producing improved cap-set constructions and work on bin packing. These were specific advances, not a solution to every version of the cap-set problem.
The associated DeepMind explanation describes a workflow in which people supply the structure of a program and the model evolves a function within it. That arrangement illustrates a useful principle: unreliable proposals can still contribute when there is a dependable way to reject bad ones and retain good ones.
Reasoning performance has improved on bounded tasks
DeepMind’s AlphaGeometry announcement reported solving 25 of 30 historical Olympiad geometry problems under competition time limits. Its design combined a language model with symbolic deduction.
In July 2025, DeepMind reported that an advanced Gemini Deep Think system solved five of six IMO problems for a score of 35 out of 42, with official IMO grading and gold-medal-standard performance. The company said the system worked from natural-language problem descriptions within the competition time limit.
Olympiad performance does not certify research originality or mastery of every mathematical field. It does provide evidence of increasingly capable reasoning on tightly defined tasks. A serious account of AI’s trajectory should consider those gains alongside its failures.
Independent research evaluations show both useful proofs and failures
First Proof’s second-batch report tested four systems on ten research problems with previously unpublished solutions and expert refereeing. Across the systems, seven problems received at least one passing grade, meaning essentially flawless or requiring minor revisions. That is an aggregate result; it is not seven successes for each system.
The report also documents weak difficult steps and attribution failures. Its funding disclosures include unrestricted donations from OpenAI and Anthropic. The important feature for assessing the findings is the published methodology and referee evidence, rather than treating the project as exempt from scrutiny.
This is the sort of evidence the debate needs. A demonstrated success can be acknowledged without turning the tool into an oracle. A failed solution can be rejected without denying useful performance elsewhere.
The earlier FrontierMath paper proposed difficult research-level problems to measure mathematical reasoning beyond easier benchmarks. Such evaluations help distinguish actual task performance from an impressive anecdote. They remain measurements of specified tasks under specified conditions.
Taken together, these examples support a practical forecast. AI can reduce the cost of exploring possibilities, finding patterns, producing candidates and completing some proof tasks. Improvements at several stages can raise the amount of useful mathematics people accomplish, even if the gains arrive unevenly.
Tao’s own AI use undermines a simple anti-AI story
There is another reason to avoid caricature. In March, an OpenAI Forum account of Tao’s remarks described him using AI for literature searches, coding, calculations, figures and checking whether approaches are worth pursuing. It reported his view that current tools save more time than they waste in mathematics and theoretical physics. This is a company-hosted report of his remarks, so it deserves that attribution.
It nevertheless points to a straightforward mechanism for acceleration. A researcher who spends less time on routine computation or locating a relevant reference can spend more time investigating the difficult step. A lower cost of testing ideas can make exploration broader.
The October 8 guest post on Tao’s blog by Álvaro Lozano-Robledo also says students should learn about current LLM capabilities, while treating the extent of their use as a personal choice. It counsels continued mathematical study rather than surrender. Again, the guest author’s position should not be silently attributed to Tao.
These sources describe a profession debating how to adapt. They do not support the claim that Tao and everyone around him oppose mathematical AI. My objection is to the restrictive demands being advanced within that debate, particularly the jump from concern about the release to withdrawal from collaboration.
OpenAI has real obligations, and real errors to answer for
OpenAI’s October 6 announcement says it consulted AGMAI, published results and some formalizations, and plans to fund activities devoted to understanding AI-produced results. It also says the work came from an internal model. Consultation is not endorsement, and a funding announcement is not evidence that a particular workshop has already happened.
The repository README says results are at different verification stages and some unformalized results may have issues. At the October 8 snapshot checked for this article, it lists 719 manuscripts in 372 families and roughly 42% of top-line results formalized. Families group related papers; 372 families does not mean 372 independently validated famous conjectures.
The versioned October 7 change log is more revealing than a launch headline. It records three withdrawn manuscripts after a sign error undermined an argument and dependent work. It also lists 14 other manuscript revisions, plus updates to dependent citations. The original release had 722 manuscripts; the current catalogue’s 719 reflects those withdrawals.
Those corrections are meaningful. Dependency errors can spread, and a polished paper can conceal a weak argument. Readers should distinguish a released claim from a checked theorem. OpenAI should make uncertainty and dependencies conspicuous wherever it promotes these results.
Corrections also illustrate why releasing inspectable material matters. A public argument can be criticized, repaired or withdrawn. The relevant measure is the quality of the resulting knowledge and the accountability of the process. A high manuscript count proves neither quality nor uselessness.
OpenAI owes researchers precise statements, stable versions, honest credit, clear verification status and enough information to assess how the work was produced. Where a model remains inaccessible, independent evaluation and a credible access plan become more important. Researchers should press for those things, including when the company would prefer a simpler success story.
There is no need to pretend the company has earned blanket trust. Its incentives can favor dramatic capability demonstrations. That is a reason to inspect the evidence and reject overclaiming. It is also a reason to build evaluation capacity outside the company.
A Lean proof is powerful evidence, with a precise scope
Lean’s proof-validation documentation explains why a successful build is not the whole audit. A verifier should inspect the theorem statement and its axioms, including whether a proof relies on a placeholder or added assumptions. The checked statement must also correspond to the intended mathematical claim.
A formal proof establishes a statement within a specified logical environment. It does not automatically establish novelty, appropriate attribution, the accuracy of an informal manuscript’s explanation, or the usefulness of the result. Those require additional assessment.
That distinction should make the debate more exact. It is possible to have strong evidence for formal correctness and still have work to do on interpretation. Conversely, an impressive informal argument can remain unverified despite sounding persuasive.
Kingy.ai’s earlier local proof-checking pilot illustrates the limits of a bounded check. The Abhyankar–Sathaye formal target received Lean-kernel acceptance twice. Affine Bernstein and Pi Exponent each reached a 7 GiB memory cap twice without a proof verdict. No human mathematician reviewed whether the formal statements matched the manuscripts.
That gives evidence for one selected formal target under the recorded conditions. It cannot validate the entire release. The four resource-limited runs also cannot be described as four rejected proofs. Reporting those distinctions makes a claim more useful to the next investigator.
Better tools can let more people participate in this process. Independent checking should become easier, reproducible and affordable. An association that wants to protect mathematical trust could help establish shared checking infrastructure, without promising that automation removes the need for expert judgment.
The history of understanding argues for doing the work
The history of the Poincaré conjecture provides a useful comparison with clear limits. As the Clay Mathematics Institute records, Grigori Perelman announced its solution in preprints in 2002 and 2003. Other mathematicians subsequently invested substantial effort in examining and explaining the arguments.
John Morgan and Gang Tian’s 2006 Ricci Flow and the Poincaré Conjecture presented expanded versions of the arguments. Publication and deeper assimilation were separate contributions.
An AI manuscript should not inherit credibility from Perelman’s achievement. The comparison concerns the sequence of intellectual work: a result can create a demanding program of verification and exposition, and that later program can make an enormous contribution.
If AI produces more sound but difficult arguments, there is a reason to fund and recognize the people who make them understandable. The fact that explanation remains necessary is compatible with faster discovery. Both activities can grow.
A profession can choose to celebrate lucid explanations, useful definitions and well-designed teaching as explicitly as it celebrates first proofs. Doing that would make the profession’s stated commitment to understanding more credible than defending scarcity as a condition of progress.
The ego problem is a claim to authority over the frontier
It is easy to speculate that individual mathematicians are jealous or frightened. The public sources do not establish those motives, and speculation would weaken the criticism. The relevant ego is institutional: the idea that a field’s established leaders can speak for its proper future and treat departures from their preferred sequence as presumptively illegitimate.
AHM’s wording about not requesting the work is evidence for that reading. It puts the community’s invitation near the center of the complaint. The argument would be stronger if it identified specific invalid proofs, attribution failures, barriers to access and costs imposed on researchers, then connected those problems to a proportionate remedy.
Tao’s stature gives the hosted statement unusual reach. That increases the importance of reading its authorship accurately, and it makes the argument fair game for substantive criticism. Prestige can help people notice a concern. It cannot settle whether the proposed response would improve mathematics.
Professional communities have interests as well as expertise. AI companies have interests too. The right response is to examine both. Corporate marketing should not determine what counts as a breakthrough. Academic status should not determine which methods are allowed to pursue one.
My reading is that the boycott rhetoric grants too much weight to the profession’s sense of control. If the output is false, identify the gap. If it is derivative, establish the missing attribution. If it is sound and novel, ask how to make it comprehensible and broadly usable. Those responses serve knowledge directly.
Making cooperation socially suspect can do the opposite. It can discourage a junior mathematician from using a helpful tool or participating in evaluation, even when they disclose the work honestly. That is a foreseeable risk, not evidence that it has already happened in a particular department.
A boycott may deepen the imbalance it seeks to correct
A researcher has every right to refuse an arrangement they consider unethical. Collective pressure can also be justified when negotiations fail. The difficulty with AHM’s broad recommendation is that it does not establish why withdrawing mathematical expertise from collaboration is the best route to better output or wider access.
Experts who engage can insist on correction procedures, help design independent tests, expose misleading claims and improve the tools researchers actually use. That work need not confer approval on the company. Its value depends on transparent terms and the freedom to publish criticism.
If careful academics leave and less demanding evaluators take their place, the company may face weaker scrutiny. If the most capable tools remain inside companies, a boycott can also leave independent researchers further from the methods they need to assess.
Those are risks to weigh against the leverage a boycott might create. An effective proposal would name the changes sought, explain how withdrawal would cause them, and define the conditions for resuming cooperation. Without that, the call can become an expression of solidarity whose costs are borne by people trying to do useful research.
The advisory group’s stated willingness to keep engaging with frontier labs offers a more constructive starting point. Insist on access and accountability while keeping channels open for independent evaluation. Make criticism harder to evade by grounding it in inspectable evidence.
AI will accelerate mathematics if we invest in the whole process
I expect the long-run effect of AI on useful mathematical work to be strongly positive. That is a forecast, not a theorem, and it does not imply that every tool, lab release or university workflow improves productivity today.
The mechanism is straightforward. Cheaper exploration makes more ideas worth trying. More capable proof assistance can remove routine obstacles. Better formalization and checking can reduce uncertainty. Better explanations can make a method accessible to people who would otherwise never use it. These gains can reinforce each other.
There will be bottlenecks. A flood of unchecked submissions can waste referees’ time. A badly presented proof can impose costs on readers. A tool that confidently invents a citation can slow down the person checking it. Some projects and professional incentives will be disrupted.
Those problems are reasons to improve the process. They do not justify treating a higher rate of discovery as inherently harmful. The practical response should include independent verification, public correction records, access for researchers outside major labs, and funding for the work that turns arguments into understanding.
Labs should supply clearly labeled releases and enough provenance for assessment. Mathematical institutions should reward verification, exposition and the creation of reusable tools. Students should learn both the subject and the limitations of the systems they use. Funding bodies should support questions chosen by researchers, including questions that do not make attractive capability demonstrations.
None of this requires accepting a company’s account of its own greatness. It requires judging particular claims and designing institutions that can handle more evidence. A strong mathematical community should be capable of doing both.
Tao and other critics are right to demand human understanding. AHM is wrong to make the absence of an invitation part of the case against discovery. The next useful step is to check the claims, credit the ideas, open the tools and fund the explanations.
Sources and further reading
Primary statements, research reports and documentation are linked throughout. The OpenAI repository figures use commit fd4aeeb2ee4fc729c18d98444fed42fd0529eeeb, checked on October 8, 2026. The collection can change after publication.
- Tao’s repost of the AHM statement, October 7, 2026.
- Association for Human Mathematics statements.
- A Severe Misalignment of AI in Mathematics, September 11, 2026.
- AGMAI membership, independence and October 6 response.
- Responsible Release of AI-Generated Mathematics, September 29, 2026.
- Terence Tao, Mathematics in the age of AI.
- William Thurston, On proof and progress in mathematics, 1994.
- Davies and colleagues, Advancing mathematics by guiding human intuition with AI, Nature, 2021.
- Romera-Paredes and colleagues, Mathematical discoveries from program search with large language models, Nature, 2023.
- DeepMind’s FunSearch explanation, December 14, 2023.
- AlphaGeometry report, January 2024.
- Gemini Deep Think’s official IMO gold-standard report, July 21, 2025.
- First Proof Second Batch, methodology, findings and funding disclosures.
- FrontierMath benchmark paper, 2024.
- OpenAI Forum’s account of Tao’s March remarks.
- Álvaro Lozano-Robledo, What should we tell our students?, guest post, October 8, 2026.
- OpenAI, Sharing AI progress in mathematics, October 6, 2026.
- OpenAI’s mathematics repository at the checked snapshot.
- OpenAI’s withdrawal and correction history at that snapshot.
- Lean, Validating a Lean Proof.
- Kingy.ai’s bounded local proof-checking pilot, with public evidence and limitations.
- Clay Mathematics Institute, Poincaré Conjecture.
- Morgan and Tian, Ricci Flow and the Poincaré Conjecture, 2006.
Trending on Kingy
Keep reading with the stories getting the most attention now.
The Kingy Brief
Get The Kingy Brief.
AI changes, original tests and one practical thing to try. Fridays at 09:00 Vancouver time.
Free · Double opt-in · Unsubscribe anytime
Signup help and newsletter schedule
Signup form provided by Beehiiv. After submitting, check your inbox for "Confirm your subscription to The Kingy Brief" and open its confirmation link. Check Spam or Promotions if you cannot find it.
Fridays at 09:00 Vancouver time: source-checked AI changes, original tests and one practical thing to try. The weekly restart begins October 9, 2026. We skip a week when there is not enough verified material. Free. Unsubscribe anytime.
