AI News

OpenAI’s Math Advisory Group: Who Controls the Future of Mathematics?

Mathematicians will advise OpenAI on the release of its mathematical discoveries. The company will retain responsibility for the decisions.

That arrangement is the most consequential detail in the new Advisory Group on Mathematics and Artificial Intelligence. Its members can challenge a company, publish recommendations, and advocate for the mathematical community. They have no decision-making power at any AI company. The group states that limitation on its own website. Read its terms and current task.

The stakes have risen with OpenAI’s claim that an internal model has resolved more than 100 long-standing open problems. Its September 21 announcement also excludes advice on the pace of internal mathematical progress from the group’s remit. OpenAI’s announcement.

The arrangement creates a way for mathematicians to influence decisions. It leaves the more difficult question open: how much influence will public expertise have over discoveries generated inside private research systems?

Analysis published September 21, 2026. The group’s stated powers are distinguished from our recommendations for how this arrangement should work.

What the OpenAI math advisory group can actually do

The group operates independently, its members accept no payment for this work, and it intends to publish recommendations. It can advise other AI companies as well as OpenAI. Its immediate assignment is helping coordinate the release of a large collection of significant results that OpenAI reports its internal model has produced.

The initial members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten, and Melanie Matchett Wood. The group’s membership and independence statement.

Those terms matter. An independent group able to make criticism public has more scope to challenge a company than a panel whose work is entirely private. Advising multiple laboratories could also help establish expectations that outlast a single company’s release schedule.

But an advisory relationship works through persuasion, expertise, and reputation. A recommendation is not a veto. Readers should assess what the company does with the advice, as well as the qualifications of the people offering it.

There are several different kinds of control

“Who controls mathematics?” can sound absurd. No company can order a valid theorem to become false. Yet there are practical decisions around discovery that have little to do with mathematical truth itself.

Here is a framework for evaluating the arrangement. The distinctions are our analysis of the decisions involved, not additional powers claimed by the group.

Decision Why it matters What meaningful scrutiny would require
Which problems receive compute It shapes the discoveries a research system attempts An account of selection criteria and the purpose of the research program.
Who can use the strongest tools It affects which researchers can compete or collaborate Clear access terms, available capabilities, and known restrictions.
When a result becomes public It affects review, attribution, and opportunities for follow-up A release rationale and enough material for independent assessment.
How a result is described It shapes public understanding of capability and significance Precise statements, prior-work comparisons, and visible corrections.
Whether the argument is accepted It determines its place in the mathematical literature Scrutiny of the actual mathematics, independent of the announcing brand.

These decisions interact. A company could publish a proof openly while keeping the system that found it private. That would allow scrutiny of the result without giving other researchers the same means of discovery. Conversely, broad model access would not by itself guarantee that a particular claimed result had been checked carefully.

Open publication and equal research access therefore deserve separate questions. Treating either one as a complete answer would leave important parts of the arrangement unexamined.

Why mathematicians object to becoming an AI benchmark

The September 11 declaration, A Severe Misalignment of AI in Mathematics, challenges the use of open mathematical problems as benchmarks for AI systems. Its authors argue that rapid production of answers can damage the slower work of developing understanding, crediting earlier contributions, and training students. Their concern is that the incentives of AI companies can diverge from the purposes of mathematical research.

That argument should be taken seriously without assuming that every AI-assisted discovery is harmful. A correct, important result is valuable. The issue is how its production and release affect everyone who must understand it, evaluate it, or build on it.

Consider a hypothetical research program in which a new method has finally made a family of problems tractable. A laboratory can use a model to finish several of them and announce the total. The count may accurately describe completed results while telling readers little about how much of the essential groundwork came from prior research.

A useful publication would identify both contributions: the earlier method and the new work required to complete the argument. A publicity narrative built only around the final answer could obscure that distinction, even if every resulting theorem were correct.

This is why attribution belongs in the technical account from the beginning. It should be possible to ask what the system contributed without either erasing the researchers who built the subject or pretending that an original machine-generated step was human work.

The authorship disagreement is already explicit

The Leiden Declaration recommends keeping authorship, credit, and responsibility with humans, alongside transparent disclosure of automated tools. OpenAI’s August ten-result announcement takes a different emphasis: it says attributing wholly AI-generated proofs to humans would misrepresent the origin of the arguments, while accepting responsibility for correctness and describing human involvement in preparing manuscripts. OpenAI’s explanation of attribution.

That is a substantive disagreement about what authorship should communicate. It will not be resolved by inserting the phrase “AI-assisted” into every paper.

Our preferred approach is to make the contribution account specific enough that a reader does not have to infer the workflow from the byline. Identify who selected the problem, supplied mathematical ideas, generated the argument, wrote the exposition, performed formalization, and reviewed the result. Where the record cannot support a confident attribution, say so.

A contribution account also needs a responsible contact who can answer questions and correct the record. Naming a model explains part of a process; it does not establish who will respond when another mathematician discovers an ambiguity six months later.

Releasing everything immediately has costs. So does waiting.

There is a strong case for prompt publication. Other researchers can inspect the work, identify mistakes, and pursue applications. Keeping useful results private can create an information advantage for the people who have seen them. A delay that appears reasonable inside a laboratory may be expensive for an outside team unknowingly working on the same question.

There is also a strong case for preparing a release properly. A large batch of difficult arguments can impose substantial review work. Without clear statements, usable files, and careful references, public availability can be more nominal than practical. Uploading material is not the same as making it intelligible.

The correct policy need not be identical for every result. A short counterexample that specialists can assess quickly presents a different release problem from a long argument spanning several technical subjects. Sensible coordination should respond to the work involved, with an explanation of any delay.

The earlier unit-distance release shows one useful ingredient: it included an account by external mathematicians alongside the AI-generated work. That offers readers more context than a capability claim alone. It does not establish that the same release process will fit every new result.

Our standard would be straightforward: make review easier and explain decisions that restrict access. “Responsible release” should be a description of what was done, with reasons a reader can evaluate.

Five things the next release should make public

The advisory group has an opportunity to turn a broad consultation into observable improvements. Here are five concrete tests we would apply to the forthcoming material. These are Kingy.ai’s recommendations, not commitments already made by OpenAI or the group.

1. A result-by-result inventory. Give each claimed advance a stable identifier, its exact mathematical statement, the prior best result, and links to the manuscript and any formalization. State whether the result resolves a problem, improves a bound, supplies a counterexample, or settles only a special case. An aggregate total cannot carry that information.

2. A review record with a defined scope. Name reviewers who consent to be named and explain what they assessed. Checking a central lemma, reading an entire argument, and rebuilding a formal project are different activities. A brief statement of scope is more useful than an unexplained assurance that “experts verified it.”

3. A reproducible formal-proof package where applicable. Publish the necessary files, dependencies, version information, and instructions. Map the principal formal statements to the claims in the paper. Lean’s documentation makes clear that its logical foundations and additional axioms are part of the system; a reviewable package should make the relevant assumptions visible. Lean documentation.

4. A contribution and correction history. Record the roles of humans and models, relevant prior work, and substantive revisions. A corrected statement should be easy to distinguish from the original. If an announcement changes, readers who encountered the earlier version should be able to discover what changed and why.

5. The company’s response to the advice. Publish which recommendations were accepted, modified, or declined, with reasons. Independence becomes easier to evaluate when readers can see where the group and the company disagree. A group may offer excellent advice while a company makes a different decision; the record should let the public identify both.

These requests would create work. The burden should be proportionate to the importance and complexity of each claim. But the alternative is to transfer the cost of reconstructing the evidence to every outside reader, repeatedly.

Who gets access after the announcement?

The release of proofs will answer only part of the access question. Researchers will also want to know whether they can investigate their own problems with tools of comparable capability.

It would be premature to attach a price to that future or assume that an internal research system is already available in a public product. The useful questions concern concrete access terms: which capabilities are offered, what limits apply, how unpublished research is handled, and whether researchers can preserve a record of their work when the service changes.

Those details determine whether a promising collaboration is usable outside a small set of well-connected institutions. A grant of access can be valuable while still leaving uncertainty about continuity. A published proof can be valuable while leaving a substantial gap in who has the resources to find the next one.

The Leiden Declaration recommends independent public research laboratories and support for less resource-intensive tools. Such alternatives would give the mathematical community more options when negotiating partnerships. They need not reproduce every frontier capability to contribute useful infrastructure, expertise, or bargaining power.

How to judge whether the group makes a difference

Membership is the beginning of the story. The test will be the public record that follows.

Do the releases become easier to inspect? Are prior contributions explained more precisely? Can readers identify the scope of external review? Are mistakes corrected visibly? Does the company explain what it did with the group’s recommendations?

The answers could establish an effective form of independent influence without granting the group formal control. They could also show that consultation leaves the consequential decisions largely unchanged. Either conclusion should rest on actions over time.

Meanwhile, advice from the group must not be mistaken for a blanket certificate of correctness. The acceptance of a mathematical argument depends on the argument and its scrutiny. For Millennium Prize claims, Clay also maintains a separate process; its Navier–Stokes page remains marked “Active” at publication. Kingy.ai’s Millennium guide and breakthrough tracker distinguish those evidence states.

A useful next milestone would be a public, itemized release accompanied by the group’s recommendations and OpenAI’s response. That would give mathematicians and readers something concrete to evaluate beyond the announcement of an advisory relationship.

Read the companion analysis: The End of Mathematics? OpenAI Says Its AI Solved 100+ Problems