Trending on Kingy
Keep reading with the stories getting the most attention now.
Anthropic says an unreleased Claude model pushed a century-old lower bound from 41.6% to 67.25%, then helped formalize the argument in Lean. Here is what was actually proved, how the AI research swarm worked, what remains unavailable, and why mathematicians should be impressed without confusing this for the Riemann hypothesis itself.
Published: August 10, 2026
Kingy AI verdict: Potentially a major theorem in analytic number theory, accompanied by unusually strong evidence for a same-day AI research announcement. It is not a proof of the Riemann hypothesis, has not yet passed conventional peer review, and cannot be reproduced end to end because Anthropic used an unidentified research model.
The short version
Anthropic’s headline is extraordinary. Its meaning is narrower than much of the online reaction.
In a paper released by Anthropic, Claude claims to prove that at least 67.250% of the nontrivial zeros of the Riemann zeta function are both simple and on the critical line, in the asymptotic sense used by analytic number theorists. The same paper claims that at least 83.625% of the zeros are distinct and gives analogous results for fixed primitive Dirichlet L-functions.
The previous unconditional critical-line record was slightly more than 5/12; the fraction itself is 41.666…%. Kyle Pratt, Nicolas Robles, Alexandru Zaharescu and Dirk Zeindler established it. If Claude’s proof survives expert scrutiny, the guaranteed proportion rises by about 25.6 percentage points in one step. That is a serious mathematical advance.
It is not 67.25% of the way to proving the Riemann hypothesis. It does not locate the remaining 32.75% off the line. It does not show that any zero lies off the line. It leaves open the possibility that every zero is on the line, as the hypothesis predicts, and the possibility of a very sparse exceptional set. Anthropic itself says the technique is unlikely to prove the full conjecture.
The result was produced through two Claude Code sessions, 31 million output tokens and a swarm of specialist agents. The decisive run coordinated roughly 60 subagents over about a day and a half. Anthropic published the manuscript, a shorter explanatory note, a large provenance appendix, process transcripts and a Lean 4 repository. That disclosure makes the claim much easier to examine than a normal model-demo anecdote.
The model, however, is described only as an unreleased research version of Claude. Anthropic has not named its checkpoint, published its weights or exposed the exact system publicly. The mathematics and formal code are public. The originating AI system is not.
What the Riemann hypothesis says, in plain English
The Riemann zeta function begins, for complex numbers with real part greater than one, as the infinite series
ζ(s) = 1 + 1/2ˢ + 1/3ˢ + 1/4ˢ + …
It can be extended to almost the whole complex plane. Its zeros include predictable “trivial” zeros at negative even integers. The mystery concerns the nontrivial zeros in the vertical strip between real parts zero and one.
The Riemann hypothesis says that every nontrivial zero has real part exactly 1/2. That vertical axis is the critical line. The zeros encode fine-grained information about how prime numbers are distributed. The Clay Mathematics Institute lists the conjecture as one of its seven Millennium Prize Problems, with a $1 million prize for a valid proof.
There are three very different statements that are easy to blur together:
| Statement | Status after Anthropic’s announcement |
|---|---|
| Many computed zeros lie on the critical line | Known for enormous finite ranges; computation cannot settle all zeros |
| A positive proportion of all zeros must lie on the line | Proved for decades; Claude claims the lower bound is now 67.25% |
| Every nontrivial zero lies on the line | The Riemann hypothesis; still open |
The word “proportion” also needs care. The new theorem is an asymptotic lower bound. As the height of the search region tends to infinity, the lower limit of the relevant ratio is at least 0.67250. It is not a survey result saying mathematicians inspected all zeros and found that fraction.
Computers have already rigorously checked every zero up to height 3 × 10¹², finding all of them simple and on the line in that finite range. Claude’s theorem addresses an entirely different scale: zeros at arbitrarily great heights, through an asymptotic guarantee. Finite computation can report 100% below a cutoff while a density theorem proves 67.25% across the unbounded population.
Even a proof that 100% of zeros lie on the critical line in density would not automatically prove the Riemann hypothesis. A sparse infinite set can have density zero, just as the perfect squares become a vanishing fraction of the positive integers while never running out. “All but a density-zero set” and “all” are logically different claims.
What Claude appears to have advanced
The manuscript, titled More Than Two Thirds of the Zeros of the Riemann Zeta Function Lie on the Critical Line and credited to Claude, states three headline bounds:
-
At least 67.250% of all nontrivial zeros are distinct zeros on the critical line.
-
At least 67.250% are simple zeros on the critical line.
-
At least 83.625% of all zeros are distinct, whether on or off the line.
The paper also claims separate results about zeros of the derivative of the completed zeta function, ξ′: at least 86.864% are simple and on the line, while 93.432% are distinct. Those figures do not describe zeros of ζ itself, so neither should replace 67.25% in the headline.
For comparison, Xiaosheng Wu’s 2015 record for distinct zeta zeros was more than 66.036%. Claude’s claimed 83.625% bound is therefore a substantial secondary advance, independent of the more visible critical-line record.
“Simple” means the zero has multiplicity one. A multiple zero is counted repeatedly when zeros are counted with multiplicity; a simple zero is counted once. The Riemann hypothesis is commonly paired with a further conjecture that the nontrivial zeros are simple, so the first two bounds address both location and multiplicity.
The historical comparison is stark:
| Milestone | Unconditional lower bound on zeros on the critical line |
|---|---|
| Hardy, 1914 | Infinitely many, not yet a positive proportion |
| Selberg, 1942 | A positive but unspecified proportion |
| Levinson, 1974 | More than one third |
| Conrey, 1989 | More than two fifths |
| Pratt–Robles–Zaharescu–Zeindler, 2020 | More than 5/12, or 41.666…% |
| Claude manuscript, 2026 | At least 67.250% |
The prior record appears in the Pratt–Robles–Zaharescu–Zeindler paper. Earlier landmarks are surveyed in Brian Conrey’s overview of the Riemann hypothesis. Records are not all direct refinements of one identical technique, but the table shows why the claimed jump commands attention.
There is an important human prehistory to the result. A 2023 preprint, published in 2024 by Siegfred Baluyot, Daniel Goldston, Ade Irma Suriajaya and Caroline Turnage-Butterbaugh, established the unconditional pair-correlation input. Follow-up work extracted horizontal information and reached a two-thirds conclusion under a narrow-box condition controlling zeros close to the critical line. Claude’s contribution, as the paper presents it, is a way to remove that extra location assumption by reading the same information through an indefinite Hermitian form.
That distinction matters. Claude did not conjure the result out of an empty room. It found a new bridge between a mature body of human mathematics and a finite-dimensional linear-algebra argument. Original mathematics often works exactly this way: the novelty is in the connection.
How the proof works, without pretending it is simple
The central problem is that zeros on the critical line and zeros off it behave differently under the zeta function’s symmetries. Off-line zeros arrive in reflected configurations. Claude’s method packages those contributions into a matrix whose positive and negative directions can be counted.
Here is the conceptual path.
First, the proof starts with Weil’s explicit formula, a family of identities linking sums over zeta zeros to sums involving primes. A carefully chosen test function lets a mathematician look at local relationships among zeros while still evaluating the relevant aggregate quantities from the prime side.
Second, the argument compresses that information into a finite-dimensional Hermitian quadratic form and its Gram matrix. Each zero on the critical line contributes a positive rank-one piece. A symmetric pair of zeros away from the line contributes an indefinite block with one positive and one negative direction.
Third, Sylvester’s law of inertia makes the signature useful. The counts of positive and negative directions do not change under a suitable change of coordinates. The proof can therefore separate how much matrix rank must come from on-line zeros even though it does not know the location of each zero individually.
Fourth, a new rank–trace inequality converts the matrix’s trace and squared size, measured through its Frobenius norm, into a lower bound on that on-line rank. Pair-correlation results supply the necessary first- and second-moment information. In less technical terms, the proof does not identify the zeros one at a time; it shows that the collective statistics cannot be produced unless enough of them sit on the line.
Finally, the test window is optimized. A basic choice gives the clean threshold two thirds. A Montgomery–Taylor-style kernel raises the numerical constant to just over 0.67250.
The manuscript is unusually candid about the method’s ceiling. With the currently available “bandwidth one” pair-correlation information, the approach tops out at roughly 0.68185. Reaching lower bounds of 70%, 80% or 90% would require pair-correlation data on wider Fourier support, approximately 1.04, 1.26 and 1.70 respectively. Those are new mathematical obstacles, not settings that can be unlocked by spending another million model tokens.
This is also why Anthropic says the route is unlikely to prove the Riemann hypothesis. The inputs average over zeros and are insensitive to a sufficiently sparse exceptional set. Comparable statistical statements can hold for related families in which the generalized Riemann hypothesis is false. The method’s strength is unconditional population-level control; its weakness is the inability to eliminate every possible exception.
How Claude found it: not one prompt, but a research organization
Anthropic’s research announcement describes a process closer to a compressed institute than a spectacular chat response.
Across a failed first Claude Code session and a successful second one, the campaign produced about 31 million output tokens. The first generated roughly 650 ideas and found no proof. The second spanned 54 hours across three calendar dates; Anthropic summarizes its intensive coordination phase as a day and a half. One coordinating instance assigned jobs to roughly 60 Claude subagents. Anthropic reports 2,400 shell commands, hundreds of Python programs and thousands of numerical checks.
The published role accounting is revealing: two agents supplied key ideas, 13 developed supporting ideas, 30 followed failed paths, 13 acted as validators and two wrote the paper. One false start tried to extract information from a matrix’s negative index. Its dual formulation produced a one-half bound. The rank–trace observation then broke the barrier to two thirds, after which test-function optimization yielded 67.25%.
Jarred Sumner, an Anthropic researcher who is not a professional mathematician, orchestrated the run. His intervention was often motivational rather than technical: keep going, try another route, believe the problem may be tractable. That does not make the result a one-line prompting trick. It shows that persistence and resource allocation can be valuable control signals when the model has enough mathematical and computational competence to act on them.
The agents also played adversarial roles. They tried to construct counterexamples, checked limiting cases, wrote numerical experiments, downloaded 54 arXiv papers during later literature work and independently re-derived the proposed result. Anthropic mathematicians Levent Alpöge and Ralph Furman then studied and validated the argument. Established analytic number theorists Brian Conrey and Daniel Goldston read the manuscript on short notice and supplied comments.
The phrase “on short notice” should stay attached to those names. It is evidence of serious expert engagement, not evidence that a journal’s refereeing process has concluded. On publication day, Kingy AI found no public independent referee report or broad expert consensus. The correct status is a high-evidence preprint claim awaiting community scrutiny.
What is public, and what is not
Anthropic has released more than the polished paper:
| Artifact | Public? | What it lets outsiders check |
|---|---|---|
| Main mathematical paper | Yes | Definitions, theorems and the complete written argument |
| Concise informal note | Yes | A faster conceptual explanation of the proof |
| Provenance appendix | Yes | Agent roles, discovery chronology, dependencies and validation trail |
| Process transcripts | Selected records only | The two decisive subagent runs plus excerpts from the coordinator |
| Lean 4 formalization | Yes | Machine checking of the formal theorem chain against Lean/Mathlib |
| Exact Claude model/checkpoint | No | Anthropic says only “unreleased research version” |
| Model weights or public endpoint | No | Outsiders cannot rerun the same system |
| Complete hidden reasoning and metadata | No | Extended thinking is not exposed; model identifiers and token counters were removed |
The Lean repository is a substantial point in Anthropic’s favor. Its audit states that the main theorem files are free of sorry, Lean’s placeholder for an unproved step, and use only the standard logical assumptions inherited from Mathlib. The project also publishes comparator configurations and reports an independent kernel replay. Kingy AI inspected the trusted definitions, theorem statements, repository tree and audit at the release commit; we did not independently build the project.
Formalization changes the evidence, but it does not end the discussion. Lean checks that conclusions follow from formal statements under a trusted kernel. Humans must still confirm that the formal statements faithfully capture the intended analytic theorem, that imported definitions match the paper and that no gap has been hidden in an incorrectly modeled assumption. Formal proof and expert semantic review are complementary.
The process record has limits too. Anthropic’s provenance appendix says the core discovery agents made no network requests and relied largely on mathematical knowledge recalled from training. Later agents searched the literature and positioned the result against prior work. The public transcript volume covers the two decisive subagents, E2 and E2-pairs, plus selected coordinator excerpts; it is not a dump of the whole campaign. Hidden thinking appears only as silent time intervals, and the released records remove envelope metadata including the exact model identifier. Most tool calls are summarized rather than reproduced.
Kingy AI also found a small documentation discrepancy. Appendix B.2 of the paper says a 31-check SymPy reference script is included in the Lean repository at tag v1.0. The complete file tree at that tag contains no Python or SymPy script. This is a reproducibility gap, not evidence of a defect in the Lean proof, but it is exactly the sort of release detail outside reviewers should record rather than silently overlook.
So the result is unusually auditable, but not fully reproducible as an AI experiment. Any researcher can read the proof and clone the Lean code. Nobody outside Anthropic can recreate the same 31-million-token model run today.
Was this “Claude,” or was it Anthropic’s human team?
The honest answer is both, in distinguishable roles.
The manuscript names Claude as author. The provenance record attributes the two central mathematical ideas to Claude agents, along with most experiments, failed approaches, internal criticism and drafting. That is stronger agency than autocomplete or a conventional computer algebra system.
Humans selected the problem, built the orchestration environment, authorized large compute, urged the system to continue, arranged literature and formalization work, checked the output and decided what to publish. Sumner is described in the paper’s acknowledgements as, in a meaningful sense, a human co-author. Alpöge and Furman provided mathematical validation. Eric Easley coordinated the Lean formalization. Conrey and Goldston provided expert comments.
Reducing this to “AI solved it alone” erases the research apparatus. Reducing it to “humans used a fancy calculator” ignores where Anthropic says the decisive bridge and rank–trace argument originated. A more accurate unit of analysis is the human-directed AI research system: model, agent scaffold, tools, compute budget, validation roles and expert review.
That framing will matter well beyond this result. Kingy AI has argued in its 2026 state of AI agents that agent performance increasingly depends on workflow design and verification, not just the base model. This case is an unusually vivid mathematical example.
What it means for mathematics
If experts confirm the proof, mathematics gains a strong unconditional theorem and a reusable proof strategy.
The number 67.25% is important, but the finite-compression and inertia method may be more durable. It offers a new way to turn averaged zero statistics into location and multiplicity information. The paper’s extension to primitive Dirichlet L-functions suggests that the technique is not a one-off numerical trick attached only to ζ(s).
The work may also shift how mathematicians divide research labor. A system that can explore dozens of branches, retain failed paths, run numerical sanity checks, trace dependencies and dispatch critics changes the economics of speculative work. Many ideas in hard mathematics die after hours of algebra or because their relation to the literature is unclear. Parallel agents can make those deaths cheaper and can sometimes recover a useful dual statement from a failed attack, as happened here.
None of that removes the need for mathematicians. It changes the bottleneck. Problem selection, taste, conceptual interpretation and source credit become more valuable. So do theorem-statement audits and independent attempts to break a proof. A formalization can absorb clerical verification, but someone still has to decide whether the formal theorem is the theorem people think they are celebrating.
There is also a credit question. The result sits on decades of human work and on contemporary pair-correlation advances that made its prime-side inputs possible. Future publication norms will need to represent model-originated ideas without turning the training literature or human orchestration into invisible infrastructure.
What it means for AI
This is one of the clearest demonstrations yet that a general-purpose language-model system can participate in open-ended mathematical research rather than merely solve a curated contest problem.
That qualification matters. Systems such as DeepMind’s AlphaProof and AlphaGeometry reached silver-medal-level performance at the 2024 International Mathematical Olympiad, and an advanced Gemini Deep Think system later achieved gold-medal-standard performance. Those were landmark results on problems with known, compact solutions and a clear success condition. Anthropic’s case involves literature navigation, conjecture formation, dead ends, code, internal reviewing and a claim that was not known beforehand.
It also punctures the idea that this capability lives inside a single magical response. The system consumed 31 million output tokens. Most agents did not find the answer. The productive shape was a portfolio: many inexpensive failures, a few feeding ideas, two key discoveries and a separate verification layer. This is evidence for scaling inference-time research effort, not evidence that every Claude chat can now advance number theory.
The unavailable model weakens the scientific lesson. Without its identity, weights, prompts in complete form and hidden reasoning, outsiders cannot measure how often the process succeeds, compare it fairly with another model or distinguish model improvement from orchestration and sheer sampling. Anthropic has published a compelling case study, not a controlled benchmark.
For organizations thinking about agent deployments, the practical lesson resembles the one in Kingy AI’s AI agent adoption playbook: separate generation from checking, preserve provenance, define escalation points and spend most of the trust budget on evidence. A confident final answer is the least interesting part of this system. The ledgers, counterexample searches, independent derivation and formal checker are what make it credible.
What happens next
The immediate task belongs to analytic number theorists. They will check the explicit-formula normalizations, the passage from local matrices to asymptotic counts, the rank–trace bounds, the optimization and the Dirichlet L-function extension. Formal-methods experts will inspect the theorem statements and dependency boundary in Lean. Historians and authors of the cited papers will assess priority and whether the manuscript describes earlier results accurately.
A confirmed theorem will not trigger the Clay prize, rewrite cryptography overnight or reveal a hidden list of primes. It will raise a famous lower bound, introduce a promising method and provide a striking proof of concept for AI-assisted discovery.
A discovered gap would also be informative. The publication package is detailed enough for critics to identify where the system failed, and the formalization creates a precise target for semantic challenges. AI research needs falsifiable cases alongside polished success stories.
The sober conclusion is still remarkable: an unreleased Claude system appears to have generated the key ideas for a credible, formally checked advance on a problem family that has occupied mathematicians for more than a century. It did so through massive parallel search and structured criticism, not a flash of chatbot omniscience. The Riemann hypothesis remains exactly where it was: unproved. The boundary of useful AI research may have moved.
Frequently asked questions
Did Claude solve the Riemann hypothesis?
No. The hypothesis says every nontrivial zeta zero lies on the critical line. Claude’s paper claims an unconditional lower bound showing at least 67.25% are simple and on that line. It leaves the rest unclassified and explicitly says the method is unlikely to prove the full hypothesis.
What was the previous record?
The previous published unconditional lower bound was more than 5/12, or about 41.67%, from Pratt, Robles, Zaharescu and Zeindler. Recent human research had reached two-thirds under additional narrow-box assumptions; Claude claims to remove that assumption.
Which Claude model made the discovery?
Anthropic has not said. Its announcement calls it an unreleased research version of Claude. The public transcript removes the model identifier, and no matching checkpoint or public API is available.
Can anyone inspect the proof?
Yes. Anthropic released the paper, informal note, provenance appendix, process transcripts and a Lean 4 formalization. That makes the mathematical claim highly inspectable. Reproducing the original AI search is not currently possible.
Has the result been peer reviewed?
Not in the conventional journal sense as of August 10, 2026. Anthropic mathematicians studied the proof, Brian Conrey and Daniel Goldston read it on short notice, and Lean checks the formal derivation. Those are meaningful safeguards, but independent community review is only beginning.
Does 67.25% mean the other 32.75% are off the line?
No. A lower bound classifies at least that many. The remaining zeros may also be on the line; the method does not establish where they are.
Is the result useful outside the zeta function?
Potentially. The manuscript states analogous bounds for fixed primitive Dirichlet L-functions, and its matrix-inertia technique may inspire other ways of extracting location information from averaged spectral data. Those extensions need the same expert scrutiny as the headline theorem.
Sources and disclosure
Kingy AI independently reviewed Anthropic’s announcement, the 35-page mathematical manuscript, the informal note, the provenance appendix, the published transcripts and the Lean repository. We did not receive access to the unreleased model, did not reproduce the 31-million-token run and did not independently build the Lean project. Anthropic’s descriptions are identified as such; release-day validation is not presented as settled consensus. Sources were checked on August 10, 2026.
- Anthropic: Learning more about Claude’s mathematical capabilities
- Claude: More Than Two Thirds of the Zeros of the Riemann Zeta Function Lie on the Critical Line
- Anthropic’s concise informal note
- Anthropic’s provenance appendix
- Anthropic’s process transcripts
- Anthropic’s Lean 4 repository
- Pratt, Robles, Zaharescu and Zeindler: More than five-twelfths of the zeros of ζ are on the critical line
- Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh: unconditional pair correlation
- Platt and Trudgian: rigorous verification of RH up to 3 × 10¹²
- Xiaosheng Wu: Distinct zeros of the Riemann zeta-function
- Clay Mathematics Institute: Riemann Hypothesis
- Brian Conrey: The Riemann Hypothesis
