Models

An Anthropic Model Pushed the Riemann Hypothesis Forward

An unreleased Anthropic model ran 60 subagents for 36 hours and widened the range where the Riemann hypothesis is known to hold. Not a proof, but the method matters.

Faruk TalmaçAugust 12, 20263 min read3 views
An Anthropic Model Pushed the Riemann Hypothesis Forward

650 ideas. 60 subagents. Roughly 36 hours of runtime and 31 million output tokens. That is what an unreleased Anthropic model spent on a problem mathematicians have been circling since 1859.

According to TechCrunch on August 11, the model raised the lower bound of the range in which the Riemann hypothesis is known to hold. The verified region got bigger. Two mathematicians on Anthropic's staff checked the result, and the argument was formalized in Lean, the open source proof assistant.

The hypothesis itself describes how prime numbers are distributed along the number line. A large amount of modern mathematics is already built on the assumption that it is true, so proving it would confirm hundreds of existing results rather than open a new field.

Half the subagents got nowhere

The result matters less than the process that produced it. The person who set the model going was an Anthropic employee without a deep mathematics background. The model then divided the work itself.

Of the 60 subagents, only two produced the key ideas. Thirteen contributed supporting work, thirty failed to reach anything useful, thirteen checked the arguments that came back, and two wrote up the paper. Half the effort was discarded, and that was how the method was supposed to work.

A bound, not a proof

This is where coverage tends to overreach. The Riemann hypothesis has not been proved. The Clay Mathematics Institute's million dollar prize is still unclaimed. Widening the verified range is a real contribution, but it does not close the problem.

The model has not been named either. It is a version Anthropic has not shipped, so there is nothing here you can go and try today.

The Lean detail deserves more attention than it usually gets. Language models are extremely good at producing arguments that sound correct. Lean forces the argument into a form a computer can check line by line, which closes the gap between convincing and verified.

Mathematicians are split on the credit question

The Leiden Declaration, published in June 2026, warned that AI could erode the attribution norms mathematics runs on. Fields Medalist Timothy Gowers pushed back, arguing that a future where theorems are no longer tied to individual names is not automatically a bad one.

The argument has moved on from whether a model can do mathematics to whose name goes on the result when it does.

Why an operations team should care

The Riemann hypothesis is not on your roadmap. The working pattern behind it probably is, or will be within a year.

Instead of asking one model a question and taking the answer, this run split a task across dozens of agents, accepted that most would fail, and paid for a separate layer whose only job was checking the work. Most companies experimenting with AI today are still at the first stage, asking a single model and hoping. The value is accumulating with the teams that build the verification layer.

Then there is the bill. Thirty-one million output tokens on one problem is a serious spend, and agent architectures are not cheap to run. When you scope one, the opening question should not be whether the system can do the task. It should be what a verified, finished result costs per run.

Sources: TechCrunch, Clay Mathematics Institute

Share This Article

Faruk Talmaç

Written by

Faruk Talmaç

Co-Founder & Editor

Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk