~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / claude-riemann-hypothesis-67-2-percent $
News

Claude Went From 41.6% to 67.2% on the Riemann Hypothesis. And GPT-5.6 Sol Answered With 0.002%

LS Lucas Souza · · 10 min read
Claude Went From 41.6% to 67.2% on the Riemann Hypothesis. And GPT-5.6 Sol Answered With 0.002%

On August 10, 2026, Anthropic published an account that reads like a sci-fi script: an unreleased version of Claude raised the best known lower bound on the fraction of zeta function zeros that satisfy the Riemann hypothesis. From 41.6% to 67.2%.

The timeline did what the timeline always does. "AI solved Riemann."

It didn't. Not even close. And the next day came the cherry on top: a claim started circulating on r/singularity that a GPT-5.6 Sol user had squeezed another 0.002% on top of Claude's result.

Let's pull these apart. In this post you'll see what was actually proven, what wasn't, how the result was produced (60 subagents, 31 million output tokens, 2,400 shell commands) and why that "0.002%" is the best maturity test this market got all month.

TL;DR

  • What it is: an unreleased Claude improved a classic lower bound in analytic number theory — from 41.6% to 67.2% of the nontrivial zeros of zeta provably simple and on the critical line.
  • What it is NOT: a proof of the Riemann hypothesis. For that the number has to be 100%, and Anthropic itself says the technique won't get there.
  • How it was done: two Claude Code sessions, ~60 subagents reviewing each other, 2,400 shell commands, 31 million output tokens, 54 arXiv papers read to check novelty.
  • Verification: full paper, public Lean formalization, two in-house mathematicians (Levent Alpöge and Ralph Furman) and two external ones (Brian Conrey and Dan Goldston). No journal peer review yet.
  • Links: Anthropic's official announcement, paper as PDF, Lean formalization.

What Claude actually proved

The Riemann hypothesis says that every nontrivial zero of the zeta function has real part equal to 1/2 — they all land on the so-called critical line. Nobody has proven it since 1859.

What mathematicians can prove are fractions. "At least X% of the zeros are there." And that fraction crawled for a century:

Year Who Bound
1914 Hardy infinitely many zeros on the line (no proportion)
1974 Levinson > 1/3 (~33.3%)
1989 Conrey > 2/5 (40%)
2020 Pratt, Robles, Zaharescu and Zeindler 5/12 (~41.67%)
2026 Claude ~67.25%

Look at the gap between 1989 and 2020: thirty-one years to move 1.7 percentage points. Claude moved 25.6 points in one shot. According to the technical read of the paper, the exact number is 67.250% of zeros simple and on the line, plus a companion result of 83.625% of distinct zeros.

The mathematical trick, in dev terms: instead of treating the zeros on the line and the zeros off it as two separate problems, Claude packed everything into a finite Hermitian matrix and used Sylvester's law of inertia to count positive and negative signatures — meaning it counts how many zeros behave without having to locate a single zero. It's a change in how the problem is represented, not brute force.

And there's a ceiling. The paper itself says the approach saturates near 0.68185 with the pair correlation data available today. Hold on to that number, it comes back in a minute.

How the result was produced: 60 subagents and 31 million tokens

This is the part that matters if you write software.

In the first session, Claude generated and tested around 650 ideas. None of them worked. Zero. If you've ever run an agent on a long task, you know that smell: the model produces a lot, converges on little, and you sit there watching the bill climb.

In the second session things took a different shape. Around 60 subagents coordinated over 1.5 days, running 2,400 shell commands and hundreds of Python scripts, with the subagents reviewing each other's work and validating numerically against known zeta zeros before attempting any formal proof.

Notice the design: it's not a genius prompt. It's orchestration, context splitting, a cheap oracle before an expensive verifier, and cross-review. It's the same anatomy as a production multi-agent pipeline — just applied to analytic number theory instead of a backlog. Orchestrating dozens of subagents with stopping criteria and verification isn't a pretty prompt, it's software engineering. That's the kind of thing we build together at Beer And Code, with devs applying it to real products instead of watching videos about it.

If you want the pattern broken down with code, I've already written about subagents splitting context so you don't blow the token budget.

The news isn't the number. It's the verification stack.

What separates this result from yet another screenshot of an LLM hallucinating a theorem is what shipped alongside it:

  1. Continuous numerical validation — a hypothesis that doesn't match known zeros dies in seconds, before it costs any proof tokens.
  2. Lean formalization — there's a public repository and the proof passes the machine checker. Lean has no opinions and no manners: it either closes or it doesn't.
  3. Named reviewers — Alpöge and Furman internally, Conrey and Goldston from outside. Real names, reputations on the line.
  4. Novelty check — 54 arXiv papers downloaded and read to confirm the result hadn't been sitting published in some corner since 2011.
  5. Re-proof from scratch — Claude redid the argument from scratch, without the trail of the first derivation.

That's the template. If you build agents that produce artifacts of value (a migration, a refactor, a technical report, a financial calculation), the transferable lesson is this stack, not the number:

1. gerar hipótese
2. matar rápido com oráculo barato (teste numérico, lint, unit test, sanity check)
3. só o que sobreviver vai pro verificador caro (prova formal, suíte completa, revisão humana)
4. checar novidade/duplicação contra o que já existe
5. reproduzir do zero, sem o contexto da primeira tentativa
6. publicar o artefato verificável junto com o resultado

An agent without step 3 is a confident text generator. An agent without step 5 is luck.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

And the "0.002%" attributed to GPT-5.6 Sol?

This is where the maturity test comes in.

One day after the announcement, threads on r/singularity started repeating that a GPT-5.6 Sol user had improved Claude's result by 0.002%. Thousands of upvotes. A screenshot making the rounds. "Already beaten."

I went looking for the primary source. As of this writing, there is no public paper, no Lean formalization, no named reviewer, and not even clarity on what the "0.002%" means — whether it's an absolute percentage point (67.25% → 67.252%) or a relative improvement (67.25% → 67.2513%). Those are different things, and nobody spreading the screenshot seemed to care about the difference.

And remember the 0.68185 ceiling I asked you to hold on to? An improvement of that order of magnitude fits comfortably inside the fine-tuning of the optimization kernel that the method itself already allows. In other words: it's plausible. It's not absurd. It could very well be true.

Except "plausible" isn't "verified."

The difference between the two cases isn't the brand of the model. It's that one side published a paper, Lean, and reviewers, and the other published a screenshot. If the repository shows up tomorrow with the formalization closing, the result stands — whether it comes from GPT, Claude, Gemini or a PhD student with insomnia. The bar is the same for everyone, and it's exactly that bar the lab wars on your timeline are trying to make you forget.

Why "AI solved Riemann" is the wrong headline

Three reasons, and none of them is nitpicking.

1. 67.2% is not a progress bar. There isn't "32.8% left to go." The hypothesis demands 100%, and a single counterexample among the remaining zeros brings the whole thing down. Going from 41.6% to 67.2% doesn't put you "two thirds of the way there" — it leaves you exactly where you were relative to the original question: without an answer.

2. The technique has a mathematical ceiling. The method saturates near 68.185% with current data, and it works by population average — it's blind to sparse exceptional sets. Precisely where a counterexample would live. Anthropic itself wrote, in so many words: "We don't expect that the techniques Claude used will lead to proving the Riemann hypothesis."

3. It hasn't gone through real peer review. Two external reviewers looking at it at the last minute is a strong signal, but it's not the mathematical community's process. And the model used wasn't identified, the weights aren't available, and nobody can reproduce the run end to end. The result is auditable; the experiment isn't.

None of this diminishes the achievement. Diminishing it would be saying that an AI system just produced the largest single jump in the history of this bound and that it changes nothing. It does. It just doesn't change what the headline says it changes.

Limitations and things to watch

  • Ghost model. "Unreleased research version" means you can't test it, compare it or replicate it. Treat the result as a scientific publication with an open artifact, not as a product benchmark.
  • Hidden real cost. 31 million output tokens, 60 subagents and 1.5 days of runtime. That has a price, and Anthropic didn't disclose it. Before you pitch a "60-agent swarm" to your team, run the numbers for your own case — on most product problems, three specialized subagents beat sixty.
  • Verifiable domain. Formal math has Lean. What does your domain have? If the answer is "human review at the end of the pipeline," the pattern doesn't transfer directly — you need to build the verifier before you scale the generator.
  • Don't generalize the hit rate. 650 ideas failed in the first session. The pretty result came from a second attempt with a different harness. That's process re-engineering, not the model getting lucky.

Quick FAQ

Does this mean the Riemann hypothesis is 67% proven? No. It means it's proven that at least ~67.25% of the nontrivial zeros are simple and on the critical line. The hypothesis demands 100%, and the remaining 32.8% is exactly where a counterexample could be hiding.

Can the experiment be reproduced? Partially. The paper and the Lean formalization are public and verifiable by anyone. The original run isn't, because the model is internal and wasn't identified.

Is the GPT-5.6 Sol claim false? You can't say that. What you can say is that, so far, it hasn't come with a paper, a formalization or a reviewer. Without a verifiable artifact, it's a screenshot — no matter which logo is in the corner.

What do I take from this into my work tomorrow? The verification architecture. A cheap oracle before an expensive verifier, subagents reviewing each other, reproduction from scratch, and an auditable artifact published alongside the result. That's what separates an agent that delivers from an agent that impresses.

Conclusion

What happened on August 10 was too good to turn into team sports. An AI system, running in a well-designed harness, produced new mathematics and delivered it with a formalized proof, named reviewers and a novelty check. That's end-to-end engineering, not magic.

What came after — the "AI solved Riemann" headline and the 0.002% claim with no artifact at all — is market noise. And the noise will always be faster than the paper.

Your competitive edge over the next few years won't be knowing which model is ahead this week. It'll be knowing how to read what shipped alongside the result. If you built a verifier, you sleep fine. If you built a screenshot, you need another screenshot tomorrow.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing