Seven hundred and twenty-two manuscripts, 372 result families, and a GitHub link dropped on the global mathematics community like an unsolicited term paper stack. OpenAI published the collection from an internal model it has not released to the public, spanning number theory, algebra, geometry, topology, logic, theoretical computer science, and mathematical physics.
Someone has to verify all of it.
What OpenAI Actually Claimed
The specific results are ambitious; broad independent confirmation has not yet been established.
Among the reported claims: a solution to the four-dimensional Kakeya conjecture, progress toward a “quasi-Riemann hypothesis,” and work touching the Hodge conjecture and the Unique Games Conjecture. To be precise about the Riemann item, the result reportedly concerns a zero-free region with a real part greater than 7/8, which is a meaningful but weaker statement than the classical Riemann hypothesis, which involves nontrivial zeros lying on the line with real part 1/2.
Broad independent confirmation of these claims has not yet been established by the wider mathematical community. The model was reportedly posed roughly 4,000 problems and used, on average, about three hours of ChatGPT Pro compute per result.
What “Verified” Actually Means Here
What Lean formalization confirms and what it does not are two very different things.
Many of the manuscripts reportedly include Lean formalizations, and the word “verified” is doing a lot of heavy lifting in coverage of this release. Lean is a proof assistant that checks whether formal logical steps compile correctly from stated premises. Think of it like spell-check for mathematical logic: it catches structural errors, but it cannot tell you whether the essay argues something true, important, or original.
Formal verification does not establish that a result is novel, significant, correctly framed, or genuinely connected to the open problem it claims to address. Notably, not every manuscript in the repository carries a Lean formalization; some remain unformalized or include explicit caveats.
The Transparency Problem
Why missing model access and prompt disclosure matter to anyone trying to reproduce these results.
OpenAI disclosed average compute statistics but did not release the exact prompts or the model itself. After controversy surrounding an earlier Navier-Stokes claim, OpenAI reportedly convened an independent advisory group. That group recommended releasing the model, the exact prompt, and compute time for each result.
OpenAI published this collection without fully adopting those recommendations. The repository does include manuscripts, supporting proof artifacts, and reasoning summaries for a subset of result families, but the unreleased model limits outside researchers’ ability to retrace or reproduce the generation process.
The accountability concern is straightforward: without access to the model and prompts, it becomes difficult to separate a reliable method from a fortunate output.
What Happens Next
Valid proofs and strategically useful proofs are not automatically the same thing.
Determining which results are correct, genuinely novel, and worth pursuing in depth could take the mathematics community considerable time. A large collection can contain technically valid but incremental or poorly motivated work, and sorting signal from noise at this volume is itself a significant research burden.
A broader question is not only whether AI can produce valid proofs, but whether it can reliably identify meaningful problems, generate durable new concepts, and communicate results in a form humans can independently understand and build on. That question remains open, and the work of answering it has largely landed with the mathematics community.




























