Who checks 720 proofs written in a week? That is the question OpenAI's mathematics release has forced on a field that normally measures progress one refereed paper at a time.
On October 6, OpenAI published "Sharing AI progress in mathematics": roughly 720 manuscripts addressing about 370 open problems, produced by what the company calls "an internal frontier model." The model is not named anywhere. The package includes Lean formalisations for part of the results (Lean is a proof assistant that verifies each step mechanically), ten summaries of the model's reasoning, and compute estimates: an average result corresponds to about three hours of ChatGPT Pro-level thinking. The Decoder reports, without OpenAI confirmation, that roughly 8,000 problems were attempted with a success rate near 5%.
One day later, three papers were gone.
The withdrawals
OpenAI's own change log on GitHub, dated October 7, records three withdrawn manuscripts, all in algebraic geometry: on the algebraicity of Weil classes on split abelian eightfolds, on Kuga-Satake correspondences for K3 surfaces, and on the rational Hodge conjecture for products of K3 surfaces. The reason, in OpenAI's words: "a sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers." One mistake, three papers. The same update revised 14 other manuscripts and added six formalisations, bringing the share of top-line results verified in Lean to 300 of 719, about 42%.
Three withdrawals out of 720 is a small ratio. What made it a story is that the error was found after publication, by the community, in a paper that two others depended on.
Is the criticism about the error?
Not mainly. The sharpest statement came from the Association for Human Mathematics, published as a guest post on Terence Tao's blog. It accuses OpenAI of violating the core norms of scientific research and of ignoring the central request of the independent Advisory Group on Mathematics and AI, which had asked companies not to test advanced open problems with proprietary models. It rejects the word "progress," calls the release "a demonstration of power," and asks mathematicians to discontinue their work with OpenAI. No source gives a count of how many have joined.
Tao's own view, expressed on social media and quoted by TechCrunch, is that problems are now "being solved autonomously by AI prompters" who do not understand the output well enough to discuss it. He argues the damage is irreversible: a solved problem cannot be unsolved, and knowing a solution exists contaminates the search for other approaches. Harvard's Melanie Wood put it plainly: "there is not human understanding of them at the point of release, and now the work begins." Scott Aaronson called the episode a "Mathocalypse." Dana Moshkovitz, who works on the Unique Games Conjecture that one paper claims to settle, said that paper was "so horribly written that it's impossible to read it without AI help."
The advisory group itself was diplomatic, saying it is up to the community to judge how well OpenAI followed its late-September guidelines. TechCrunch's audit: OpenAI did release promptly and did publish reasoning, but only ten of 719 manuscripts include it, 58% are not formalised, and there is no funding for human verifiers. A paper from Cambridge and King's College London found at least two discrepancies between OpenAI's natural-language proof and its Lean code on a Navier-Stokes-derived problem, and concluded such proofs "should not prima facie be trusted" without peer review.
Why this is not only a mathematics story
Strip away the field and the structure is familiar. A system produces output far faster than anyone can verify it. Part of the output is checked by machine and an error surfaces within a day. The rest sits in natural language, waiting for a human who has weeks, not hours.
Every business deploying AI faces the same ratio. A generated contract clause, financial summary or code change is only useful if it can be checked at something close to the speed it was produced. In our project work the teams that get value from AI are the ones who design the verification step first: a test suite, a reconciliation rule, a human sign-off with a clear scope. OpenAI's release shows both halves of that lesson at once. Where verification was mechanical, the system corrected itself in 24 hours. Where it was not, the backlog is now the field's problem.
Sources: OpenAI, math repository change log, TechCrunch, The Decoder, Retraction Watch




