OpenAI has made public a collection of 722 mathematics manuscripts generated by an internal AI system, claiming the work stemmed largely from a single prompt. While the release showcases a remarkable volume of output, the community notes that a small portion has been formally validated, leaving many results unverified.

Scale of the release

The company uploaded the papers to a public GitHub repository, organizing them into 372 thematic groups. Each group may contain a primary theorem together with related arguments or alternative demonstrations, meaning the manuscript count does not equal the number of distinct problems solved. OpenAI estimates it presented roughly 4,000 challenges to the model and retained those it deemed noteworthy. The average computation time per result was reported as about three hours of ChatGPT Pro processing.

Formal verification status

Out of the 722 documents, 162 have been translated into Lean, a proof-assistant language that checks each logical step mechanically. This represents roughly 22% of the collection. A Lean check confirms that the proof follows from the formal statement, but it does not guarantee that the statement matches the original problem or that the result is novel. OpenAI acknowledges that the remaining, unformalized papers could contain errors.

Calls for reproducibility

Mathematicians have expressed caution, emphasizing that the claim of solving hundreds of problems with a single AI prompt remains unproven until the model and its inputs are disclosed. Andrew Sutherland of MIT described the assertion as “unverified” without access to the underlying system. An advisory panel from the Institute for Advanced Study also urged OpenAI to share the model name, prompts, reasoning chains, compute costs, and timing for each result—details that were omitted from the release.

Mixed reactions from the scholarly community

Opinions among experts vary. Some, such as Professor Abhishek Saha, view the event as a notable milestone for the field, noting that many of the results fall within incremental advances rather than revolutionary breakthroughs. Others, like Daniel Litt of the University of Toronto, argue that withholding the answers defeats the purpose of scientific openness. The Institute for Advanced Study warned that AI can now produce mathematical arguments beyond the immediate comprehension of the human prompt engineer, underscoring the need for human oversight.

Comparison with other AI efforts

Anthropic recently posted a Lean-checked proof of Fermat’s Last Theorem, releasing the full 13-million-line codebase. Unlike OpenAI’s claim of novel discoveries, Anthropic’s work formalized an existing theorem, providing a contrast in transparency and scope.

Why it matters

The episode highlights a growing tension between rapid AI-driven research and the traditional safeguards of mathematical verification. If AI can generate plausible proofs at scale, the burden on human mathematicians to validate and interpret those results will increase. Transparency about model architecture, prompts, and computational resources will be essential to integrate AI contributions responsibly into the scientific record and to maintain trust in peer-reviewed mathematics.