Skip to content
NewsResearch

OpenAI posts 722 math manuscripts from an unreleased model on GitHub

· by Pondero Newsdesk

The short version

OpenAI published Lean-verifiable proofs and compute disclosures from an internal frontier model on October 6, following weeks of mathematician criticism over earlier, less transparent claims.

OpenAI posts 722 math manuscripts from an unreleased model on GitHub

OpenAI published results from an internal, unreleased frontier model on October 6, choosing a GitHub repository with computer-checkable proofs over a traditional paper, after weeks of mathematician criticism over how the company handled earlier claims about the same model.

What

The repository holds 722 manuscripts covering 372 result families, according to The Verge's review of the GitHub release. Many of the proofs come with formalizations in Lean, a language that lets software check a proof's logic rather than resting on human peer review alone. OpenAI also posted 10 written summaries of the model's reasoning, estimates of the compute spent, and counts of how many problems the model attempted; per OpenAI's release post, the average result used roughly three hours of ChatGPT Pro-equivalent "thinking." The company built this release protocol with input from the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), an independent panel of mathematicians based at the Institute for Advanced Study.

Lean proofs answer the self-reporting complaint, not a product question

AGMAI had published its own recommendations days earlier, urging AI labs to name the model, publish prompts and compute cost, and route releases through established academic channels rather than, in the group's words, treating open-problem solutions as marketing for their models. The panel formed after OpenAI said in September that an internal model, the same one that had already resolved the Navier-Stokes Millennium Prize problem, had solved more than 100 long-standing open problems across most of mathematics, per OpenAI's September announcement. That framing and the pace of releases drew pointed pushback, including a Verge piece headlined "OpenAI keeps bulldozing mathematicians". The Lean formalizations give outside mathematicians a way to check a result's logic rather than take OpenAI's word for the claim, which is the specific gap AGMAI's recommendations called out. For AI-tool buyers, there is nothing to test or deploy here: the model behind these results has no public API, no benchmark release, and no stated ship date tied to this announcement.

What to watch next

Mathematicians working with AGMAI have not yet said which of the 372 result families they consider validated, and OpenAI's post commits to funding workshops and conferences on understanding AI-generated math results without naming dates. Both will show whether the new disclosure habits hold once the field starts checking the Lean proofs line by line.

Sources