Summary

  • OpenAI has shared 722 math papers on GitHub, generated by an internal, unreleased AI model.
  • Only 162 of these papers contain a verified main result, and OpenAI cautions that some unverified findings may be flawed.
  • MIT mathematician Andrew Sutherland emphasizes that claims of solving problems with a single prompt remain unproven until the model is publicly accessible.

OpenAI has published a collection of 722 mathematics manuscripts on GitHub, all originating from an internal model that has yet to be made public. According to an OpenAI representative, nearly all results were produced from a single prompt given to one AI agent, although some may have required multiple tries.

This assertion marks a potentially groundbreaking moment in mathematics, but it has not been met with universal enthusiasm.

“Until they provide access to the model and allow for replication of their findings, any assertions regarding solving problems with a single agent should be considered unverified,” stated Andrew Sutherland, an MIT mathematician, in an interview with Scientific American. He added, “We should ask for receipts.”

The manuscripts are categorized into 372 families of related results, where each family may include a main theorem along with supporting arguments, implications, or alternative proofs. This means that the count of 722 pertains to manuscripts, not necessarily distinct solved problems. OpenAI reports that it posed around 4,000 problems to the model, selecting outputs deemed significant enough for publication.

The average result was derived from approximately three hours of computational time using ChatGPT Pro, according to OpenAI. This contrasts with a previous claim regarding the Navier-Stokes equations, which involved 10,000 agents working for 88 hours.

OpenAI has provided condensed reasoning summaries for 10 of the results; however, only 162 out of the 722 papers include a computer-verified main result, as indicated in the formalization catalog. This accounts for about 22% of the total, verified using Lean, a software that mechanically checks logical steps.

OpenAI acknowledges that not all manuscripts possess Lean formalizations and warns that “some of the unformalized results could have issues,” implying that many published findings might be incorrect.

A Lean verification confirms that the proof follows from the statement as recorded in Lean but does not guarantee that the statement accurately reflects the original problem or that the result is novel or significant—this is where mathematicians must apply their judgment.

This has led to skepticism among researchers. Dmitry Rybin noted, "I was trying to read the OpenAI proof regarding the chromatic number of a plane being >= 6, but it appears to be incomprehensible alien math? The model seemingly linked any K-coloring to 'weakly measurable' K-coloring, which feels arbitrary."

“The openai/math repository has Issues disabled and has never accepted any pull requests, which is disappointing. If you publish 722 manuscripts and solicit Lean formalizations, it’s essential to have a way for people to contribute,” commented Keith Adler, who is currently formalizing OpenAI's proof of Saxl's Conjecture in Lean 4.

The Institute for Advanced Study in Princeton, New Jersey, expressed concerns, stating, “AI can now generate mathematical arguments in contexts that the prompting human cannot comprehend, verify, or take responsibility for. We emphasize the importance of human understanding in mathematics and question how we can develop a new paradigm that integrates human comprehension into responsible academic output.”

Conversely, some, like Professor Abhishek Saha, are optimistic, declaring, “It is a monumental day for mathematics,” while acknowledging that most of the problems fall under the category of “exceptional advancements within existing frameworks” rather than groundbreaking discoveries. Only one problem out of the 722—namely, the Quasi-Riemann Hypothesis—could be considered truly significant.

“If I were to classify theorems that mathematicians prove and publish by their groundbreaking nature, I would roughly categorize them into four groups,” Saha tweeted, sharing his insights on the 372 results released by OpenAI.

Additionally, the release did not fully comply with recommendations made by an advisory group at the Institute for Advanced Study, which had called for the model name, prompting details, a summarized thought process, time spent, and computational costs for each result. While OpenAI has released average computational figures and 10 reasoning summaries, it has yet to provide the prompts and is still working on a responsible model release.

In contrast, Anthropic took a different approach with its Lean-verified Fermat's Last Theorem proof, which was made publicly available on GitHub, containing 13 million lines of code, formalizing a theorem initially published by Andrew Wiles in 1995 without asserting new results. OpenAI has stated it will add Lean formalizations as they become available; currently, 162 of the 722 manuscripts include one.