sözaltı news Science
Science
EN AZ
OpenAI has dumped 722 maths papers – now it must clean up the mess

OpenAI has dumped 722 maths papers – now it must clean up the mess

newscientist.com 07.10.2026 18:22 5 views
OpenAI is using mathematics as a test site for its most capable AI models, and that means it has a responsibility to deal with the fallout, rather than just leaving it to mathematicians, says Jacob Aron

Yesterday, OpenAI fired an answering volley in the form of 722 maths papers dumped onto the code-sharing website GitHub, the output of a mass campaign against 4000 open problems across a broad range of mathematics. Mathematicians are still digesting the release and are likely to be doing so for some time. It is easy to think that this is an unheard-of situation, but actually I think there are some historical precedents that could prove useful in unpacking what has just happened, and what should happen next.

AI has been growing ever more capable in mathematics and OpenAI has now released hundreds of papers at the same time – to the astonishment of mathematicians Before I get on to that, there are two big questions to answer here: what does this mathematical torrent mean for AI, and what does it mean for mathematics? On the first question, we now have further confirmation that leading AI models can essentially produce seemingly research-level mathematics at the push of a button. OpenAI used an unnamed internal model and says that each result used the equivalent of an average “three hours of ChatGPT Pro thinking compute” – Pro being its highest-tier product, offered at $100 to $500 a month.

Of course, “average” is doing a lot of work here, and OpenAI’s subscription services are heavily subsidised, so it is hard to give a true cost figure for each of these results, and OpenAI hasn’t provided one. It seems unlikely, however, that the average problem required anywhere near the effort that went into the firm’s Navier-Stokes result last month, which OpenAI says took around 88 hours at a reported cost of $15 million. The company is doing this not because it wants a mathematics machine, but because it is reaching for problems that stretch its models – mathematics is just one of many test sites.

Does this apparent proficiency in maths mean that OpenAI is close to achieving similar performance across other sciences or its stated goal of artificial general intelligence, an AI model that can do anything a human can? Mathematics has a crucial component that has been essential to AI success: verifiability. There is no objective measure to confirm whether an AI can, say, write plays that rival William Shakespeare, but there is such a measure in maths because a proof is either true or it isn’t.

By formalising a proof – turning its logical steps into computer code called Lean that can be mechanically checked by a computer – AI companies can demonstrate that they have actually achieved what they claim. This provides a handy loop for improving an AI’s mathematical ability: have it produce a proof, formalise it and reward the model for accuracy. The lack of such a loop in other areas seems likely to make it harder for AI models to progress.

What is particularly interesting with OpenAI’s latest release is that its formalisation is incomplete. Quantifying this is slightly tricky: its catalogue of Lean proofs lists only 162 of the 722 papers as having had their main results formalised in this way, but a list of the latest results that are accompanied by any Lean code comes to 235 “families” out of a total 372 families across the 722 papers. However you slice it, the job isn’t done.

Extract — continue reading at the source.

Read full story