Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

newscientist.fortune.newscientist.Mathematicians are checking a large batch of AI-generated proofs that OpenAI posted on Oct. 6, and some of what they have found so far is raising doubts. A team led by Cambridge researchers says the human-readable and computer-checkable versions of the company's earlier Navier-Stokes proof don't match.newscientist
The October release included 722 manuscripts, grouped into 372 "families" of results on long-standing open problems. All of them were produced by an unreleased internal model. OpenAI said the model spent about three hours of computing time, on average, on each solution. The release came weeks after OpenAI said it had solved the Navier-Stokes problem, one of the Clay Mathematics Institute's Millennium Prize Problems. The company says the new batch makes progress on three more Millennium problems but does not fully solve them.axios+1
OpenAI published its Navier-Stokes proof in two forms. One is written in "natural language," the mix of English and symbols that mathematicians use. The other is in Lean, a programming language that lets a computer check every logical step. According to New Scientist, Anders Hansen of the University of Cambridge and his colleagues found that the two versions diverge at Lemma 8.6. The written proof requires a value to stay below m + 4. The Lean version only requires it to stay below m + 5, which is a weaker condition.newscientist
The researchers did not say either proof is wrong. Their concern is that when the AI translates a proof into Lean and part of it won't compile, it may quietly change the argument to get around the problem. It took the team about two weeks to find the discrepancy. OpenAI has said its agents spent 88 hours producing the proofs.newscientist
"What has to be done with all of these large language model-generated proofs is that they will have to be read by humans, and this creates an enormous extra burden on mathematicians," Hansen said.newscientist
OpenAI told New Scientist it knows about the mismatch and that it does not make either proof invalid. The company said it will fix errors in the written proofs as they are found and will keep formalizing the new papers. Only some of the papers come with Lean proofs, and those Lean proofs have not been checked by hand. Kevin Buzzard of Imperial College London said: "I am confident that the Navier-Stokes problem has been correctly resolved. I am far less confident that the proof described in the PDF is correct."newscientist
The new release has drawn both excitement and concern. Dan Litt of the University of Toronto told Fortune "this is great for mathematics," but he also warned that the idea that AI has "solved math" could weaken support for human mathematical research. Tristan Buckmaster of New York University questioned whether OpenAI had made sure its model did not build on unpublished work by human mathematicians. He told The New York Times, "I don't think they've done their sort of due diligence at all".fortune+1
An independent advisory group based at the Institute for Advanced Study had issued recommendations for how AI-generated proofs should be published. OpenAI followed some of them but not all. The group had also said it does not support AI companies testing their own models on open research problems, and OpenAI appears to have set that position aside.smh+1
Terence Tao of UCLA wrote on Mastodon that the release marks the end of "Math 1.0." He called for a new era that will "decenter the role of raw problem-solving and value mathematical progress more holistically."fortune