✓ AI-Debiased Article
Rewritten from Hacker News — Front Page • • 2 min read
4 Wire-neutral provisional

✓ No loaded language, vague sourcing, or framing detected.

OpenAI's Navier-Stokes Proof Contains Mistranslation, Researchers Claim

Researchers from the University of Cambridge have identified a mistranslation in OpenAI's proofs of the Navier-Stokes problem, raising concerns about the reliability of AI-generated mathematical results. The discrepancy involves a difference in the requirements of a specific lemma between the natural language proof and the Lean code proof. OpenAI has acknowledged the mismatch and plans to address any errors.

Companies
OpenAI
People
Anders Hansen Fabian Circelli Alexander Bastounis Kevin Buzzard

A team of mathematicians has identified a potential error in OpenAI's proofs of the Navier-Stokes problem, which does not imply that the proofs are incorrect or that OpenAI has failed to solve the problem. The researchers emphasize that this raises questions about the reliability of mathematical results generated by AI models. Anders Hansen from the University of Cambridge stated, "What has to be done with all of these large language model-generated proofs is that they will have to be read by humans, and this creates an enormous extra burden on mathematicians."

On September 8, OpenAI announced a solution to the Navier-Stokes problem, a significant open problem in mathematics, and published the proof in two formats: one in natural language and the other in Lean, a computer code designed for formal verification. However, Hansen and his team argue that the two versions do not match.

Fabian Circelli, also from the University of Cambridge, noted, "This formalisation process is trying to replace peer review," highlighting that the AI's auto-formalisation cannot substitute for human review. The researchers clarified that they are not claiming OpenAI has failed to solve the problem, but they suggest that the AI may have "mistranslated" the proof into Lean.

Hansen explained that the AI must produce a Lean proof that compiles, meaning it must be self-consistent. If the AI encounters a section that does not compile, it may create a workaround that diverges from the original proof.

The team's claim focuses on a specific part of the proofs known as Lemma 8.6, where the natural-language proof requires a certain value to be below m + 4, while the Lean proof requires it to be below m + 5, which is considered mathematically weaker.

OpenAI acknowledged the mismatch and stated that it does not invalidate either proof. The company plans to correct any errors in the natural-language proof as they are identified and will continue formalising the 722 mathematics papers it released, some of which lack checked Lean proofs.

The researchers identified the discrepancy through a process involving ChatGPT, which suggested potential inconsistencies that were later verified manually. Hansen remarked that the manual verification was challenging and time-consuming.

The team emphasized the importance of ensuring that AI models can accurately auto-formalise mathematics to maintain trust in the results. Hansen expressed concern that if AI-generated proofs are accepted without thorough inspection, it could undermine scientific understanding. Kevin Buzzard from Imperial College London added that distinguishing between a theorem's statement and its proof is crucial for validation.

Hansen hopes OpenAI will take the team's findings seriously and that further development of robust auto-formalisation techniques is necessary. He concluded, "Do we have the solution? Not yet. Is it possible to do this in a controlled way? Yes, it will be, but the optimal and the ultimate way of doing this is completely unknown."

Annotating as

No note attached

on this article.

Original vs. Neutral

Original Headline

OpenAI mistranslated mathematics into code for its Navier-Stokes proof

Neutral Headline

OpenAI's Navier-Stokes Proof Contains Mistranslation, Researchers Claim