No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.
OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…
Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.
Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.
It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.
No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.
OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…
Even worse, the software-based proofs are not the same as the plain text ones that they’re giving out.
There’s literally no reason to trust them, because it’s completely divorced from the actual text.
Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.
Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.
It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.