• Crozekiel@lemmy.zip
    link
    fedilink
    English
    arrow-up
    21
    arrow-down
    1
    ·
    15 hours ago

    No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.

    OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…

    • a_non_monotonic_function@lemmy.world
      link
      fedilink
      English
      arrow-up
      8
      ·
      14 hours ago

      Even worse, the software-based proofs are not the same as the plain text ones that they’re giving out.

      There’s literally no reason to trust them, because it’s completely divorced from the actual text.

    • nialv7@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      arrow-down
      5
      ·
      14 hours ago

      Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.

      Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.

      • postscarce@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        4
        arrow-down
        1
        ·
        7 hours ago

        It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.