• threeonefour@piefed.ca
    link
    fedilink
    English
    arrow-up
    150
    arrow-down
    1
    ·
    2 days ago

    This seems to be the mathematical equivalent of using AI to submit 200 unverified pull requests to an open source project and then telling the maintainers it’s their job to figure it all out.

    The one mathematician saying he’s not going to spend hours of his time to verify if a slop report is true, let alone do it for dozens of reports, reminds me of all those projects making rules that unverified AI pull requests will be trashed.

    • 42yeah@eviltoast.org
      link
      fedilink
      English
      arrow-up
      47
      ·
      2 days ago

      I think OpenAI has retracted some of them already. For OpenAI it’s a really low-risk thing: if it’s wrong, then just retract the paper. If it’s right though, the fame all goes to OpenAI. Meanwhile, people who spent their whole life researching on this topic, needs to confirm this manually for OpenAI.

    • a_non_monotonic_function@lemmy.world
      link
      fedilink
      English
      arrow-up
      24
      ·
      2 days ago

      I mean, this is a step worse actually. We’ve already seen mathematicians claiming that these systems have actively scooped them. Lots of academics are using these systems regularly.

      At this point, every time I see a paper being “published” by an AI company, I’m wondering who they stole the result from.

    • badgermurphy@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      edit-2
      1 day ago

      It definitely is in the very same family of rudeness to the system. Like with everything anymore, we have to design every system to account for bad faith actors, because they are too abundant to ignore and will quickly inundate any good faith, mutual trust, or professional courtesy based system with their “flood the field with bullshit” approach to everything.

      I think in this case, as in many others, we need to shift the burden of contribution more onto the contributor so that spurious contributions are more costly to the contributor than the system they’re contributing to. Much like with bots now disrespecting robots.txt, we have to add an element of discomfort to that violation of trust, since the violators lack any senses of community, respect, or shame to keep them acting in good faith.

      A system that achieves the same goal as the Nepenthes Web project that punishes and wastes AI resources that disrespect web sites’ bot policies may be needed for software and scientific contributions to raise the bar on contributions such that it is not so easy for prompt engineers to copy and paste some LLM output and call themselves mathematicians.

      Excuse me while I go take a shower after saying “prompt engineer”.

    • Venator@lemmy.nz
      link
      fedilink
      English
      arrow-up
      12
      ·
      2 days ago

      Also makes me wonder if it found and exploited (or got caught out by) some bugs in the Lean programming language…

      (Not saying Lean is buggy, but finding bugs seems more likely to me, as a programmer who knows not much about mathematical proofs since I haven’t looked at anything like that since uni, about a decade ago)

    • nialv7@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      12
      ·
      edit-2
      2 days ago

      At least half of these are formally verified. Although they did retract 3 papers. Out of about 700

      • Crozekiel@lemmy.zip
        link
        fedilink
        English
        arrow-up
        25
        arrow-down
        1
        ·
        2 days ago

        No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.

        OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…

        • a_non_monotonic_function@lemmy.world
          link
          fedilink
          English
          arrow-up
          11
          ·
          2 days ago

          Even worse, the software-based proofs are not the same as the plain text ones that they’re giving out.

          There’s literally no reason to trust them, because it’s completely divorced from the actual text.

        • nialv7@lemmy.world
          link
          fedilink
          English
          arrow-up
          12
          arrow-down
          5
          ·
          2 days ago

          Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.

          Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.

          • postscarce@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            6
            arrow-down
            1
            ·
            2 days ago

            It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.