• Zedstrian@sopuli.xyz
    link
    fedilink
    English
    arrow-up
    71
    arrow-down
    7
    ·
    16 hours ago

    It’s not free; doing so provides Anthropic with free access to training data for its models.

    • DrCake@lemmy.world
      link
      fedilink
      English
      arrow-up
      69
      ·
      15 hours ago

      I think it’s safe to assume that any public code on the internet is already in their training dataset.

      At this point anything that is public in general, probably all the posts in the “fediverse” have been taken already

      • dan@upvote.au
        link
        fedilink
        English
        arrow-up
        17
        ·
        edit-2
        6 hours ago

        probably all the posts in the “fediverse” have been taken already

        I don’t doubt this, as it’s trivial to do (by design). You don’t even have to scrape anything.

        New posts and comments in a Lemmy community (or Mastodon profile or any other ActivityPub-based Fediverse app) get pushed to every server where federation is allowed and at least one person is subscribed to it. You can spin up your own server, subscribe to a bunch of communities, wait a while, and over time your PostgreSQL database will be populated with profiles, posts, comments, etc.

    • dan@upvote.au
      link
      fedilink
      English
      arrow-up
      28
      arrow-down
      1
      ·
      15 hours ago

      The code they’re analyzing is free and open-source though, so they can already use it this way. One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.

      Isn’t it good that they’re sharing data with the projects instead of just keeping it for themselves?

      • Zedstrian@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        7
        ·
        11 hours ago

        One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.

        Given that Anthropic and OpenAI derive their profits from their models being closed source rather than open source, I don’t think they should be assisted in devouring freely-accessible data in their endeavor.

        Not all developers may see it this way, but those contributing code to LLMs are training them for free at the expense of career opportunities for fellow developers, all the while Anthropic, OpenAI, and other AI companies rake in money for maintaining access to content that isn’t theirs.

        • dan@upvote.au
          link
          fedilink
          English
          arrow-up
          2
          ·
          edit-2
          5 hours ago

          I don’t think they should be assisted in devouring freely-accessible data in their endeavor

          How is this assisting them, though? I see it as two different options:

          1. They use the open-source code without giving anything back; or
          2. They use the open-source code and help improve the project by reporting security issues.

          In any case, a fundamental feature of both free and open-source software is that anyone can use it, regardless of if you like them or agree with them or not, without discrimination. You could instead use a license that forbids AI training, but then your code would no longer be free or open-source.