probably all the posts in the “fediverse” have been taken already
I don’t doubt this, as it’s trivial to do (by design). You don’t even have to scrape anything.
New posts and comments in a Lemmy community (or Mastodon profile or any other ActivityPub-based Fediverse app) get pushed to every server where federation is allowed and at least one person is subscribed to it. You can spin up your own server, subscribe to a bunch of communities, wait a while, and over time your PostgreSQL database will be populated with profiles, posts, comments, etc.
The code they’re analyzing is free and open-source though, so they can already use it this way. One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.
Isn’t it good that they’re sharing data with the projects instead of just keeping it for themselves?
One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.
Given that Anthropic and OpenAI derive their profits from their models being closed source rather than open source, I don’t think they should be assisted in devouring freely-accessible data in their endeavor.
Not all developers may see it this way, but those contributing code to LLMs are training them for free at the expense of career opportunities for fellow developers, all the while Anthropic, OpenAI, and other AI companies rake in money for maintaining access to content that isn’t theirs.
I don’t think they should be assisted in devouring freely-accessible data in their endeavor
How is this assisting them, though? I see it as two different options:
They use the open-source code without giving anything back; or
They use the open-source code and help improve the project by reporting security issues.
In any case, a fundamental feature of both free and open-source software is that anyone can use it, regardless of if you like them or agree with them or not, without discrimination. You could instead use a license that forbids AI training, but then your code would no longer be free or open-source.
It’s not free; doing so provides Anthropic with free access to training data for its models.
I think it’s safe to assume that any public code on the internet is already in their training dataset.
At this point anything that is public in general, probably all the posts in the “fediverse” have been taken already
I don’t doubt this, as it’s trivial to do (by design). You don’t even have to scrape anything.
New posts and comments in a Lemmy community (or Mastodon profile or any other ActivityPub-based Fediverse app) get pushed to every server where federation is allowed and at least one person is subscribed to it. You can spin up your own server, subscribe to a bunch of communities, wait a while, and over time your PostgreSQL database will be populated with profiles, posts, comments, etc.
The code they’re analyzing is free and open-source though, so they can already use it this way. One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.
Isn’t it good that they’re sharing data with the projects instead of just keeping it for themselves?
Given that Anthropic and OpenAI derive their profits from their models being closed source rather than open source, I don’t think they should be assisted in devouring freely-accessible data in their endeavor.
Not all developers may see it this way, but those contributing code to LLMs are training them for free at the expense of career opportunities for fellow developers, all the while Anthropic, OpenAI, and other AI companies rake in money for maintaining access to content that isn’t theirs.
How is this assisting them, though? I see it as two different options:
In any case, a fundamental feature of both free and open-source software is that anyone can use it, regardless of if you like them or agree with them or not, without discrimination. You could instead use a license that forbids AI training, but then your code would no longer be free or open-source.