They are hungry for active projects, want to distill better datasets.
I am sorry, to enroll a primary maintainers need to open a Pull Request with a Yaml file containing their email address?
And people do that? Here is an example PR: https://github.com/anthropics/oss-scanner/pull/139/changes
Did Anthropic just created a largest ever publicly available collection of email addresses of primary maintainers of OSS projects matching this criteria
established projects that have a critical impact on infrastructure and user security
(Quote from Anthropic)
Is it an invitation for every bad actor to scrape pull requests of this single repo, extract email addresses and target those with every fishing/hacking/account takeover attack imaginable?
Is not that insane?
One of the uses i can fully stand behind - we all profit from secure software, and since it’s open source anyways, there’s nothing to lose by using this.
The thing to lose is human expertise in finding vulnerabilities, which is needed to feed the machines. It seems likely fewer people will go into the field if it is viewed as ‘solved’ by a robot.
That depends. If an coding LLM finds a vulnerability, would you have been able to find it yourself? And the more important part is: do i blindly merge whatever i get thrown at me or do i critically engage with it, analyze why this is a security issue and how to prevent it in the future?
The thing is: an LLM can NEVER replace secure design. The best LLM can’t do shit if you have designed your software without security in mind from the start. And this also means this area will never be fully solved by machines, outside of them becoming able to code the thing securely from start to finish (and even then the design which is the basis of that software has to take security into account, or whatever you get will be insecure by default). This is one of the areas that can be massively enhanced by AI, but it can’t replace people that are able to design secure systems.
It’s not free; doing so provides Anthropic with free access to training data for its models.
I think it’s safe to assume that any public code on the internet is already in their training dataset.
At this point anything that is public in general, probably all the posts in the “fediverse” have been taken already
probably all the posts in the “fediverse” have been taken already
I don’t doubt this, as it’s trivial to do (by design). You don’t even have to scrape anything.
New posts and comments in a Lemmy community (or Mastodon profile or any other ActivityPub-based Fediverse app) get pushed to every server where federation is allowed and at least one person is subscribed to it. You can spin up your own server, subscribe to a bunch of communities, wait a while, and over time your PostgreSQL database will be populated with profiles, posts, comments, etc.
The code they’re analyzing is free and open-source though, so they can already use it this way. One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.
Isn’t it good that they’re sharing data with the projects instead of just keeping it for themselves?
One of the freedoms of free software is that people can study it and use it however they want, for whatever purpose they want.
Given that Anthropic and OpenAI derive their profits from their models being closed source rather than open source, I don’t think they should be assisted in devouring freely-accessible data in their endeavor.
Not all developers may see it this way, but those contributing code to LLMs are training them for free at the expense of career opportunities for fellow developers, all the while Anthropic, OpenAI, and other AI companies rake in money for maintaining access to content that isn’t theirs.
I don’t think they should be assisted in devouring freely-accessible data in their endeavor
How is this assisting them, though? I see it as two different options:
- They use the open-source code without giving anything back; or
- They use the open-source code and help improve the project by reporting security issues.
In any case, a fundamental feature of both free and open-source software is that anyone can use it, regardless of if you like them or agree with them or not, without discrimination. You could instead use a license that forbids AI training, but then your code would no longer be free or open-source.
Loling at the previous comments here… As if AI labs don’t already vacuum all open source code they can find
“The outputs of this opt-in vulnerability scanner will be fully model-generated, without human review or triage,” Anthropic explained. “This will enable faster and more frequent scanning, but means that it is possible reports will be incorrect or invalid.”
When I saw the headline, I was wondering about this specifically. This may make this service not super useful.
My experience with AI security reviews is that they’re fantastic at finding faults, but they always find a list of things to complain about. If there are no real/serious faults they’ll start finding things that kind of have the same shape as a security issue, but really aren’t if you dig into them. I’ve regularly had an LLM generate a list of 10-15 issues ranging in severity from “nits” to “critical” where none of them were actual issues.
Periodic reviews seem like they could get annoying really quickly, becoming more of a maintenance burden than a help.
Yeah. I did use an open weight model LLM to scan an app that I use and it came up with a critical severity issue and I checked on it first to see if I could reproduce the problem before ever submitting it to the developer as a problem.
Turns out it was actually a legitimate security problem that has now been fixed.
Yeah, there is some fatigue to be expected if you run this thing on the regular.
But the library ecosystem is changing constantly, and that might mean that while it doesn’t find a real issue at one point that there won’t be one some time in the future. So i would say: Scan it semi-regulary for issues and review what was found, but don’t go overboard. Somewhat similar to getting an MRT every year instead of once a decade.
AI security reviews is that they’re fantastic at finding faults, but they always find a list of things to complain about.
Same experience for me. At work, I always review the output because a lot of times it finds something that can be noteable but not relevant to the scope of what is being reviewed. Even with guardrails like instructions to not go beyond the stated scope, a human absolutely still needs to read and check the findings at the end.
In return your code is theirs.
and they haven’t already scraped the whole of OSS already?
Assuming yes, it would’ve been outdated, and this give them access to fresh new code for a foreseeable future in the case of Microsoft blocking access from other AI to make it exclusive for Co-Pilot.
They don’t have to offer this to get new versions, that’s kinda the point of foss software, it’s there for anyone to download and use however they like whenever they want.
Hot how ever they like though, that would be code under MIT ant the like, not all FOSS. Where the other licenses actually fall?
No, the GPL doesnt put any restrictions on what you can do with code, just on that code and derivatives of it have to be released (source code included) under the GPL as well.
The courts have tended to find that training LLMs on data you have legal access to is fair use and transformative.
It‘s pointless. The pace at which games are decompiled with AI in the modding community tells me that each and every software will be figured out and broken into within a few years or even months. The digital world will be turned upside down. All data could be leaked, entire hedge funds and banks could get wiped with a mouse click and governments could be toppled over night when pension funds of millions suddenly disappear. The AI ouroboros is a vortex that could swallow the entire internet including everything we put in it. Moving fast and breaking things may just break our way of life forever.
I do agree with your fears regarding AI, although i am very sure that the real dangers lie with an fully automated surveillance state where even punishment of perceived transgressions are outsourced to AI and not in the decompilation of games.
Securely designed software has no problem with it’s source being available. This is ensured not by hidden code, but by mathematics.
De-compiling isnt magic that lets you do whatever you want with software. It just works well for games because DRM has the insane model of “we give you the encrypted payload and the key to unlock it, but you can only unlock it in the way we want” which is fundamentally insecure. Most software isnt like that.
Wait you mean just because I find the deposit api for my bank I can’t send it a fake 100m deposit into my account? 😜
Things take years to build, minutes to erase.







