Hugging Face crawls as many GitHub repositories as it can with the help of the Software Heritage Archive so as to create their Stack v2 LLM for code. And of course without consent of rights holders.
Opt-out is possible.
Oh, scrapes.
My code was scraped according to this website but let’s not kid ourselves that any kind of opt-out will prevent this. Any corpo knows that you can outsource all of your shady dealings to a couple of layers of third party vendors and pretend you didn’t know any better. It’s safe to assume that if you use an LLM that you didn’t train yourself you’re violating copyleft of some GPL project.
The opt-out procedure is on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard’.
I checked and 52 of my repos are in this.
The datasets are stored in a git repository, so even when your opt-out request is “honored”, the “deleted” data remains in the repository history.
Hugging face open sourced a dataset every big AI company already has privately. We need to get mad about this because hugging face is the last target to take down, to finish the regulatory capture started at the behest of openAI and company.
Even if scrapping GitHub was made illegal, does anybody here really think the money would go to them and not Microsoft?
It’s data brokers and the copyright mafia vs open source and you guys are choosing to root for the former?
I’m supposed to be upset about ai using opensource code now?
The second part of my first paragraph is sarcastic. I realize I worded it weirdly.
That being said, they need us upset so it’s easier to pass laws that make developing AI models impossible for smaller players.
Most are begging for stronger copyright laws. It’s nuts.






