China is attempting to mirror the entire GitHub over to their own servers, users report

0x815@feddit.org · 1 year ago

China is attempting to mirror the entire GitHub over to their own servers, users report

bionicjoey@lemmy.ca · 1 year ago

Solution: create a GitHub repo with Markdown articles outlining human rights abuses by the CCP and have a large number of GitHub users star and fork the repo.

Asherah@lemmy.world · 1 year ago

Maybe we should consider the same for the US government instead of being afraid of the big Chinese boogeyman across the sea? Because I guarantee you the US has just as many, if not more. But China bad. 🙄

bionicjoey@lemmy.ca · edit-2 1 year ago

I was making a joke about abusing Chinese censorship in order to stop them cloning GitHub repos (assuming that was something you wanted to do). The joke being that the CCP suppresses information about their human rights abuses. That is not true of the US. You could absolutely make a GitHub repo detailing the crimes of the US government. Nobody will stop you.

bufalo1973@lemmy.ml · 1 year ago

Tell that to Julian Assange

Doom@ttrpg.network · 1 year ago

Is that what you think got him in trouble?

☂️-@lemmy.ml · edit-2 27 days ago

deleted by creator

raspberriesareyummy@lemmy.world · 1 year ago

With the obligatory “fuck everyone who disregards open source licenses”, I am still slightly amused at this raising eyebrows while nearly no one is complaining about MS using github to train their copilot LLM, which will help circumvent licenses & copyrights by the bazillion.

kava@lemmy.world · 1 year ago

If I look at a few implementations of an algorithm and then implement my own using those as inspiration, am I breaking copyright law and circumventing licenses?

sugar_in_your_tea@sh.itjust.works · 1 year ago

That depends on how similar your resulting algorithm is to the sources you were “inspired” by. You’re probably fine if you’re not copying verbatim and your code just ends up looking similar because that’s how solutions are generally structured, but there absolutely are limits there.

If you’re trying to rewrite something into another license, you’ll need to be a lot more careful.

kava@lemmy.world · 1 year ago

What’s the limit? This needs to be absolutely explicit and easy to understand because this is what LLMs are doing. They take hundreds of thousands of similar algorithms and they create an amalgamation of it.

When is it copying and when it is “inspiration”? What’s the line between learning and copying?

sugar_in_your_tea@sh.itjust.works · edit-2 1 year ago

I disagree that it needs to be explicit. The current law is the fair use doctrine, which generally has more to do with the intended use than specific amounts of the text/media. The point is that humans should know where that limit is and when they’ve crossed it, with motive being a huge part of it.

I think machines and algorithms should have to abide by a much narrower understanding of “fair use” because they don’t have motive or the ability to Intuit when they’ve crossed the line. So scraping copyrighted works to produce an LLM should probably generally be illegal, imo.

That said, our current copyright system is busted and desperately needs reform. We should be limiting copyright to 14 years (as in the original copyright act of 1790), with an option to explicitly extend for another 14 years. That way LLMs can scrape comment published >28 years ago with no concerns, and most content produced >14 years (esp. forums and social media where copyright extension is incredibly unlikely). That would be reasonable IMO and sidestep most of the issues people have with LLMs.

China is attempting to mirror the entire GitHub over to their own servers, users report

China is attempting to mirror the entire GitHub over to their own servers, users report

Still (@still@infosec.exchange)