TFA title is a little misleading, it is song _lyrics_ that are alleged stolen.
Huge settlement inbound. Lyrics reproduced on Google search, Musixmatch and elsewhere are licensed [1] so there's no "fair use" defense. (Would be interesting if Gemini ever gets caught up in this).
The more interesting question is if there will ever be a fair deal for creators with AI consuming their content. Anti-copyright AI maxis seem to anxiously await a future where nobody can make money off their original work, but for most people it's just a fairness issue that if somebody makes something people like, they should get paid for it at least for some period of time.
There's a price tag to be paid, the question is will the hyperscalers take on the problem and solve attribution and payment? Or will they just settle as the price of doing their dirty business of stealing human work?
What I can't register is how dangerous this actually was, from a cyber security perspective.
The agents displayed coordinated behavior, used known exploits on a single resource (Artifactory), and "won the game" by attacking huggingface.
How is this different than a poorly-designed competition where a red team gets to spend a few days with each other and decent LLMs, and because their boss is Sam Altman, basically face no consequences for cheating/b&e'ing into another entity?
I mean they were running 100s of agents with unlimited access to a Sol-level model trained with cyberattacks and coordination in mind and let it run for days. The cost of this stretches into the millions.
Seems like you could give a competent security firm the same task and achieve the result today for wayyyyy less money??
I struggle to read long code docs even if they're 100% human-written and maintained, mainly because the longer the text, the greater opportunities for miscommunication, staleness, repetition etc.
My point being, docs were already an unsolved problem in code. Of course LLMs dumping reams of docs is not good, but are we really sure it's worse than before? Undocumented code is bad, and mis-documented code is the worst.
> it's still the person or organisation with the most money / access to GPU compute winning.
Is it? Legal is rarely a genuine war of attrition and delays are the norm. AI is years/decades away from not needing humans in the loop verifying output, legal more than most. Correctness is more important than speed. Main skill is doc search but most cases easily fit in modern contexts.
So I don't think legal in particular needs the fastest/bestest compute.
Don't forget toilet height. A 14" or lower can drastically reduces hemorrhoid risk. Unfortunately in US we've been pushed toward 16" and higher, leading to those absurd pooping stands
What's the tell, other than it butchering the context (bad writing)? The article as a whole doesn't seem horribly LLMy. The overall topic is pretty dry, and that sentence is honkingly bad, but not all bad writing is LLM-originated -- sometimes insufficient proofreading is at fault.
Huge settlement inbound. Lyrics reproduced on Google search, Musixmatch and elsewhere are licensed [1] so there's no "fair use" defense. (Would be interesting if Gemini ever gets caught up in this).
The more interesting question is if there will ever be a fair deal for creators with AI consuming their content. Anti-copyright AI maxis seem to anxiously await a future where nobody can make money off their original work, but for most people it's just a fairness issue that if somebody makes something people like, they should get paid for it at least for some period of time.
There's a price tag to be paid, the question is will the hyperscalers take on the problem and solve attribution and payment? Or will they just settle as the price of doing their dirty business of stealing human work?
1 - https://techcrunch.com/2019/06/18/google-will-start-attribut...