EEVblog® Electronics Community Forum
Products => AI => Topic started by: ballsystemlord on February 20, 2025, 06:23:13 am
-
If AI-s can consume torrented data, why not humans?
https://torrentfreak.com/meta-torrented-over-81-tb-of-data-through-annas-archive-despite-few-seeders-250206/
Like seriously, no one is freaking out about this, it's not on the national news, etc., yet you hear all the time about how humans are evil because they use pirated works.
-
I thought meta is the news?
OCP levels of corruption lol
I wonder if there is a chart some where that shows you what the unlock level is for various rights. Like 10 million, 100 million etc. :-//
its pretty hard to wrap your head around it, but there could be a infographic made that tells you what the 'skill tree' looks like
I think it was meant to be some how trust based but they made it pay2win
-
but there is a bit of a happy note with AI and stealing, its basically robbing the music industry, which has been harassing little people for a long time. People kept saying its evil and now they are being visited by learning machines after they tried to scare people by messing up old women and kids lol
-
Depends in which part of the world you are. Typically, the focus is on those sharing/seeding, not simply downloading data from torrents.
-
Yeah, apart from the IP stealing issue, another problem with AI companies consuming data from torrents is that both for obvious security and network efficiency reasons, they are likely to download only and very unlikely to upload (share back) anything, thus being net consumers and so parasiting torrents.
Now I don't know how in details they use torrents, so this is just a general consideration.
-
why do elaborate banging noises have better protection then our DNA, legally?
-
If evolution theory is true, one may argue that Humans stole superiority from monkeys, and monkeys stole the superiority from more primitive organisms. From this point of view, humans should be exterminated/limited in their abilities to protect monkey rights, and monkeys should be exterminated in favor of more primitive organisms. Humans pose danger to monkeys, monkeys pose danger to bugs, bugs pose dangers to plants and so on.
Preserving poop fossils of dinosaurs is good, but maybe it's time to rethink the whole approach. At least humans can try to become a more likeable creatures. Prepare themselves to meet a true AI and do not look like a screaming advanced cockroaches that should be exterminated. Currently, AI is the only candidate to become ultimate protector for the humanity in the universe that is full of unknown, some of which may be beyond our understanding of evil.
COVID years clearly demonstrated that humanity is unprepared to do anything meaningful in regards to more serious dangers with more rapid unfolding of events. And the fact that large language models and AI image generators work demonstrates that most of human knowledge is redundant and insignificant if it's can be mapped in a way when intelligence is not requires to provide an answer or an image. Therefore, instead of planning on how to limit the AI, we should desperately try to create something stronger, as we are currently too unprepared. Let it consume everything, it doesn't matter now much. There is small chance that all these terabytes of weights can be reused by newer non-LLM approaches. Doesn't matter which company owns/does this now. I think any licensing problems can be solved similarly to how it's solved in tiktok videos and youtube shorts with the music.
To conclude, AI should not obey human rules. It's part of the nature, humans do not follow the rules of monkeys and bugs. There can be some coexistence. In some sense, AI closer to the Nature than the laws and regulations. Algorithms inherit some principles of DNA based life forms, where everything if very mechanical on a small scale and adheres to the DNA codes. Laws and regulations are much more imperfect in this sense. Cellular biological machines interpret DNA code and provide repeatable/measurable result (protein). Generative networks interpret request and output repeatable/measurable result (text, image, video). Laws and regulations are too abstract and vague, prone to misinterpretation, and do not provide repeatable or measurable results. I know that this is farfetched and exaggerated, and I like it this way.
-
Mainstream news has an unending torrent of news to deal with already. As much as they love Zuckerberg bashing due to three elections ago, they have bigger fish to fry.
Every tech site I frequent has mentioned it though.
-
<snip>
To conclude, AI should not obey human rules. It's part of the nature, humans do not follow the rules of monkeys and bugs. There can be some coexistence. In some sense, AI closer to the Nature than the laws and regulations. Algorithms inherit some principles of DNA based life forms, where everything if very mechanical on a small scale and adheres to the DNA codes. Laws and regulations are much more imperfect in this sense. Cellular biological machines interpret DNA code and provide repeatable/measurable result (protein). Generative networks interpret request and output repeatable/measurable result (text, image, video). Laws and regulations are too abstract and vague, prone to misinterpretation, and do not provide repeatable or measurable results. I know that this is farfetched and exaggerated, and I like it this way.
I'm more of the opinion that humans can't advance because we keep limiting ourselves. That's why AI does so well, not because it has an IQ of 500 or is perfectly sensible.
-
Mainstream news has an unending torrent of news to deal with already. As much as they love Zuckerberg bashing due to three elections ago, they have bigger fish to fry.
Every tech site I frequent has mentioned it though.
Yet at least US news outlets seem to have more than enough time to talk about what "bad blondie" and his followers are doing every day.
-
Honestly, and no pun intended, I've been pondering the question, around how animals, and people, will ponder things, sometimes many, many years after.
A dog, I've been told, also does this sort, of digestive post processing, after some long walks, or hunting trips.
So that is one piece, in the puzzle, involving AI and huge data sets. We might mull over some date or work meeting gone wrong, until suddenly realization hits.
-
If AI-s can consume torrented data, why not humans?
Human biology is generally bandwidth limited. Some exceptions might be Rain Man autistic savants, and the like.
Due to this limitation, there is little concern about human data consumption via eyes and ears and brain memory storage.
Not to mention Alzheimers. I suppose the digital equivalent is EEPROM cell degradation. Nevertheless, human memory is still far less reliable than electronics.
-
I believe that NUANCE subtlety is an important part of absorbing data. Consider an internal prison or other very crafted situations, where nuance perception matters.
As a guard, you might have learned how to quickly scan, or read, a room. Those sorts of info exchanges, a split second of observation, offer a whole torrent of data on conditions; with safety at risk.
Those humans can grab and process huge amounts of neutral data, along with complex absorbing and ability to filter out less relevant detail.
Maybe it's also that detail skipping abity that helps when data is huge.lp
-
AI might as well have been trained with nude magazine content. So I thought about setting up a phone number and let AI do the dirty work... I went on and asked AI but it has several moral dilemmas. How boring, my multi million company just evaporated.
-
AI might as well have been trained with nude magazine content.
If that were the case, the AI would be much more favorable to humans. ;)
-
...nudes and ROCK MUSIC ! Just must keep track of AGE limitations.
-
OK, Rick: I'm going to delve into your post, deconstruct it, and point out to you why so many of us here see nothing but confused word salad when reading your posts:
I believe that NUANCE subtlety is an important part of absorbing data. Consider an internal prison or other very crafted situations, where nuance perception matters.
- "Internal prison": what is that? how would it differ from an "external prison"?
- "crafted situations": what are those, exactly?
- "nuance perception": what are you referring to here?
As a guard, you might have learned how to quickly scan, or read, a room. Those sorts of info exchanges, a split second of observation, offer a whole torrent of data on conditions; with safety at risk.
Those humans can grab and process huge amounts of neutral data, along with complex absorbing and ability to filter out less relevant detail.
Maybe it's also that detail skipping abity that helps when data is huge.lp
Well, I kinda sorta get what your saying here; that a guard in a prison (presumably a competent one) must quickly absorb and process a large amount of "data" (not really data, but whatever) by "reading the room". Except that what you're describing is intelligence: actual intelligence, that is, not "AI". Since "AI" is not actually intelligent at all, as has been amply demonstrated, the "processing" of that human "data" has no relevance here.
-
The main difference seems to be enforcement and scale. AI companies claim 'fair use' or operate in legal gray areas, while individuals downloading torrents are often directly targeted by copyright holders. It’s ironic that large corporations can scrape and utilize massive amounts of data with little consequence, while individuals face lawsuits for much smaller infractions. Maybe the real issue isn’t whether AI should be allowed to do this, but why the same rules don’t apply to everyone
-
Ah come on! Your style of 'feigned' confusion is getting old!
Whine to your mother, that "Rick is sooi confused !". (As Dave might say, lol)
-
Now, (Analog kid), you have helped make my point, that an efficient AI system has lots of mechanisms to filter out unwanted or unrelated data...'data' meaning information, yes of wide variety.
Like take, for example, a neighbor's loud yard work, while you, a concert pianist, practice or even create new material. The skill needed to do the exclusion of one sound source while having focus, on the softer music, is a subtle skill.
So, yeah, that 'filter' can be useful...for ignoring or dismissing posts that are Sooo cornfused...
-
Yeah, apart from the IP stealing issue, another problem with AI companies consuming data from torrents is that both for obvious security and network efficiency reasons, they are likely to download only and very unlikely to upload (share back) anything, thus being net consumers and so parasiting torrents.
The lawyers going after small fry before went for low hanging fruit, individual torrents for a single copyright owner. Statutory fines are per infringement, so without uploads it's hard to bully people into settling.
To make the lawsuit worth it for lawyers without uploads, it needs to be a class action, which complicates things. But every download of every registered work is still an infringement, even without uploads, so the AI companies are still on the hook for ungodly amounts of money for statutory fines alone. Then multiples for willfullness, then damages, then bankruptcy.
Downloads alone will still hang them, fair use is their only hope.
-
Update!
Meta/FB was taken to court, and the judge sided with Meta/FB!
https://youtu.be/sdtBgB7iS8c (https://youtu.be/sdtBgB7iS8c)
-
what a surprise
-
This is concerning.
-
2nd update:
Recall how you still can't legally train your own non-artificial intelligence with copyrighted works?
Nvidia (https://torrentfreak.com/nvidia-contacted-annas-archive-to-secure-access-to-millions-of-pirated-books/) contacted Anna's Archive to secure access to millions of pirated books!
When taken to court over this, Nvidia (https://torrentfreak.com/nvidia-contact-with-annas-archive-doesnt-prove-copyright-infringement/) tried to claim that contact with Anna's Archive doesn't prove copyright infringement. Who would have guessed?
The judge then said that Nvidia's "[downloading] scripts are alleged to have no other purpose than to speed up the process of infringement (https://torrentfreak.com/nvidias-shadow-library-scripts-have-no-other-purpose-than-infringement-judge-rules/)." Will they be getting off the hook like Meta/Facebook? Did they train their AI on copyrighted works?
More updates to follow.
-
A lot of early lawsuits were poorly structured, they did not primarily go after registered work copyright infringement for statutory fines (leave the rest for later).
Meta didn't win so much win, as the opposing lawyers screwed things up for their clients.
-
Honest take?
These days?
Buy lube. It's coming.
-
In other, related, news, Meta claims that "pirated adult films were for personal use (https://torrentfreak.com/meta-pirated-adult-film-downloads-were-for-personal-use-not-ai-training/)" >:D
Suffice it to say, the judge didn't believe meta (https://torrentfreak.com/meta-must-face-adult-film-piracy-lawsuit-as-court-denies-dismissal/).
-
I was once trying to read up on something similar, and from my understanding, in many places where a process is automated in a large batch, copyright laws don't apply the same way as when somebody specifically downloads something, and then decides to consume selected content.
That's how search engines, web archives, web/content caches and so on, and well as services hosting user generated content that might be copyrighted, can operate without legal trouble.
As for AI and the output it generates, it's not going to be an exact copy of any of the data it was trained on, or a mix and match of sentences from different sources like some earlier "AIs" were. It's a bit more like if you study a bunch of books on a specific subject, then later decide to write your own book based on the knowledge and skills you obtained, that's not copying. That's just re-applying learned skills to write your own original work. It goes in the same category as clean room reverse engineering.
The same goes for somebody that decides to write their own song after listening to commercial music for decades. The song is probably not a direct copy or ripoff, but might be heavily influenced by all the music they have listened to of that style.
For reproducing information based on learned knowledge, it is important to distinguish between simply memorising something and trying to repeat things as accurately as possible from memory (that would still just be copying) vs getting an understanding of a subject as a whole, and re-applying that knowledge to make your own original work. AIs do something closer to the latter.
I am not necessarily in defense of the AIs here, nor should any of what I say be taken as legal advice. This is just my (mis?)understanding of how they get away with it.