Tech companies are increasingly adopting controversial methods to feed their data-hungry artificial intelligence (AI) models, harvesting content from books, websites, photos and social media posts, often without the knowledge of the creators.
A recent investigation by Proof News has revealed that some of the wealthiest AI firms in the world have used subtitles from thousands of YouTube videos to train their AI, despite YouTube’s explicit rules against such practices.
The Pile
A research paper published by EleutherAI alleged that the dataset is part of a compilation the non-profit released called the Pile.
The developers of the Pile included material not just from YouTube but also English Wikipedia, the European Parliament and a trove of Enron Corporation employees’ emails released as part of a federal investigation into the firm.
The findings underscore the ethical and legal complexities surrounding the use of publicly accessible data for AI training and raise important questions about consent, compensation and the future of content creation in the digital age.
Proof News discovered that subtitles from 173,536 YouTube videos, sourced from more than 48,000 channels, were used by prominent Silicon Valley companies including Anthropic, Nvidia, Apple and Salesforce.
These videos span educational channels like Khan Academy and MIT, media outlets like The Wall Street Journal and BBC, and entertainment shows including The Late Show With Stephen Colbert and Jimmy Kimmel Live.
Popular YouTube creators, such as MrBeast, Marques Brownlee, Jacksepticeye and PewDiePie, also had their content included in the AI training datasets.
Flat-earth theory promoted
Some of this material even promoted conspiracy theories, such as the 'flat-earth theory'.
Dave Wiskus, CEO of Nebula, a creator-owned streaming service, labelled the unauthorised use of content as “theft” and “disrespectful”.
He voiced concerns about the potential exploitation of artists through generative AI, which could replace creators in various stages of content production.
Major companies like Apple, Nvidia and Salesforce have publicly acknowledged using the Pile for AI training.
Jennifer Martinez, a spokesperson for Anthropic, clarified that their use of the Pile is distinct from direct usage of YouTube’s platform, adhering to the dataset’s terms.
Salesforce highlighted that the dataset was publicly available for academic and research purposes.
That said, their research flagged potential issues with biases and safety concerns, given the presence of profanity and discriminatory language within the dataset.