OpenAI, Microsoft Execs Called AI Scraping 'Theft'
Unsealed documents reveal Microsoft and OpenAI executives privately called AI data scraping 'theft'.
"AI giants got caught calling their own data scraping 'theft.' Talk about a self-own. This isn't just a lawsuit; it's a look behind the curtain."
Newly unsealed court filings in a copyright lawsuit reveal internal communications from Microsoft and OpenAI executives describing AI training practices as "theft." These documents indicate that both companies scraped paywalled content from news outlets, including The New York Times, to build datasets. Internal warnings also suggested that these practices would significantly impact publishers.
The unredacted filings detail how content was allegedly obtained by bypassing paywalls and stripping copyright notices from training data. A Microsoft executive reportedly called the practice the “largest theft of labor in human history.” OpenAI's leadership also acknowledged that its AI models posed an “existential threat” to publishers and journalists whose work was used for training.
These admissions challenge the "fair use" defense often cited by AI companies. Microsoft's own data showed its Copilot "answer engine" caused a significant drop in click-through rates for The New York Times' domain. A Microsoft document also noted a "real risk" of generative AI disrupting employment for those who created the training data.
This development highlights the growing legal and ethical challenges surrounding AI training data. Businesses relying on AI models must be aware of potential copyright infringement risks and the implications for content creators, as these cases could reshape data acquisition practices and intellectual property rights in the AI landscape.
Relevant tools
Find the right AI tool for your business
Chat with Insta and get matched to the right tool in seconds.
Try Insta Tool Finder →