Newly unsealed court documents reveal tech executives privately called AI data scraping "the largest theft of labor," undermining fair use claims
Newly unsealed court documents from the ongoing copyright lawsuit filed by The New York Times and other news outlets against OpenAI and Microsoft have revealed striking internal admissions from top tech executives.
As reported, publishers argue that these leaked communications eviscerate the tech giants' legal defense of "fair use," proving that both companies knowingly exploited copyrighted journalism.
Microsoft Director of Applied Science Brent Hecht internally characterized the data scraping required for AI training as "an astonishing theft of unprecedented proportions" and potentially "the largest theft of labor in human history."
Hecht further noted that winning a fair use defense on these grounds would "make a complete mockery of the idea of fair use."
OpenAI Head of ChatGPT Nick Turley admitted in internal chats that their products are "largely substitutive, period," warning that commercial AI models pose an "existential threat" to the survival of news publishers by diverting readers and cutting off traffic.
Court filings claim that OpenAI leadership utilized technical workarounds and "hacks" to bypass paywalls to extract training data, while Microsoft CEO Satya Nadella testified that he would have forced OpenAI to retrain its models from scratch had he known paywalled content was being scraped.
Internal Microsoft presentations outlined a self-defeating "doom loop," acknowledging that degrading or bankrupting the news industry would ultimately hurt the quality of the web and data supplies that future LLMs rely on.
To successfully claim fair use under U.S. copyright law, tech companies must prove that their use of copyrighted material is transformative and does not severely harm the market value of the original work.
News outlets argue that these unsealed documents serve as a smoking gun—proving that tech executives privately knew their chatbots served as direct substitutes for traditional reporting, caused massive drops in referral traffic, and relied on unauthorized, systemic scraping of copyrighted material rather than legally protected transformation.