Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

techcrunchtechcrunchtechcrunchNewly unredacted filings in the New York Times' three-year-old copyright lawsuit against OpenAI and Microsoft reveal that senior executives at both companies privately acknowledged their AI training practices amounted to theft and posed an existential threat to the publishers whose work powered their models.
The unsealed documents, made public on Thursday, show that Microsoft's Director of Applied Science, Brent Hecht, described the companies' mass scraping of news content as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history" in a January 2023 internal memo. In a separate presentation from January 2024, Hecht warned that Microsoft's AI-driven approach had created a "doom loop" that would "hurt the performance of our models and the entire web at the same time".wsj+1
Microsoft's own data showed its Copilot "answer engine" caused click-through rates for the Times' domain to drop as much as 93% compared to traditional Bing search. OpenAI's Head of ChatGPT, Nick Turley, wrote internally that publishers face an "existential threat" from products like the chatbot, which he described as "largely substitutive". OpenAI President Greg Brockman called the models "excellent at news," while Microsoft CEO Satya Nadella testified under oath that conversing with chatbots "has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source".techcrunch
The filings reveal for the first time the scale of the alleged copying: OpenAI's mid-training datasets alone contained more than 91,692 copies of works published by the Times, the Daily News, and the Center for Investigative Reporting, while a Common Crawl-derived dataset included more than 2 million documents from nytimes.com. The companies also allegedly assembled training data through initiatives called Project Taxi and Project Mango, the latter containing copies of at least 160,903 unique works from the news publishers.techcrunch
The filings describe how OpenAI employees allegedly devised methods to circumvent paywalls without detection. When researcher Nick Ryder told Brockman about a "hack to get around nytimes paywall," Brockman replied: "ah nice". Employees also allegedly stripped copyright notices from training data, since researchers "wouldn't want model outputting" such notices to users.techcrunch
The admissions could undermine OpenAI's fair use defense, particularly the requirement that use does not harm the market for the original work. Nadella himself testified that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training". Earlier this month, the Trump administration filed a brief in defense of OpenAI's unlicensed use of copyrighted material, according to TechCrunch. OpenAI and Microsoft did not return requests for comment.techcrunch