
Unsealed filings: Microsoft researcher called AI scraping "the largest theft of labor in human history"
Newly unsealed filings in The New York Times' lawsuit against OpenAI and Microsoft quote a senior Microsoft researcher calling AI training "theft" and show that Copilot cut clicks to The Times' domain by as much as 93%.
Newly unsealed filings in The New York Times' copyright lawsuit against OpenAI and Microsoft show that a senior Microsoft researcher privately described the companies' AI training practices as theft, while OpenAI's own leadership called chatbots an existential threat to the publishers whose work trained them.
What the documents say
The filings quote a January 2023 internal memo by Brent Hecht, Microsoft's director of applied science, who called news scraping "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." A presentation he wrote in January 2024 described an internal "doom loop": Microsoft's own data showed that its Copilot answer engine cut click-through rates for The Times' domain by as much as 93% compared with traditional Bing search, a decline the document said would "hurt the performance of our models and the entire web at the same time."
A separate Microsoft document warns of a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained." OpenAI's head of ChatGPT, Nick Turley, wrote internally that publishers face an "existential threat" from products that are "largely substitutive" and "will get more and more substitutive as they get better." OpenAI president Greg Brockman called the models "excellent at news."
Scale of copying and paywalls
The documents state that OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by The Times, the Daily News and the Center for Investigative Reporting, while a Common Crawl-derived dataset included more than 2 million documents from nytimes.com. Data assembled under the companies' "Project Mango" initiative contained at least 160,903 unique works from the news publishers.
The filing also describes plans to bypass paywalls without detection. When OpenAI researcher Nick Ryder told Brockman about a "hack to get around nytimes paywall," Brockman replied: "ah nice." Microsoft CEO Satya Nadella testified in a deposition earlier this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training," and said he would have required OpenAI to retrain its models had he known about paywalled training data.
Why it matters
The disclosures cut against pillars of OpenAI's fair-use defense, which requires that a use not substitute for the original work or harm its market. Courts have so far leaned toward AI companies on fair use. Some caution is warranted: much of the new material comes from The Times' own brief rather than the underlying exhibits, which remain sealed, and the quotes are presented without their original context. OpenAI and Microsoft did not return requests for comment.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.