“An astonishing theft of unprecedented proportions.” That is how Brent Hecht, Microsoft’s Director of Applied Science, described his own company’s AI training practices in a January 2023 internal memo, later adding that the mass scraping of journalism amounted to “largest theft of labor in human history.”
These are not a plaintiff’s accusations. They are a company’s own words, surfaced September 17, 2026, through newly unredacted filings in The New York Times’ copyright lawsuit against OpenAI and Microsoft, a case joined by the New York Daily News and the Center for Investigative Reporting.
What the Documents Actually Say
Internal communications describe the harm in language no public relations team would ever approve.
Hecht’s memo went further than calling the practice theft. A separate Microsoft document warned of a “real risk” that generative AI could significantly disrupt the employment of the very people who generated the training data. The filing directly links that scraping to threats against journalists’ jobs.
On the OpenAI side, Nick Turley, head of ChatGPT, internally characterized publishers as facing an “existential threat” from products he described as “largely substitutive” and likely to become more substitutive as they improve. Greg Brockman, OpenAI’s president, is quoted in the filings as saying the models are “excellent at news.”
That single word, “substitutive,” is the legal and ethical center of gravity here.
It is worth noting that much of this language appears in The Times’ summary-judgment brief, and some underlying exhibits remain sealed, so full original context is not always visible in reporting.
The Traffic Numbers and the Doom Loop
Microsoft’s own data predicted this outcome; the companies shipped the products anyway.
Internal Microsoft data cited in the filing shows Copilot’s answer engine produced up to 93% fewer referrals to The New York Times domain compared with traditional Bing search. Hecht’s January 2024 internal presentation, according to the filings, named this a “doom loop“: AI products suppress publisher traffic, weakening newsrooms’ capacity to produce content, which in turn degrades the content supply chain those same AI models depend on for quality.
The companies’ own analysts described this outcome plainly before it happened.
The Paywall Problem and What Nadella Said Under Oath
A deposition, a dataset, and an “ah nice” that may prove difficult to walk back.
The filings allege that OpenAI bypassed paywalls to obtain content. Copyright notices were then stripped from datasets before training, so models would not reproduce them in outputs.
As presented by plaintiffs, OpenAI researcher Nick Ryder informed Brockman of a “hack to get around [the] nytimes paywall,” to which Brockman replied “ah nice.”
Microsoft CEO Satya Nadella, in sworn deposition testimony, agreed that chatbot conversations can substitute for visiting original publisher sites. He stated that paywalled content used for training or grounding should be licensed. Nadella also testified he would have required OpenAI to retrain its models had he known otherwise.
Project Mango, a Microsoft-OpenAI collaboration, allegedly assembled a training dataset containing at least 160,903 unique works from the news publishers involved.
The Fair-Use Contradiction and Who Pays the Price
Publicly, these are transformative uses; privately, executives chose a different word entirely.
OpenAI and Microsoft publicly defend their training practices as fair use, arguing the process is transformative and does not replace the original works. Internally, their own executives used the word theft and acknowledged the substitutive impact on news audiences.
A federal government brief filed this month supports OpenAI’s fair-use position, adding a political dimension to a question courts have so far treated cautiously. Microsoft’s spokesperson characterized Hecht’s views as personal, not corporate policy.
Counsel for the New York Daily News said the evidence shows OpenAI and Microsoft “knew that what they were doing was wrong.” Both companies declined to comment on the specifics of the unredacted filings, according to multiple outlets.
The internal record, as presented by plaintiffs, points toward licensing frameworks rather than legal maneuvering as the more durable path forward. When your own documents call the practice theft, a better legal argument is not the fix; a different business model is.




























