Skip to content
SUNDAY, SEPTEMBER 20, 2026

Independently reported.

Tech

Microsoft Director Called AI Scraping the Largest Theft of Labor in History. It's Now a Court Exhibit.

An unsealed filing in The New York Times' copyright suit against Microsoft and OpenAI puts a name and a date behind a line that had circulated for weeks. Microsoft says it does not speak for the company. Its own click-through numbers, cited in the same filing, tell a harder story.

By Mara Voss, Technology

· 4 min read · Updated

A dim server room corridor with a laptop open on a desk in the foreground, no people, no visible text.
Illustration: Trestlewire

Key Takeaways

  • Microsoft's director of applied science, Brent Hecht, called AI web scraping "the largest theft of labor in human history" in a January 2023 internal memo, unsealed September 17, 2026.
  • OpenAI's mid-training datasets contain more than 91,692 copies of New York Times, Daily News, and Center for Investigative Reporting articles, per the unsealed filing.
  • Microsoft's own data shows Bing Copilot summaries cut click-through to nytimes.com by as much as 93 percent, a figure that roughly matches Satya Nadella's sworn testimony of a drop of more than 90 percent.
  • OpenAI's Nick Turley called chatbot products an "existential threat" to publishers and "largely substitutive" for journalism in internal communications from 2023 and 2024.
  • Microsoft spokesperson Alex Haurek says Hecht's memo does not represent company policy and points to Microsoft's court filings defending its practices as fair use.

An unredacted court filing made public on September 17, 2026 confirms who said it: Brent Hecht, Microsoft's director of applied science, wrote in a January 2023 internal memo that AI companies scraping the web for training data amounted to "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The line had circulated in paraphrase for weeks. Now it has a name, a date, and a paper trail attached, courtesy of The New York Times' copyright lawsuit against Microsoft and OpenAI.

The short answer

Brent Hecht, Microsoft's director of applied science, called AI training on scraped web content "the largest theft of labor in human history" in a January 2023 internal memo. The quote became public on September 17, 2026, in an unsealed filing from The New York Times' copyright suit against Microsoft and OpenAI. Microsoft says the memo reflects one employee's opinion, not company policy, and points to its own court filings on fair use.

The paper trail behind one sentence

The filing, unsealed as part of a summary judgment motion, lays out how OpenAI and Microsoft acquired New York Times content well before the lawsuit was filed. According to TechCrunch, OpenAI's mid-training datasets alone hold more than 91,692 copies of articles from the Times, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset pulled in more than two million documents from nytimes.com by itself, and a Microsoft-OpenAI data-sharing effort called Project Mango assembled at least 160,903 unique works from news publishers.

93%

Drop in NYT click-through via Bing's Copilot summaries

Figure comes from Microsoft's own internal data, cited in the unsealed filing, comparing Copilot-generated answers against a traditional Bing search results page.

A 'doom loop,' documented a year later

Hecht did not let the subject drop. In a January 2024 internal presentation, he warned that Microsoft's AI content strategy had triggered a "doom loop" that would hurt the performance of its models and the entire web at the same time. The same document called the situation highly unusual: an end product competing with the economic foundations of the suppliers it depends on. Separately, CEO Satya Nadella testified under oath that clicks to news sites fell by more than 90 percent on Bing after AI summaries took over, a figure 404 Media reported lines up with the 93 percent number in Microsoft's own filing.

It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain.

Brent Hecht, Director of Applied Science, Microsoft, internal document, January 2024

OpenAI's own paper trail

OpenAI executives left a similar record. Nick Turley, who leads ChatGPT, wrote in mid-2023 that publishers faced an "existential threat" from products that are "largely substitutive" for journalism, then followed up in February 2024, per Futurism, to say the substitution problem would only grow as the models improved. OpenAI president Greg Brockman called the models "excellent at news" in an internal exchange and, according to the filing, replied "ah nice" after a researcher told him about a way to bypass the Times' paywall undetected.

Microsoft's defense: one employee, not the company

Asked about Hecht's memo, Microsoft did not disown the document, just the conclusion. "Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism," spokesperson Alex Haurek said, in a statement reported by Engadget. Neither Microsoft nor OpenAI answered a separate request for comment from TechCrunch about the filing itself. Steven Lieberman, counsel for the New York Daily News, took the opposite view in a statement to TechCrunch: "The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong."

Real friction, not a slam dunk

None of this settles the underlying legal question. Several AI copyright cases have already gone the industry's way, with judges ruling that training on copyrighted material counts as transformative fair use rather than infringement, per Engadget. Those same judges have been explicit that the law here is unsettled, not resolved in AI companies' favor as a rule. The Times' case is treated as a bellwether for exactly that reason: it is the one most likely to test whether a researcher's own words, calling the product 'a product that destroys its supply chain,' can survive a fair-use defense once a jury reads them.

For a company that has spent three years telling advertisers, publishers, and regulators that generative AI creates more value than it destroys, having a senior researcher's own phrasing turn into a court exhibit is a specific kind of bad. The 91,692 copied articles are a discovery problem. The doom-loop memo is a strategy problem. The line about the largest theft of labor in human history is the one that will keep getting quoted, mostly because Microsoft wrote it first, in its own building, three years before anyone outside subpoenaed it.

  • Microsoft
  • OpenAI
  • AI copyright lawsuit
  • New York Times
  • web scraping
  • fair use

About the reporter

Mara Voss

Technology Reporter, Trestlewire

I spent seven years as a product manager at a mid-size SaaS company before I ever wrote a sentence for pay, which means I have sat through more roadmap reviews than most people would tolerate in a lifetime. I watched a scheduling feature get rebranded three times before it shipped, and I watched a launch date slide past four straight quarters while the slide deck stayed exactly the same. That is where the question I still ask every day came from: does this actually ship, or is it a demo.

Read full bio and all stories →