Microsoft exec: AI scraping may be the largest labor theft

Published · AI Daily — AI-assisted deep research, methodology & disclosure

Ars Technica reports that a summary judgment motion from news plaintiffs led by The New York Times has been unsealed. It quotes Microsoft applied science director Brent Hecht warning that scraping news for AI training may be the "largest theft of labor in human history." Microsoft says the documents reflect one employee's view. The plaintiffs also cite steep click-through declines.

What happened

A motion for summary judgment from the news plaintiffs, led by The New York Times, was unsealed on Thursday. According to Ars Technica, the filing exposes internal Microsoft and OpenAI documents that the two companies had fought to keep out of public view. The news organizations argue that these documents show how both firms understood the threat to news before they launched products such as ChatGPT and Copilot.

The most striking line comes from Brent Hecht, Microsoft's Director of Applied Science. The news organizations say he warned repeatedly that scraping news for AI training was "an astonishing theft of unprecedented proportions", and that it was perhaps the "largest theft of labor in human history". In another document, they say, he wrote that the plan to scrape news widely made "a complete mockery of the idea of 'fair use.'" Fair use is the exact defense that Microsoft and OpenAI rely on in this case.

Microsoft disputes how these documents should be read. Its spokesperson says the Hecht documents "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views." OpenAI had not responded to a request for comment when Ars published.

Key facts from the filing

The article reports the following points, all of them as claims made by the news plaintiffs or as quotes from documents they cite.

  • **A "doom loop".** A Microsoft document described a "doom loop" that "will hurt the performance of our models and the entire web at the same time." The same document said: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"
  • **"Existential threat".** OpenAI's head of ChatGPT, Nick Turley, wrote in an internal message that publishers would face an "existential threat" from commercial products trained on news that can substitute for news providers.
  • **Click-through data.** Ars reports that Microsoft recorded drops in click-through rates of 83 to 93 percent for some news plaintiffs, and 51 to 94 percent for others.
  • **Compensation.** Hecht acknowledged in a Microsoft document that "almost no one intended for content they created to be used in this fashion, nor are they compensated for its use."
  • **The paywall exchange.** Satya Nadella testified under oath that AI companies should not violate news sites' terms of use by dodging paywalls. Yet OpenAI's internal messages show that when Nick Ryder told President Greg Brockman that "a hack" had been found for OpenAI crawlers "to get around" the NYT paywall, Brockman replied, "Ah, nice."
  • **Substitution.** Nadella also acknowledged that chatbots act as substitutes, taking clicks by "giving you the information right there on the website on the AI platform versus needing to go to the underlying source." An OpenAI software engineer wrote that "no matter how prominently we show the links, users won't click." Turley said there is "no good reason to click" when the chatbot provides the information. He called chatbots "largely substitutive, period", and predicted they "will get more and more substitutive as they get better."

How the plaintiffs tested for verbatim output

The motion also describes how the news groups probed the products. They went beyond the early method of asking a chatbot "what's the next line?" again and again. In some cases, chatbots produced long excerpts of articles when users asked for summaries. Other outputs came from requests for the key bullet points of an article. Prompts that asked chatbots to "rate the bias" of an article were particularly successful. Chatbots also reproduced parts of articles when users asked them to pick any article from a site's homepage.

The plaintiffs ask the court to rule only on articles where the outputs "demonstrate extensive verbatim overlap". They say they are confident that the "substitutive purposes of defendants' copying weigh against fair use". Legal concerns about other articles would be raised at trial.

The plaintiffs also allege that the two firms built a filter that made it harder for news groups to test the chatbots. Hecht suggested that such a filter could be seen as an "accidental cover up", because it would leave "people who have a right over the content having less visibility into what was used for training."

Licensing and data sourcing allegations

The news groups say the insiders' warnings did not lead Microsoft and OpenAI to license news content. They allege that this let the firms usurp the publishers in another market in ways the publishers could not anticipate. Two specific claims appear in the motion. First, Microsoft allegedly violated "industry norms" by selling a dataset it bought for Bing as training data for OpenAI, without asking news groups that would not have approved of that reuse of their consent to ordinary search-engine crawling. Second, OpenAI allegedly "acted improperly" by obtaining a NYT dataset of 1.8 million articles from a third party that was bound by an agreement not to allow commercial use. OpenAI employees knew it "would not be appropriate" to use that data "to train a model", the plaintiffs allege, and did it anyway.

These are allegations in a motion. The article does not report any court finding on them.

Background: why the substitution argument matters

Fair use is a flexible legal test. A central question in it is whether a copy harms the market for the original work. That is why the plaintiffs focus on substitution. If a chatbot replaces the news site for the reader, and if it also repeats article text word for word, the plaintiffs think the harm to their market is clear. The filing says that "courts have rejected claims that copying news articles to provide a product that substitutes for demand for news is fair use."

The plaintiffs also frame the issue as a collective action problem. In their words, AI companies "remain powerless to break out of this 'doom loop,'" because "each individual company is better off taking content for free while others pay." They say a finding that copying news for AI is not fair use "would solve this prisoners' dilemma by putting all AI companies, OpenAI and Microsoft included, on an even footing." They point to Google's AI Overviews, which arrived soon after ChatGPT launched and absorbed more of the traffic that once went to news sites.

Our analysis

We read this as a story about evidence, more than about any single quote. The "theft" language is vivid, and it will get the headlines. But Microsoft's reply is a fair legal point: an internal opinion is not a legal conclusion. What may matter more in court is the pattern across many documents, from different people at two companies, that describe the same effect. Engineers, product heads and a CEO are quoted saying that users do not click and that chatbots substitute for the source.

The click-through numbers are the hardest part of the case for the defendants to wave away. They come from the companies' own data, according to the article. A court still has to decide whether lost clicks come from copying, from a new way of finding information, or from both. Nadella's testimony, as Microsoft describes it, dealt with "changes underway in how people find and consume information", not with copyright law. That gap between a market trend and a legal wrong is where the fair use fight will be decided.

The paywall exchange is a separate issue. It concerns how data was gathered, not how outputs look. It could affect how a judge views the good faith of the defendants, though we cannot say how much weight a court will give it.

Limits and open questions

This article is one reporter's account of one side's motion. The plaintiffs chose which documents to highlight and how to frame them. Microsoft and OpenAI have their own filings, and Ars did not quote OpenAI's position at all.

Some quotes come to us through the plaintiffs' description, so their full context is not visible. The motion asks for a ruling on only a subset of articles, and the article does not say when the court will decide or whether the case will go to trial. Nothing here is a finding of infringement.

Takeaways for readers

  • **Policy watchers:** watch how the court treats internal opinions and usage data. This ruling could shape whether licensing becomes the norm for news content.
  • **Publishers:** the case suggests that click-through and substitution data may become key evidence in disputes with AI firms.
  • **AI product teams:** internal documents can become public in litigation.

Statements about risk, data sourcing and paywalls may be read by a judge later.

  • **Everyone else:** treat leaked one-sided quotes with care. Wait for the response filings and the ruling before you draw firm conclusions.

Sources

FAQ

What did Brent Hecht say?

The news plaintiffs say Microsoft's Director of Applied Science repeatedly warned that scraping news for AI training was "an astonishing theft of unprecedented proportions" and perhaps the "largest theft of labor in human history." In another document he reportedly said it made "a complete mockery" of fair use.

How does Microsoft respond to the documents?

A Microsoft spokesperson says the Hecht documents "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views." The company defends its AI products as transformative fair use that does not substitute for news sites. OpenAI had not responded when Ars published.

What click-through data does the motion cite?

Ars reports that Microsoft recorded click-through drops of 83 to 93 percent for some news plaintiffs and 51 to 94 percent for others. The plaintiffs use this, with internal statements, as evidence of substitution.