Brent Hecht, Director of Applied Science at Microsoft, privately characterized the unlicensed bulk use of others' copyrighted content to train AI models as "the largest labor heist in human history." Around the same time, Nick Turley, who leads ChatGPT at OpenAI, described ChatGPT in internal communications as an "existential threat" to publishers. These statements weren't leaked by reporters — they came from documents Microsoft itself submitted, which The New York Times' legal team then pulled out of court filings. One face for the public, preaching fair use; another face for the inside, admitting theft — both laid side by side before a judge.

Imagine you sneak into your neighbor's orchard, pick all the apples, and open a juice shop to sell them. Once business takes off, you publicly proclaim "this is technological innovation, a win-win for everyone." But at a family dinner, your dad leans over to your mom and mutters, "We're running the biggest apple heist in human history." Even better, your sales director marches up to the orchard owner and says, "If you don't work with us, your orchard will be bankrupt within three years." That's the awkward scene Microsoft and OpenAI left behind in the NYT lawsuit filings. Enough with the analogy — the real difference is this: the neighbor's apples have been swapped out for billions of articles, books, and papers; the juice shop is now an AI company worth hundreds of billions of dollars; and the transcript of that "family dinner" is being page-by-page unearthed during federal court discovery.
The Incident

Microsoft's Applied Science Director's private "largest-scale theft" comment gets dug up

OpenAI and Microsoft tell one story in court, and a different one in their own internal documents. The New York Times' side just pulled the latter out of legal filings.

The discovery phase (the evidence-exchange stage) of The New York Times v. OpenAI lawsuit has surfaced a new batch of documents. In them, Brent Hecht, Director of Applied Science at Microsoft, privately calls AI companies' practice of bulk-scraping others' copyrighted content without permission to train models "the largest labor heist in human history."

Hecht got more specific: the scale is "unprecedented, because the target is the labor output of every country, every language, and every field." A global content harvest — other people's writing, reporting, and research stuffed into AI models, without asking first and without paying after.

2
Legal filings drawn from Microsoft and OpenAI's own internal records
Source: Tom's Hardware
July 9, 2026
The publisher coalition led by the NYT has asked the court to sanction OpenAI
Source: Tom's Hardware
December 16, 2025
OpenAI formally filed its motion for summary judgment
Source: Tom's Hardware

The two faces sit side by side before the judge: publicly, Microsoft champions "fair use" (i.e., legally permitted unlicensed use) for AI training; privately, its own director says the business runs on theft. The defense arguments on one side, the emails on the other.

Both statements came from internal communications and entered the record through the defendants' own submissions — to deny them, they'd first have to deny their own documents.

On July 9, the NYT legal team separately moved for sanctions against OpenAI in Manhattan federal court, accusing the company of "choosing to obstruct" and repeatedly making "false statements" for two years on matters of data destruction and ChatGPT logs (i.e., records of user-AI conversations).

So what new card does this hand them? Microsoft's own people have admitted the theft — and the other side is still digging in at trial.

Why It Matters

Internally calling it theft, publicly calling it fair use — the gap is the whole case

The publishers' real target isn't the scraping. It's the "willful misconduct" claim: OpenAI knew the practice was legally contested, said so in its own documents, and still told the public it was fair use.

That's what the internal records buy the media coalition — the word "knew," in the defendants' own words, on the record.

AI companies' practice of bulk-scraping news, books, and papers to train models(the process of using massive datasets to teach AI to speak and write) is the standard approach for this generation of AI products.

What the coalition led by the New York Times wants isn't to shut AI down — it's to put a price on this practice: either pay for licensing, or face legal liability. Every inch the evidence tips the scale affects not just the damages owed by the two companies, but how — and at what cost — every future AI product gets its content.

If fair use is the honest defense, why does the word "theft" only show up in private?