Skip to content

Unsealed Court Filings Show Microsoft Executive Privately Called AI Training “The Largest Theft of Labor in Human History”

Getting your Trinity Audio player ready...

Unsealed Court Filings Show Microsoft Executive Privately Called AI Training “The Largest Theft of Labor in Human History”

A newly unsealed court filing in The New York Times’ long-running copyright lawsuit against OpenAI and Microsoft has surfaced a striking internal admission: a senior Microsoft executive privately described the industry-wide practice of scraping web content to train AI models as nothing less than history’s largest theft of labor.

The quote comes from Brent Hecht, Microsoft’s Director of Applied Science, who wrote the assessment in an internal company memo back in January 2023, nearly a year before the Times filed its lawsuit in December of that year. According to the unredacted filing, made public in Manhattan federal court on September 17, Hecht didn’t hedge in his own internal writing, describing the practice as an astonishing theft of unprecedented proportions and, in his own words, the largest theft of labor in human history. The timing matters here: Hecht wrote that assessment well before either company was facing formal legal exposure over it, which is part of why plaintiffs’ attorneys are now leaning on it so heavily as evidence.

The filing also points to comments attributed to OpenAI leadership, including cofounder and president Greg Brockman, characterizing the company’s own AI models as posing what’s described as an existential threat to the publishers and journalists whose work fed into training those very systems. Taken together, the picture emerging from the unsealed material is one where executives at both companies were privately acknowledging, in their own internal communications, a level of risk and ethical concern that looks difficult to square with the public legal defense both companies have mounted in the case.

Real More:  Sepehran Airlines Boeing 737 Suffers Mid-Air Structural Failure

Beyond the headline quote, the filing lays out new specifics about the scale of the alleged copying. According to the unsealed documents, OpenAI’s mid-training datasets alone contain more than 91,692 copies of articles published by the Times, the Daily News, and the Center for Investigative Reporting. Separately, a dataset derived from Common Crawl, a nonprofit that regularly scrapes and archives large portions of the public web, reportedly included more than 2 million individual documents pulled from nytimes.com alone. The filing also details specific methods the companies allegedly used to acquire that content, including scraping material indexed through Bing, bypassing paywalls without detection, and stripping copyright notices from documents before they were folded into training data.

One additional Microsoft document cited in the filing reportedly acknowledged what it called a real risk that generative AI could significantly disrupt the employment of the very people who generated the data used to train the underlying models, a striking admission given that Microsoft has spent the past several years positioning its own AI products, including Copilot, as tools meant to enhance rather than replace human work. Separate reporting on the filing has also pointed to a later internal warning suggesting Copilot was actively reducing the referral traffic that news outlets rely on to sustain reader revenue and advertising income, a dynamic that goes to the heart of the Times’ broader argument that AI chatbots don’t just use news content, they actively compete with and undercut the outlets that originally produced it.

Real More:  UAE Pushes to Exit Foreign Travel Warning Lists as Dubai Hotel Occupancy Hits Highest Level Since Iran War Began

It’s worth being precise about exactly what’s been revealed here and what hasn’t. Much of this new detail comes from the Times’ own legal brief rather than from the underlying exhibits and internal documents themselves, which remain under seal. That distinction matters, since a legal brief is inherently an advocacy document, written by one side to make the strongest possible case against its opponents, and the full context surrounding Hecht’s memo, along with any response Microsoft or OpenAI might offer explaining or contextualizing it, isn’t part of what’s currently public. Neither company has issued a detailed public response addressing Hecht’s specific quote as of the filing’s unsealing, and nothing in the unsealed material amounts to a legal finding or court ruling on the underlying copyright claims themselves.

The lawsuit itself dates back to December 2023, when the Times became one of the highest-profile publishers to sue OpenAI and Microsoft directly, alleging the companies used millions of the paper’s copyrighted articles without permission or compensation to train the large language models powering products like ChatGPT and Copilot. The case has become something of a bellwether for the broader legal reckoning playing out across the media industry, as numerous other publishers, authors and content creators have filed similar suits against AI companies over the past few years, all wrestling with the same fundamental question the Times case is trying to answer: whether training an AI model on copyrighted material without a license constitutes fair use, or whether it crosses into straightforward infringement at a scale the legal system has never really had to grapple with before.

Real More:  25 Years Since 9/11: How Four Hijacked Planes and Nearly 3,000 Lives Lost Reshaped America Forever

What makes this particular disclosure notable isn’t just the content of Hecht’s quote, but where it came from. Rather than relying purely on external analysis or third-party investigation to make its case, the Times’ legal team is now able to point directly to internal admissions from the defendants’ own senior staff, effectively letting Microsoft’s own applied science director make part of the argument for them. Whether that admission ultimately shapes the eventual outcome of the case, or gets contextualized away as an offhand internal comment that doesn’t reflect the companies’ formal legal position, is something that will play out as the litigation continues, but for now it’s given the Times’ side of the argument a considerably sharper edge heading into the next stages of the case.

Leave a Comment