“An astonishing theft of unprecedented proportions”: Court records show what Microsoft and OpenAI actually thought about AI training

1 hour ago 1
Add to circle

Would you call large language models training on an internet’s worth of content “an astonishing theft of unprecedented proportions”? Or perhaps you prefer the “largest theft of labor in human history”?

Take your pick because those two options — and many more — appear in a newly unredacted court filing in the ongoing New York Times v. OpenAI/Microsoft case. The motion was filed by a group of news publishers led by The New York Times. The most inflammatory language was not written by the Times’ lawyers, however, but by OpenAI’s and Microsoft’s own executives.

“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” wrote Brent Hecht, Microsoft’s director of applied science, in the unearthed documents. “Almost no one intended for content they created to be used in this fashion, nor are they compensated for its use,” he wrote.

AI companies have publicly argued that their training falls under the legal doctrine of fair use. Publishers, however, have long warned that the AI companies’ plan to take content without compensation will severely undermine the news industry whose work trains those models. It appears at least some within the AI industry agree with the assessment.

Internal Microsoft communications identified the “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” A separate policy document put it like this: “LLMs are a product that destroys its supply chain.”

The current approach has created a “doom loop,” according to another internal Microsoft document.

“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time,” the document reads. “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

Elsewhere, Nick Turley, OpenAI’s head of ChatGPT, acknowledged that publishers face an “existential threat” from OpenAI’s products. The AI tools “are largely substitutive, period” and “will get more and more substitutive as they get better.”

OpenAI president and co-founder Greg Brockman told his colleagues: “[W]e are excellent at news btw. every time i do any generative stuff on NYT it seems to predict the next sentence pretty well.” Brockman also wrote that OpenAI’s model “seems to be particularly good at predicting text of news articles like whenever i have it complete in the middle of a sentence in a NYT article, it seems to complete the sentence on point.”

The filing also shows that OpenAI employees went out of their way to evade news paywalls to scrape websites. (When told about “a hack to get around nytimes paywall,” Brockman responded “ah nice.”)

Previously, these statements were heavily redacted or under seal. Some portions remain sealed.

“Easy to see why the AI companies had so many redactions,” wrote Digital Content Next CEO Jason Kint, who first shared the unredacted documents on X. “Their own people wrote the lede for NYT.”

You can read the court documents here.

Read Entire Article