Unsealed filings in the Times lawsuit quote a Microsoft executive calling AI scraping 'theft' of labor
The New York Times' brief also says OpenAI leaders privately called their models an 'existential threat' to publishers, though the underlying exhibits remain sealed.

Newly unredacted material in The New York Times' three-year-old copyright lawsuit against OpenAI and Microsoft includes an admission that AI training data scraping amounted to theft, TechCrunch reports. According to the filing, a top Microsoft executive privately described the companies' training practices as "theft", and OpenAI's own leadership said its models posed an "existential threat" to the publishers and journalists whose work trained them. The unsealed material also describes how the companies allegedly obtained and used that content, including bypassing paywalls undetected, building training datasets through mass scraping, and deliberately stripping copyright notices from training data. TechCrunch notes that much of the new information comes from the Times' own brief rather than the underlying exhibits, which remain sealed, and that the quotes are presented without their original context. The Times originally alleged that the firms violated copyright law by training generative models on its content. Whether training on copyrighted material is legal has no clear answer, but judges have so far been largely favorable to AI companies' argument that training constitutes fair use, the doctrine that permits unlicensed use in cases such as parody, news reporting and criticism. The unredacted statements give the Times new ammunition on the question of intent and could shape how other publishers pursue their own claims.