The New York Times has moved to amend its copyright lawsuit against Microsoft and OpenAI, alleging that Microsoft deliberately constructed a bespoke supercomputing system to facilitate the unauthorized use of the newspaper’s copyrighted works.
Legal shift and amended claims
The motion follows a Supreme Court ruling that raised the bar for contributory infringement claims, requiring plaintiffs to prove intentional inducement of illegal conduct. The NYT now seeks to align its claim against Microsoft with this new standard, arguing that the supercomputer was not a generic cloud service but a purpose-built machine designed to train large language models on copyrighted material.
“Today, we asked the court for permission to file an amended complaint that further strengthens our case, clarifying our claim of contributory infringement against Microsoft based on new law and new evidence uncovered during discovery,” a NYT spokesperson stated. The newspaper also agreed to voluntarily dismiss two claims of contributory copyright infringement and trademark dilution against all defendants.
Microsoft’s supercomputer as infringement tool
The amended complaint specifies that Microsoft’s supercomputer—ranked among the world’s most powerful—was tailor-made to help OpenAI infringe. The NYT alleges that the system was built to train AI on essentially the entire internet, with Times articles disproportionately weighted to produce high-quality outputs.
“Microsoft specifically designed it for the purpose of using essentially the whole Internet—curated to disproportionately feature Times Works—to train the most capable LLM in history,” the NYT alleges. The newspaper further claims that Microsoft’s deployment of Times-trained models across its product line boosted its market capitalization by a trillion dollars in the past year.
Evidence of market substitution and hallucinations
Discovery revealed significant evidence of market harm, including ChatGPT sessions where users bypassed paywalls and obtained near-verbatim excerpts of copyrighted articles. The complaint includes side-by-side comparisons showing models reproducing NYT content without authorization.
Equally damaging are hallucinations where AI systems falsely attribute fabricated content to the NYT. Examples include Bing Chat citing fake quotes and ChatGPT inventing a non-existent article linking non-Hodgkin’s lymphoma to orange juice. “Users who ask a search engine what The Times has written on a subject should be provided with neither an unauthorized copy nor an inaccurate forgery,” the NYT argued.
Fair use defense and stakes
OpenAI maintains that training on publicly available data constitutes fair use. “Our models empower innovation, are trained on publicly available data, and are grounded in fair use,” a spokesperson said. However, the NYT’s evidence of market substitution directly challenges this defense. A federal judge has previously suggested that proving market harm could be a winning argument against fair use claims in AI training cases.
If the NYT prevails, the most severe outcome could require OpenAI and Microsoft to wipe and retrain their models. The newspaper also seeks permanent injunctive relief and extensive damages, arguing that defendants have “wrongfully profited from copyrighted works that they do not own.”
This case represents a pivotal test of whether AI training on copyrighted material constitutes fair use or infringement, with implications for the entire generative AI industry. The outcome will likely shape how technology firms build and deploy large language models, particularly regarding the sourcing and weighting of training data.
