Major Publishers Sue Meta Over Llama AI Training Data Copyright Infringement
Five major publishing houses, including Elsevier, Cengage, Hachette Book Group, Macmillan Publishers, and McGraw Hill, along with bestselling author Scott Turow, have initiated a class-action lawsuit against Meta Platforms Inc. and its CEO, Mark Zuckerberg. Filed on May 5, 2026, in the U.S. District Court for the Southern District of New York, the suit alleges that Meta engaged in willful copyright infringement by using pirated works to train its Llama artificial intelligence models.
The core of the complaint centers on two main theories of liability. Firstly, the plaintiffs assert that Meta committed willful copyright infringement by torrenting over 267 terabytes of copyrighted books and journal articles from notorious pirate websites. They further claim that Meta masked its IP addresses to avoid detection during these illicit downloads and subsequently stripped the copyrighted materials of their copyright management information (CMI) to conceal their origins. This alleged removal of CMI is presented as a separate violation of copyright law.
Secondly, the lawsuit highlights the issue of demonstrable market harm. Unlike some previous AI copyright cases that struggled to prove direct harm, the publishers argue that Meta's Llama models are capable of producing full-length scientific papers, journal articles, replacement chapters for academic textbooks, and study guides. These outputs, they contend, directly substitute for the plaintiffs' original works, thereby circumventing existing licensing markets for AI training materials and directly depriving publishers of revenue. The complaint alleges that Meta abandoned legitimate licensing negotiations with publishers at Zuckerberg's direction, opting instead for unlawful sourcing.
This litigation is significant because it builds upon themes from recent copyright suits against AI developers but introduces new dimensions. For the first time, the plaintiffs include major publishing companies, not just individual authors, and the defendants extend to both the AI company and its CEO. The case could significantly influence how courts evaluate fair use in AI training contexts, particularly when plaintiffs can present compelling evidence of licensing market disruption and AI-generated content that directly competes with original copyrighted works. The publishers are seeking statutory damages, injunctive relief, and the destruction of infringing copies.
Read original source