Major Publishers Sue Meta Over Llama AI Training Data Copyright Infringement
In a significant legal challenge that could redefine the landscape of AI development, a consortium of prominent publishing houses, including Elsevier, Cengage, Hachette Book Group, Macmillan Publishers, and McGraw Hill, alongside bestselling author Scott Turow, has initiated a copyright infringement lawsuit against Meta Platforms, Inc. and its CEO, Mark Zuckerberg. The complaint, filed in the U.S. District Court for the Southern District of New York, accuses Meta of illicitly using vast quantities of copyrighted material to train its Llama large language models (LLMs).
The plaintiffs contend that Meta engaged in willful copyright infringement by torrenting millions of copyrighted books and journal articles from notorious pirate websites. Furthermore, the lawsuit alleges that Meta deliberately removed copyright management information (CMI) from these works, a move designed to obscure the original sources and ownership of the content. This action, according to the publishers, constitutes a violation of copyright law and demonstrates a clear intent to bypass legitimate licensing processes.
A central tenet of the publishers' argument is the assertion of demonstrable market harm. They claim that Meta's Llama models are capable of generating full-length scientific papers, journal articles, replacement chapters for academic textbooks, and study guides, among other materials. These outputs, the plaintiffs argue, directly substitute for their copyrighted works, thereby directly depriving them of revenue and circumventing existing licensing markets for AI training materials. This aspect of the case is particularly crucial, as establishing market harm has been a significant hurdle for plaintiffs in previous AI copyright suits.
The lawsuit builds upon themes that have emerged in earlier copyright litigation against AI developers but introduces new dimensions. Notably, the plaintiffs now include publishing companies, not solely individual authors, which could strengthen the argument regarding organized market disruption. The case is expected to scrutinize two key fair use issues: the legality of sourcing training data from unlawful channels and the extent to which AI-generated content directly competes with and harms the market for original copyrighted works. The outcome of this litigation could set a precedent for how courts evaluate fair use in the context of AI training, especially when evidence of licensing market disruption and AI-produced substitutes for original works is presented.
Read original source