Can training an AI model on copyrighted books ever qualify as fair use? That question of AI training fair use sits at the heart of one of the most consequential copyright disputes of the AI era. On July 20, 2026, the Northern District of California approved a class action settlement awarding a class of copyright holders $1.5 billion against Anthropic PBC, the creators of the Artificial Intelligence Large Language Model (“LLM”) Claude.
This litigation arose out of a dispute involving Anthropic’s use of the various authors’ original works and books; Anthropic purchased and pirated millions of books to create a large central library to train and perfect Claude. While this may seem like a victory for the authors, strengthening copyright safeguards and protecting human created works in a society ever dominated by AI, decisions throughout this litigation suggest the opposite.
Prior to approving the largest copyright settlement in history, the Court decided that Anthropic’s purchasing and digitization of copyrighted material for storage in a large central library to then copy the material for the purpose of training their LLM was quintessentially transformative and protected fair use under Section 107 of the Copyright Act.
Background
In August 2024, authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson filed suit against Anthropic PBC for allegedly infringing on their individually copyrighted books. While creating Claude, Anthropic wished to create a digital library to permanently store “all the books in the world.” To achieve this, Anthropic pirated over seven million books from multiple websites including “LibGen” and “PiLiMi,” well known online pirated material libraries, to then store in a centralized library for possible future use. Additionally, Anthropic later purchased millions of print books which were then digitally scanned with the physical copy discarded. Once in the central library, some books were copied and modified into “tokenized” versions, allowing the AI to measure the statistical gaps between word stems and language use, ultimately training the LLMs to perfectly mirror human speech.
In response to this joint suit, Anthropic moved for summary judgment, claiming their use of the copyrighted material, regardless of any pirating, constituted protected fair use.
Anthropic’s AI Training Fair Use Defense
Anthropic’s motion for summary judgment claimed that inputting copyrighted material into their LLM for training purposes, along with the digitization and retention of both the pirated and purchased books in a central library, was fair use. The Court considered the following four factors when addressing each of Anthropic’s three disputed uses to determine whether they constituted fair use:
- the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
- the nature of the copyrighted work;
- the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
- the effect of the use upon the potential market for or value of the copyrighted work.
Training Claude Was Transformative; Pirating Books Was Not
First, the Court held that inputting copyrighted material into Claude for training purposes, even if resulting in multiple copies or “tokenized” versions being made, was “quintessentially transformative,” and thus fair use. No contentions were made that Claude would output the copyrighted works word-for-word to users; if this were the case the Court explained their ruling would have turned out differently. Rather, the LLM studied language patterns and speech principles to create something new. While the authors fear this could allow Claude users to create works that can compete in the market with their own, the Court explained that copyright does not extend to methods of operation or various speech principles illustrated in these protected books but rather allows for the creation of a new and transformative project. To reference an analogy used by the Court, think of this AI training process like teaching children to write using copyrighted books; similarly, AI uses copyrighted materials to establish a foundation that is then built upon into something bigger and better.
Second, the Court held that making digital copies of purchased books to create a central library that could possibly be used for training purposes, while destroying the original physical copy, was also fair use. The Court explained that so long as the digital copies are not shared, shown, or sold outside of this central library, and no contentions were made that this was the case, one who fairly purchases this material can use it as they see fit. However, the books pirated by Anthropic for this same use were not protected, regardless of any eventual fair use down the line.
Ultimately, the Court partially granted Anthropic’s motion for summary judgment, holding the training of the LLM using copyrighted material, and the digitization of purchased books into a central library, to be fair use. However, the authors’ claims that Anthropic digitized and retained pirated copies of their works, and allowed additional copies to be created of the books stored in the central library by employee engineers for other training uses, survived summary judgment.
Class Certification
Despite Anthropic’s small victory at the summary judgment stage, a class of copyright owners was certified on July 17, 2025. The certified 23(b)(3) class represented copyright holders, authors, publishers, and any beneficiary with the right to make reproductions, whose books were pirated from the “LibGen” and “PiLiMi” sites only. Overall, notice was sent to 594,945 potential class members, encompassing over 480,000 books on the identified pirated works list; as of April 2026, over 91% of the books on the works list were claimed. Currently, the statutory minimum damages for ordinary copyright infringement are $750; with such a large class certified and the potential for statutory damages to exceed the $750 minimum per instance of infringement, a massive payout was potentially on the horizon.
The Largest Copyright Settlement in History Prevents Lengthy and Costly Litigation
Ultimately, following the settlement’s preliminary approval in September 2025, the $1.5 Billion settlement was approved on July 20, 2026. With 91% of the works currently being claimed on the class works list, roughly $3,000 will be paid to each copyright holder per work on the list, 4 times the statutory minimum for copyright infringement. While a strong case likely existed to receive more than $3,000 per work, this payout wasn’t guaranteed; a settlement avoided the risks of outright losing or enduring a long and costly trial while ensuring Anthropic’s financial capability for the payout. Although the copyright holders released Anthropic from all present claims, the settlement agreement contained future looking protections. Per the agreement, Anthropic must destroy all pirated copies stored in the central library and remains vulnerable to claims for future instances of infringement for similar actions.
The Future of AI
Despite the settlement in Bartz v. Anthropic representing the largest copyright settlement in history, this case may establish future protections for AI companies and LLMs. The Copyright Act aims to protect and encourage originality. However, in the Court’s view, training artificial intelligence using original copyrighted material is “quintessentially transformative,” opening the door to AI being trained by virtually any material. For copyright owners, yes, companies can use protected material to improve their AI, but they must do so carefully, making sure to remain transformative to maintain the integrity of the fair use system. As AI training fair use continues to be litigated across the industry, businesses that build or train AI models—and the authors and publishers whose works may be swept in—should watch this space closely. If your organization has questions about AI training fair use, copyright risk, or a licensing strategy for AI development, our intellectual property team can help you navigate what comes next.