top of page
wood_phillips_hero_gradient_135deg.png

The $1.5 Billion Distinction: Why Anthropic Settled the Books It Downloaded, Not the Models It Trained

  • Writer: JGordon
    JGordon
  • 4 hours ago
  • 5 min read

Most copyright headlines this week reported the same number: $1.5 billion, the largest copyright class action settlement on record. The number is real, and it is worth pausing on. But the number is not the lesson.


On Monday, July 20, 2026, U.S. District Judge Araceli Martínez-Olguín granted final approval to the class settlement in *Bartz v. Anthropic*, resolving claims covering 482,460 books and paying participating authors and publishers roughly $3,000 per work. The case began in 2024, when thriller novelist Andrea Bartz and two co-plaintiffs sued the AI company over the books behind its Claude chatbot. It ends with a payout, a destruction order, and a distinction that every client on either side of the AI economy should understand.


That distinction is this: **Anthropic did not pay $1.5 billion for training an AI model on books. It paid $1.5 billion for how it got the books.**


## Two Questions the Court Kept Separate


The now-retired Judge William Alsup, who handled the case through preliminary approval in September 2025, issued a ruling last year that cut the dispute cleanly in two.


On the first question — whether training a large language model on copyrighted books infringes — he found for the company. Training was transformative enough to sit within fair use. Anthropic's deputy general counsel highlighted exactly that holding in the company's statement on final approval, and it is fair for them to do so. It is a meaningful ruling.


On the second question, the court went the other way. Anthropic had downloaded millions of books from LibGen and PiLiMi, two repositories of pirated material, and retained them in a permanent internal library. The court declined to let a lawful downstream use launder an unlawful acquisition. Fair use protects what you *do* with a work you legitimately possess; it is not a warrant for how you came to possess it in the first place. A damages trial on the pirated copies was on the calendar when the parties settled.


So the settlement resolves the acquisition claim. The training claim was already decided — in the defendant's favor — and is not what the class was paid for. The remedy reflects that: alongside the money, Anthropic must destroy the files pulled from the pirate libraries.


## What Actually Makes This AI Copyright Settlement Notable


Set aside the headline figure and three features deserve attention.


• **The claim rate is extraordinary.** More than 91% of covered works — 440,490 of them — had been claimed as of April 16. Consumer class actions routinely see participation in the low single digits. A 91% rate tells you this was a well-defined, identifiable, professionally represented class working from registration and ISBN records rather than a diffuse group of consumers who never opened the notice. Courts notice this, and it is part of why the settlement cleared the fairness standard.


• **The injunctive component may outlast the money.** The destruction requirement establishes that an unlawfully assembled training corpus is not a sunk cost a defendant simply pays down and keeps. For any company holding scraped data of uncertain provenance, that is the more durable exposure.


• **The per-work figure is now a reference point.** Approximately $3,000 per book, before fees and costs, is not a legal standard. But it is the first large-scale number in this space, and both plaintiffs' counsel and licensing negotiators will cite it. Expect it to anchor demands and term sheets that have nothing to do with these facts.


## What This Case Does Not Decide About Fair Use and AI Training


A district court approval order is not a national rule, and clients should resist the temptation to read this settlement as the answer to the AI copyright question.


The fair use holding here was one judge, on one record, about one category of work and one method of use. Other courts have reached different conclusions on different records, and the pending litigation against other model developers involves distinct theories — output-side substitution, market dilution, DMCA claims over stripped copyright management information — that were not the center of gravity in *Bartz*. A settlement, by design, also forecloses the appellate review that would have turned this into precedent. The law here remains genuinely unsettled, and anyone selling you certainty about it is selling you something.


What *Bartz* does establish, durably, is the framing: courts are separating **acquisition** from **use**, and the acquisition side is where the money has been.


## Practical Steps for AI Developers and Rights Holders


**If you build or deploy AI systems:**


1. **Treat data provenance as a litigation record, not an engineering detail.** The question that cost $1.5 billion was not "what did the model learn?" It was "where did this file come from?" Maintain acquisition records — purchase orders, license agreements, scan logs — with the same discipline you apply to invention disclosures.


2. **Audit for retained corpora.** Anthropic's exposure ran in part to material it kept, not merely material it trained on. Inventory what your organization is holding, and adopt a retention and deletion policy for datasets whose provenance you cannot document.


3. **Understand that lawful acquisition is available and defensible.** Buying and scanning books fared differently in this case than downloading them. Where a compliant path exists, the cost of taking it is almost always lower than the cost of the alternative.


4. **Do your diligence on acquired data and acquired companies.** Training data liability travels with the asset. Provenance representations and indemnities belong in your data-supply and M&A documents.


**If you own copyrights:**


1. **Register, and register on time.** Statutory damages and attorney's fees under the Copyright Act generally require timely registration, and registration records are what made a class of nearly half a million works administrable in the first place. Unregistered works are considerably harder to include in a proceeding like this one.


2. **Do not ignore class notices.** The authors in the 9% who did not claim have works on the list and money on the table. Establish an internal process for reviewing IP-related class notices rather than routing them to the same place as junk mail.


3. **Inventory your catalog for training value.** Know which of your works are attractive as training material and decide your licensing posture deliberately — before someone else decides it for you.


4. **Distinguish your claim.** If you are evaluating an infringement theory, understand which side of the acquisition/use line your facts fall on. That is where the leverage currently is.


## The Line the Court Drew


There is nothing novel about the principle underneath this case. A library that buys a book and a library that steals one may put it to identical use, and the law has never treated them identically. What is novel is the scale at which the question is now being asked, and the fact that the answer arrived in a class action rather than a single infringement suit.


The courts will spend years working out what fair use means for machine learning. In the meantime, they have said something simpler and more immediately actionable: however that question resolves, you still have to be able to say where you got it.


*Related reading: [AI's Hidden Ownership Question: What a New "Worldview" Benchmark and the USPTO's Revised Inventorship Guidance Mean for Innovators](https://www.woodphillips.com/post/uspto-ai-inventorship-guidance) and [Look What You Made Me Trademark: Taylor Swift's New Layer of Defense Against Generative AI](https://www.woodphillips.com/post/look-what-you-made-me-trademark-taylor-swift-s-new-layer-of-defense-against-generative-ai).*


---


*Jennifer Gordon is a partner at Wood Phillips, where she counsels clients on intellectual property, technology, and data privacy. For guidance on AI training-data provenance, copyright registration and enforcement strategy, or data-licensing agreements, contact the author or your Wood Phillips attorney.*


*This post is provided for general informational purposes only, does not constitute legal advice, and does not create an attorney-client relationship. It reflects developments as of July 2026. The district court rulings discussed are not binding precedent outside that court and may be affected by subsequent decisions in related litigation.*


*Sources: Order granting final approval, Bartz v. Anthropic PBC, N.D. Cal. (July 20, 2026); Associated Press wire coverage (July 21, 2026); prior orders of Judge William Alsup on fair use and preliminary settlement approval (2025).*

 
 
 

Comments


bottom of page