🚨BREAKING: A LANDMARK JUDGEMENT FOR THE AI INDUSTRY. US Federal...

US Federal Judge ruled Anthropic may train its AI on published books without authors’ permission.
This is the first court endorsement of fair use protecting AI firms when they use copyrighted texts to train LLMs.
AI may study what it buys, not what it grabs from pirate sites.
---------
"First, Authors argue that using works to train Claude’s underlying LLMs was like using works to train any person to read and write, so Authors should be able to exclude Anthropic
from this use (Opp. 16). But Authors cannot rightly exclude anyone from using their works for training or learning as such. Everyone reads texts, too, then writes new texts. They may need
to pay for getting their hands on a text in the first instance. But to make anyone pay specifically for the use of a book each time they read it, each time they recall it from memory,
each time they later draw upon it when writing new things in new ways would be unthinkable.
For centuries, we have read and re-read books. We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing
problems."
The court file is such an interesting read.
🧵 Read on 👇
Using complete books to map token relationships is “spectacularly transformative.” No verbatim outputs reach users, and the system’s purpose—generating fresh text—is orthogonal to selling the originals.
That satisfies factor 1 and, with no market substitution, factor 4 as well.
He explains that a closer match is a past case where an AI tool was trained on court opinions so it could draft fresh legal text, and that earlier court called the training fair use.
He adds that the AI’s new writing stands far enough from the source material that copyright holders cannot claim control over it.
Citing a Supreme Court precedent and the Constitution, he stresses that this kind of reuse helps knowledge grow without hurting authors’ incentive to create.
He ends by calling Anthropic’s training “quintessentially transformative,” because the model learns from books only to produce fresh, non-substituting text.
Because Anthropic paid for each print copy, it can dispose of that copy as it likes, including destroying it after making one digital version.
The format shift adds no extra copies, only makes storage easier and searching possible, so the judge calls the change transformative and therefore fair use.
The authors cannot show that Anthropic shared these scans outside the company, so the court sees no harm to the market.
The judge chooses a simpler route and says the mere switch from print to digital is fair use, because it only changes how the text is stored and searched, not the expressive content itself.
He cites earlier rulings on microfilming journals, scanning books for search, and recording broadcasts for later viewing, all of which treated private format shifts as transformative when they do not hurt sales.
The outcome is that Anthropic may keep one digital replacement for every print book it bought, and that step does not violate copyright.
Grabbing pirate-site files when legal copies were for sale “plainly displaced demand” and served nothing more than convenience.
Keeping them after deciding they would never train models showed the library was an independent, non-transformative use.
Every fair-use factor cuts against Anthropic here.
The judge rejects the authors’ claim that model training destroys their sales; that loss is no different from new writers entering the field.
Anthropic wins the right to train on paid books, yet still faces a jury over its pirate cache.
Anthropic must return to court for a jury trial that will look only at the 7 million pirate books.
Judge Alsup has already said the focus will be actual or statutory damages, including willfulness penalties that jump when bad faith is proven.
💰 How the numbers climb
Copyright law sets $750 as the rock-bottom fee for every infringed book, and it reaches $150 000 when willfulness appears.
Multiplying the floor figure by 7 million puts the minimum exposure north of $5 billion, even before any punitive multiplier.
📚 Licensing versus scraping
Anthropic already spends millions buying used books to scan, showing there is a real market for lawful training copies.
If licensing deals emerge, they could offset courtroom risk and become the industry’s new normal.
→ 🗞️ rohan-paul.com
Includes:
- Top 1% AI Industry developments
- Influential research papers/Github/AI Models/Tutorial with analysis
📚 Subscribe and get a 1300+page Python book instantly.









