
The $1.5 Billion Reckoning: Anthropic Pays Up for Pirated Book Datasets
Anthropic can finally start writing checks to a massive group of authors and book publishers. A federal judge gave final approval to the AI lab landmark $1.5 billion copyright settlement, closing out a major class action battle.
Judge William Alsup of the US District Court for the Northern District of California issued preliminary approval last year after ruling that Anthropic illegally downloaded and stored millions of copyrighted books. Following Alsup retirement, Judge Araceli Martinez-Olguin signed off on the final settlement terms on Monday.
Under the agreed settlement structure, rights holders receive $3,000 per work across roughly 500,000 books. Authors and publishers who hold rights share those payments directly. While legal analysts view this payout as the largest in US copyright law history, many writers and creators refuse to see it as a true victory.
Their frustration comes down to how the court handled the core legal argument. Judge Alsup sided with Anthropic on the main AI training question. He ruled that using copyrighted text to train an AI model counts as fair use, setting a precedent that AI companies desperately wanted. However, the judge drew a sharp line at how Anthropic obtained those books in the first place.
Anthropic built its original training library using two distinct channels. It legitimately purchased and scanned physical books, which the court found completely legal under fair use guidelines. But it also downloaded massive torrent files from pirate sites like Library Genesis and Pirate Library Mirror. Judge Alsup ruled that downloading pirated copies was illegal on its own terms, sending that specific issue toward a jury trial. Rather than risking unpredictable jury damages over piracy charges, Anthropic chose to settle.
Because Anthropic settled before an appeals court weighed in, this case does not create binding legal precedent across the entire country. Other federal judges remain free to make their own decisions in similar cases involving OpenAI, Meta, Google, and Midjourney. In fact, a group of publishers and authors including Hachette, Cengage, and Elsevier recently launched a fresh class action lawsuit against Google, accusing the search giant of scraping copyrighted books to train its Gemini AI model.
Anthropic settlement puts cash directly into the hands of affected creators, but the broader war between artificial intelligence builders and copyright holders is far from over. Silicon Valley companies will keep training models on public text, while authors continue fighting in court rooms across the nation to protect their intellectual property.







