
Book Pirates or Training Ground: The Legal Battle Over Copyrighted Data
Major tech platforms like OpenAI, Google, and Anthropic train their models on vast libraries of published books, online articles, and academic papers without asking for permission or paying authors. While authors argue that using their work without consent destroys their livelihoods, federal courts view the legal reality much differently.
Intellectual property attorney Cathy Gellis explains that copyright law in this sector remains extremely complex. Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to authors whose books trained its models. However, Judge Alsup also ruled that using books to train software models is legally permissible under existing fair use standards. He compared a model studying published text to a human student reading books to learn how to write, noting that software aims to create new works rather than copy existing stories line by line.
Gellis points out that this legal interpretation favors tech firms, especially when a $1.5 billion court settlement represents a small expense for a sector projecting $200 billion in annual revenue by 2030. Copyright law focuses heavily on whether software duplicates existing text, rather than how systems ingest and read data during training stages.
Current copyright statutes date back to 1976, forcing federal judges to interpret fifty-year-old guidelines for modern technology. Attorney Jason Henderson notes that uncertainty around legal rules creates widespread concern across creative industries. Courts decide these disputes by evaluating fair use guidelines, specifically checking whether new software creates a distinct purpose or competes directly with the original creator’s market.
For example, Thomson Reuters sued research firm Ross Intelligence for copying legal database records to build a competing research platform. Judge Stephanos Bibas ruled that copying database records to build a direct rival product was unfair because it directly stole market share. While authors argue that chatbots compete with writers by generating synthetic books, that argument has not won in court yet.
Courts also separate training rules from generated outputs. In the Thaler v. Perlmutter case, federal judges confirmed that purely machine-generated content cannot hold a copyright. Gellis notes that using basic software tools like spellcheck never stripped authors of their ownership rights, but modern software forces society to re-examine where human creation ends and machine generation begins.
Most technology firms remain locked in ongoing court battles over data scraping, meaning federal courts will not settle these questions anytime soon. Early trial rulings shape how tech firms operate today, but future appeals could shift rules entirely. For now, tech companies continue training models while watching federal court dockets closely.







