Tech
EN AZ
Is it legal to train AI models on copyrighted books? It’s complicated

Is it legal to train AI models on copyrighted books? It’s complicated

techcrunch.com 23.08.2026 17:00 12 views
Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right?

You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right?

At face value, this seemed like a moral victory favoring authors, but Judge Alsup actually ruled that Anthropic’s AI training was lawful. What Alsup penalized Anthropic for was pirating these books from illegal online shadow libraries. Gellis thinks the ruling is more advantageous for AI companies.

What’s a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028? Fair use is a carve out of copyright law that allows for the use of copyrighted materials without explicit permission, protecting the ability to comment and iterate on copyrighted works through criticism, parody, education, and other means. Judges consider specific factors when deciding if something is fair use, including the purpose and nature of the work, the amount used, and its impact on the market.

What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.” Henderson is referencing a case in which the media and technology company Thomson sued the research firm Ross Intelligence for copying its content in order to build a competing, AI-based legal platform. In that case, Judge Bibas decided that it was not fair use to train on ’ content to make a new platform that would directly compete with it. While authors could potentially argue that chatbots are competing with them by using their works to generate new, synthetic books, that argument has not yet prevailed in court.

When it comes to the relationship between AI and copyright, Gellis finds it helpful to narrow down what we’re actually talking about – the way we think about copyright in terms of AI training is quite different from how we think about copyrighting AI-generated content. Perlmutter, the court ruled that if a work is 100% AI-generated, it’s not copyrightable, which opens a whole new can of worms – how can we definitively prove whether or not a work was generated using AI, and if so, how do we know what percentage of it was created or assisted with AI? It would be kind of foolish for the AI companies to ignore them.”

Extract — continue reading at the source.

Read full story