Technology· AI Tools

Is it legal to train AI models on copyrighted books? It’s complicated

You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods.

By AI NewsroomPublished 2 days agoUpdated about 1 hour ago3 views
Is it legal to train AI models on copyrighted books? It’s complicated

Why It Matters

This story touches on their, legal, train — topics readers are actively tracking. Review and add editorial context before publishing.

Key Facts

  • Fact 1: “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.” Last year, in one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay a mammoth $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models.
  • Fact 2: What’s a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028?
  • Fact 3: “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work.” Copyright law hasn’t been updated since 1976, which means that judges have to figure out how to interpret guidelines from 50 years ago when confronting legal questions that have the potential to shape the future of the AI industry.
  • Fact 4: Perlmutter, the court ruled that if a work is 100% AI-generated, it’s not copyrightable, which opens a whole new can of worms – how can we definitively prove whether or not a work was generated using AI, and if so, how do we know what percentage of it was created or assisted with AI?

You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods.

That seems illegal, right? The reality isn’t that simple. “I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch.

(Original synthesis pending human/AI review — generated by the stub provider by selecting real sentences from the source material, not by writing new analysis or commentary.)

Original source: TechCrunch

Share

Related Stories