Skip to main content

AI Training on Copyrighted Books: A Legal Tightrope?

The artificial intelligence models underpinning popular chatbots such as ChatGPT, Gemini, and Claude are extensively trained on vast datasets of publi

5 min read12 views5 tags
Originally reported bytechcrunch

The artificial intelligence models underpinning popular chatbots such as ChatGPT, Gemini, and Claude are extensively trained on vast datasets of published works, encompassing millions of books, online articles, academic papers, and a wide array of internet content. This process often involves the use of material from published authors without their explicit knowledge or consent, leading to concerns that these very AI tools could undermine their livelihoods. The immediate question arises: is this practice illegal?

However, the legal landscape is far from straightforward.

Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, highlighted the complexity to TechCrunch, stating, “I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on. It’s very complex and there are a lot of raw feelings about what is happening, both for and against.”

A notable ruling occurred last year when Judge William Alsup ordered Anthropic to pay a substantial $1.5 billion copyright settlement to a group of writers whose works were utilized for training its AI models. While this initially appeared to be a significant victory for authors, Judge Alsup's decision actually affirmed the lawfulness of Anthropic's AI training process itself. The penalty was levied against Anthropic for illegally acquiring these books from unauthorized online "shadow libraries."

The judge articulated his reasoning, writing, “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” drawing a parallel between an LLM's ingestion of vast textual data and a human writer's study of literature.

Gellis perceives this ruling as largely beneficial for AI companies, questioning the true impact of a $1.5 billion fine on a company projected to reach approximately $200 billion in annual revenue by 2028.

She elaborated, “I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work. Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work.”

A significant challenge stems from the fact that copyright law has not been updated since 1976. This places judges in the position of interpreting decades-old guidelines to address novel legal questions posed by AI, decisions that could profoundly shape the future of the entire industry.

Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch that there is widespread concern because "the law is all over the place, and it’s because of this question. They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.”

These legal inquiries frequently revolve around fair use doctrine, specifically whether the utilization of copyrighted material is sufficiently "transformative" to be legally permissible.

Fair use is a critical exception within copyright law, permitting the use of copyrighted works without explicit permission for purposes such as criticism, parody, education, and other forms of commentary or iteration. Judges assess several factors when determining fair use, including the purpose and character of the use, the amount of the work used, and its effect on the potential market for or value of the copyrighted work.

Henderson emphasized that “Copyright is always about protecting and growing the market.” He noted the inconsistency in judicial reasoning in AI cases, observing a trend: “What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.”

Henderson's observation refers to a case where Thomson Reuters, a media and technology company, sued the research firm Ross Intelligence. Ross was accused of copying Reuters' content to develop a competing AI-based legal platform.

In that specific case, Judge Stephanos Bibas ruled last year, stating, “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s.”

Judge Bibas concluded that training an AI on Reuters' content to create a directly competing platform did not qualify as fair use. While authors could argue that chatbots compete by using their works to generate new, synthetic books, this specific argument has yet to achieve success in court.

When discussing AI and copyright, Gellis finds it beneficial to distinguish between two separate issues: the legal considerations surrounding AI training data and the copyright status of AI-generated content itself.

In a case known as Thaler v. Perlmutter, the court determined that a work generated entirely by AI is not copyrightable. This ruling opens up a new set of complex questions, such as how to definitively prove the extent or percentage of AI involvement in a creative work.

Gellis used an analogy to illustrate the challenge: “If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel.” She added that “[AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.”

Currently, most AI companies are embroiled in pending litigation concerning these issues, suggesting that definitive resolutions to these multifaceted problems are not imminent.

Gellis concluded, “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail. But in the meantime, all these decisions are shaping everything that’s happening. It would be kind of foolish for the AI companies to ignore them.”

#AI News#AI Training#Copyright Law#Anthropic Case#Legal Tightrope
ES
Editorial StaffEditor

The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.

View all posts
Reader feedback

What did you think of this story?

User Comments

Filter:
No comments yet. Be the first to comment!
Continue reading
View all news