Artificial intelligence models, which power tools like ChatGPT, Gemini, and Claude, were developed using vast databases containing hundreds of millions of books, articles, academic papers, and internet content. A central point of contention is that most of the authors whose works comprise these datasets were never consulted or gave their consent for such use.
The legal complexity of this issue was addressed by Cathy Gellis, a lawyer specializing in intellectual property, copyright, and technology, in an interview with TechCrunch. She observed that the legal and technological fields are undergoing intense changes, generating strong opinions both for and against the practice.
In a pioneering case, Judge William Alsup ordered Anthropic to pay a collective of writers whose works were used in training the company's AI models $1.5 billion. However, Alsup concluded that the act of training itself was legal, penalizing only the illegal acquisition of books from unauthorized digital libraries.
In this decision, the judge drew an analogy between the process of ingesting trillions of words by a language model and the reading done by a student studying pre-existing works to create something new. Alsup stated that Anthropic's LLMs trained on works not to replicate or replace them, but rather to overcome conceptual challenges and generate distinct results.
Gellis assesses that the verdict is more beneficial to AI corporations than to the authors themselves. For a company projecting annual revenues close to $200 billion by 2028, a $1.5 billion fine carries a different financial weight. She commented that it is positive for AI development that the court recognized the similarity to reading copyrighted material, instead of treating it as mere copying.
According to Gellis, American copyright legislation has not been revised since 1976, forcing judges to interpret guidelines from fifty years ago to resolve dilemmas that may define the future of the AI industry. Much of the litigation focuses on the concept of fair use, which defines whether the use of a protected work is sufficiently 'transformative' to be legally permitted. This is a copyright exception that allows use without explicit permission for purposes such as criticism, parody, and education, analyzing factors such as the purpose of the use, the quantity used, and the impact on the original market.
Jason Henderson, senior attorney and founder of the IP and Media Practice at JWL International, told TechCrunch that copyright always aims to protect and expand the market. He added that courts have sanctioned cases where the training of a work aims to directly compete with it, but tend to accept uses that do not create this direct competition.
One example of this was the lawsuit filed by Thomson Reuters against Ross Intelligence, which had copied content to develop an AI-based legal platform with potential competition. Judge Stephanos Bibas considered that the use was not transformative because it lacked a 'additional purpose or different character' compared to the Thomson Reuters content.
Although authors may argue that chatbots compete with them by using their works to produce new synthetic books, this argument has not yet succeeded in court. Gellis emphasizes that the relationship between AI and copyright covers two distinct topics: the use of works in model training and the safeguarding of AI-generated content. In another US ruling, the court decided that a creation entirely generated by AI cannot be protected by copyright, introducing new uncertainty about how to prove the origin of the work and what percentage of it derives from human authorship or machine assistance.
As AI forces a reevaluation of previously neglected concepts, Gellis concludes that most AI companies still face lawsuits. She warns that initial decisions are influential but can be overturned by other courts, indicating that the legal landscape will continue to evolve.

