Copyright Debate: Use of Books for Artificial Intelligence Training
Read more
Olhar Digital
olhardigital.com.br

Copyright Debate: Use of Books for Artificial Intelligence Training

Artificial intelligence models, which power tools like ChatGPT, Gemini, and Claude, were developed using vast databases containing hundreds of millions of books, articles, academic papers, and internet content. A central point of contention is that most of the authors whose works comprise these datasets were never consulted or gave their consent for such use.

The legal complexity of this issue was addressed by Cathy Gellis, a lawyer specializing in intellectual property, copyright, and technology, in an interview with TechCrunch. She observed that the legal and technological fields are undergoing intense changes, generating strong opinions both for and against the practice.

In a pioneering case, Judge William Alsup ordered Anthropic to pay a collective of writers whose works were used in training the company's AI models $1.5 billion. However, Alsup concluded that the act of training itself was legal, penalizing only the illegal acquisition of books from unauthorized digital libraries.

In this decision, the judge drew an analogy between the process of ingesting trillions of words by a language model and the reading done by a student studying pre-existing works to create something new. Alsup stated that Anthropic's LLMs trained on works not to replicate or replace them, but rather to overcome conceptual challenges and generate distinct results.

Gellis assesses that the verdict is more beneficial to AI corporations than to the authors themselves. For a company projecting annual revenues close to $200 billion by 2028, a $1.5 billion fine carries a different financial weight. She commented that it is positive for AI development that the court recognized the similarity to reading copyrighted material, instead of treating it as mere copying.

According to Gellis, American copyright legislation has not been revised since 1976, forcing judges to interpret guidelines from fifty years ago to resolve dilemmas that may define the future of the AI industry. Much of the litigation focuses on the concept of fair use, which defines whether the use of a protected work is sufficiently 'transformative' to be legally permitted. This is a copyright exception that allows use without explicit permission for purposes such as criticism, parody, and education, analyzing factors such as the purpose of the use, the quantity used, and the impact on the original market.

Jason Henderson, senior attorney and founder of the IP and Media Practice at JWL International, told TechCrunch that copyright always aims to protect and expand the market. He added that courts have sanctioned cases where the training of a work aims to directly compete with it, but tend to accept uses that do not create this direct competition.

One example of this was the lawsuit filed by Thomson Reuters against Ross Intelligence, which had copied content to develop an AI-based legal platform with potential competition. Judge Stephanos Bibas considered that the use was not transformative because it lacked a 'additional purpose or different character' compared to the Thomson Reuters content.

Although authors may argue that chatbots compete with them by using their works to produce new synthetic books, this argument has not yet succeeded in court. Gellis emphasizes that the relationship between AI and copyright covers two distinct topics: the use of works in model training and the safeguarding of AI-generated content. In another US ruling, the court decided that a creation entirely generated by AI cannot be protected by copyright, introducing new uncertainty about how to prove the origin of the work and what percentage of it derives from human authorship or machine assistance.

As AI forces a reevaluation of previously neglected concepts, Gellis concludes that most AI companies still face lawsuits. She warns that initial decisions are influential but can be overturned by other courts, indicating that the legal landscape will continue to evolve.

Similar stories

Amazon acquires and mass destroys books in Las Vegas for Artificial Intelligence training
Read more
tecnoblog.net

Amazon acquires and mass destroys books in Las Vegas for Artificial Intelligence training

A recent surge in the purchase of physical books by artificial intelligence companies has gained attention this week. An investigation by the portal 404 Media tracked a rare copy that was sent in July within a batch of a thousand books, discovering that the package arrived at an Amazon logistics warehouse in Las Vegas, United States.

This facility is reportedly used by Amazon specifically to digitize and subsequently destroy millions of books, with the goal of extracting human-created texts to feed AI models. This method allows corporations to obtain large quantities of high-quality data, especially given the increasing difficulty in finding reliable content for training on the internet.

The experiment was motivated by an unusual growth in large batch sales in the publishing market. Suspecting this sectoral boom, the website's team sent an Apple AirTag device hidden inside a book in an order of a thousand units, placed through Biblio, an important online marketplace for independent publications.

The tracker's journey culminated in the hands of VGT3, a team operating within an Amazon logistics complex. According to local worker reports, this unit's exclusive function is to transform physical books into digital datasets. To speed up the insertion of pages into high-performance industrial scanners, employees proceed to cut all spines and bindings, resulting in the damage of the original material immediately after digitization.

The team's visual identity does not attempt to disguise the procedure; according to 404 Media, the VGT3 logo features a dinosaur holding a book, symbolizing the act of tearing and consuming the material.

Publications edited and released before the peak of AI in 2022 are considered a source of purely human, clean, and guaranteed language. Using current web content to train next-generation algorithms carries the risk of incorporating machine-generated texts, which, in the long run, can diminish the logical capacity of these models.

Beyond data quality, financial and legal factors drive this practice. Previously, large technology companies obtained digital files for free through clandestine internet libraries. However, judicial scrutiny restricted this activity. Court documents indicated that Anthropic managed an internal initiative to digitize all world literature. This endeavor encountered copyright obstacles, culminating in an agreement where the owner of Claude would pay $1.5 billion to maintain a database containing seven million pirated books.

In parallel, Meta was also subject to lawsuits following the leak of an illegal collection containing nearly 82 terabytes of literary works. Due to billions of dollars in fines imposed for copyright infringement, the legal acquisition of vast physical collections has become a more economical and legally safer option for corporations.

A recent ruling by the American federal court suggested that using legitimately purchased copies to train AI models could be classified as 'fair use,' creating legal openings for the continuation of this practice. When contacted about the Las Vegas facility, Amazon merely confirmed in a statement that it acquires publications through legitimate commercial means to 'help develop and improve the products and services that customers use.'

AI-Generated Books Flood Amazon and Other Bookstores
Read more
olhardigital.com.br

AI-Generated Books Flood Amazon and Other Bookstores

Users visiting online e-book stores such as Amazon are encountering a large number of works created or written using artificial intelligence (AI). These works often feature standard covers and template plots. Today, authors use AI-based systems that promise a ready-made e-book in less than 30 minutes, allowing them to release hundreds of titles annually under numerous pseudonyms.

This influx of new publications is impacting the publishing industry. Readers complain about the low quality of 'robotic content,' while human authors are losing visibility in search engines and seeing a decline in their income. If the advent of the e-book previously simplified distribution, AI has now significantly lowered production costs, but the market has become oversaturated. The question arises as to how to find clarity amidst so much noise.

Maurício Pinheiro, a professor of technology and arts from Sesc Piracicaba, explains to the publication Olhar Digital that AI has changed the scale of independent authorship. Now, one person can conduct research, structure, write, edit, and format a work.

This allows the author to operate almost like a micro-publishing house, testing various covers, descriptions, and niches at a speed impossible under the previous model.

This simplicity has transformed the logic of the digital market. Previously, producers aimed to create bestsellers, but the focus has now shifted to overall sales volume. Since the cost of text generation has dropped sharply, publishing dozens of titles consecutively has become a profitable business, even if each book sells little or the income depends on pay-per-read in subscription services.

Industry data confirms the impact of this industrial production tactic. For instance, a study by the National Bureau of Economic Research (NBER) shows that the number of e-books released on Amazon monthly has nearly tripled between 2022 and the end of 2025.

Another study conducted by the State University of New York found that the number of books with registered sales increased 19.2 times during this period, while total market revenue only grew by 8.9 times. This means that far more works are competing for readers' attention, leading to a decrease in average income per book.

According to the professor, who is also a software analyst, the reduction in production barriers has eliminated traditional filters that screened books before they reached the public. The specialist notes: 'For a long time, publishing was complex, and this complexity served as a filter before the book reached the audience. Now the filter happens later.'

The virtual stores themselves are now struggling to control this flow of synthetic content. For example, on the Rakuten Kobo self-publishing platform, 45% of uploaded e-books were rejected in one year, with 85% of these rejections being due to the books being AI-generated. Despite blocks and analysis, automated systems continue to bypass detection and add titles to stores daily.

Beyond market consequences, mass production of AI-generated e-books faces a legal hurdle in Brazil. Lawyer Marcelo Gargano, a digital law specialist at ABE Advogados, explains to Olhar Digital that Brazilian copyright law focuses on a natural person and requires human creativity to protect a work.

Gargano asserts: 'An e-book entirely generated by AI without significant human creative input does not have a natural person who could be recognized as the author of the creation, and without an author, there is no copyright protection.'

This implies that these books can be freely copied by third parties and also carry other risks for those who publish them, as there is a danger that the tool might reproduce copyrighted excerpts from other authors.

Furthermore, there is a discussion about the reader's right to know the origin of the purchased product. However, the country does not yet require platforms to indicate the use of technology on covers. The lawyer points out: 'It can be assumed that the consumer has the right to adequate and clear information about various products and services in accordance with the CDC, and therefore, the mentioned authors and platforms that publish them should inform consumers that this work was created using AI.'

Amazon serves as an example. On this marketplace, the books Tumbleweeds & Tequila: A Forced Proximity Romance (Shots of Love Book 2) by Sydney Marsh and Fighting Love’s Flame (Alaska Rugged Heart Series Book 3) by Dana Winston carry the following notice: 'This story was produced using AI tools, directed by the author' (free translation). However, it should be noted that this warning is hidden at the bottom of the description, behind the 'read more' button.

This legal vulnerability is accompanied by a change in the very value of the book, beyond its monetary equivalent. Since algorithms are capable of generating stories in large volumes, the ability to create acceptable text is no longer a distinguishing feature. AI is very competent at creating what seems plausible, familiar, and good enough. But there is a curious consequence: in an ocean of 'good enough' things, 'good enough' ceases to be a rare quality.

In this scenario of abundance, the value of the independent market shifts towards what cannot be automated. Factors such as personal style, repertoire, life experience, and the certainty that there is a real author behind the pages gain more importance for the reader's choice. The Sesc Piracicaba professor analyzes: 'If the text becomes practically infinite, then what cannot be produced infinitely begins to be valued higher: trust, reputation, perspective, repertoire, lived experience, taste, and recognizable relationship between the work and someone who is willing to take responsibility for it.'

Ultimately, self-publishing has turned into a battle against being overshadowed in an infinite catalog. Pinheiro concludes: 'I do not think the main fight will simply be between the human author and the artificial author. The fight will be between signal and noise.'

For the reader, the task of choosing the next book to immerse oneself in becomes more difficult. For the human writer, creating something outstanding now requires a 'secret ingredient' that a machine cannot replicate.

Popular