Artificial intelligence (AI) models are turning to printed books as a data source for their training. The company ISBNdb, which holds an extensive literary database, provides this content for the development of AI models.
These systems require high-quality data, and there is a preference for physical books published before 2022, as they are considered 'cleaner' material, meaning free from superficial content generated by other AIs.
The massive acquisition of these copies raises concerns, given the risk that the books may be discarded or destroyed after being processed by machines. There is also a lack of clarity regarding these transactions, especially concerning rare works.
Currently, AI models are obtaining information from an unusual source: printed books. ISBNdb acts as a digital repository for these works, capturing material from various suppliers, both small and large. Although it is crucial for AI training, there remains a fear that these books are being consumed or, worse, destroyed after automated reading.
According to ISBNdb, this methodology represents the ideal way to train Large Language Models (LLMs), as the material is considered 'clean,' avoiding what is called AI Slop, which is shallow content produced by generative artificial intelligences. Using older data resolves the potential problem of declining information quality when using more recent materials.
As reported by the website 404 media, AI developers prefer publications prior to 2022, the year models like ChatGPT and Claude gained prominence in the market. It is worth noting that the main sources of information for LLMs include platforms such as Reddit and Wikipedia, which contain user-created texts.
The purpose of this search for physical books is to elevate the quality of the data used in AI models, serving as a strategy to combat AI Slop, a practice that has been targeted online. Recently, LinkedIn implemented a feature to report this type of content. Superficial texts, generated entirely or partially by AIs, pose an obstacle for social media users as well as corporations, a phenomenon known as Workslop.
For LLMs to generate more natural dialogues, a solid base of real interactions is necessary, which justifies the importance of training with social media data like Reddit. Similarly, Wikipedia offers a vast collection of compiled information, facilitating access to a large volume of simple writing.
The main hurdle lies in obtaining these printed books on a large scale. To make this possible, ISBNdb makes data from bookstores, small merchants, and even individually sold titles available in its catalog. The company claims it is possible to organize purchases amounting to one million units.
The increase in the purchase of physical books occurs during a period of massive content digitalization, which draws attention. Marketplaces observe this movement cautiously, as supply does not match such high demand. Some sellers point to the lack of transparency in these deals, as there is no clear information about who is acquiring these large lots or why. ISBNdb, for its part, states that negotiations are kept confidential.
A positive aspect for sellers, besides the financial benefit, is the release of titles that were stagnant on shelves, according to an interviewee to 404 Media. However, there remains the fear that the works will be destroyed after scanning. This is suggested by a lawsuit filed against Anthropic in the United States, where the digitization method requires cutting the pages for mechanical reading, which would speed up the process. Thus, precious materials, such as rare and low-circulation books, run the risk of becoming mere 'food' for AI.