A recent surge in the purchase of physical books by artificial intelligence companies has gained attention this week. An investigation by the portal 404 Media tracked a rare copy that was sent in July within a batch of a thousand books, discovering that the package arrived at an Amazon logistics warehouse in Las Vegas, United States.
This facility is reportedly used by Amazon specifically to digitize and subsequently destroy millions of books, with the goal of extracting human-created texts to feed AI models. This method allows corporations to obtain large quantities of high-quality data, especially given the increasing difficulty in finding reliable content for training on the internet.
The experiment was motivated by an unusual growth in large batch sales in the publishing market. Suspecting this sectoral boom, the website's team sent an Apple AirTag device hidden inside a book in an order of a thousand units, placed through Biblio, an important online marketplace for independent publications.
The tracker's journey culminated in the hands of VGT3, a team operating within an Amazon logistics complex. According to local worker reports, this unit's exclusive function is to transform physical books into digital datasets. To speed up the insertion of pages into high-performance industrial scanners, employees proceed to cut all spines and bindings, resulting in the damage of the original material immediately after digitization.
The team's visual identity does not attempt to disguise the procedure; according to 404 Media, the VGT3 logo features a dinosaur holding a book, symbolizing the act of tearing and consuming the material.
Publications edited and released before the peak of AI in 2022 are considered a source of purely human, clean, and guaranteed language. Using current web content to train next-generation algorithms carries the risk of incorporating machine-generated texts, which, in the long run, can diminish the logical capacity of these models.
Beyond data quality, financial and legal factors drive this practice. Previously, large technology companies obtained digital files for free through clandestine internet libraries. However, judicial scrutiny restricted this activity. Court documents indicated that Anthropic managed an internal initiative to digitize all world literature. This endeavor encountered copyright obstacles, culminating in an agreement where the owner of Claude would pay $1.5 billion to maintain a database containing seven million pirated books.
In parallel, Meta was also subject to lawsuits following the leak of an illegal collection containing nearly 82 terabytes of literary works. Due to billions of dollars in fines imposed for copyright infringement, the legal acquisition of vast physical collections has become a more economical and legally safer option for corporations.
A recent ruling by the American federal court suggested that using legitimately purchased copies to train AI models could be classified as 'fair use,' creating legal openings for the continuation of this practice. When contacted about the Las Vegas facility, Amazon merely confirmed in a statement that it acquires publications through legitimate commercial means to 'help develop and improve the products and services that customers use.'

