Analysis
An investigation by 404 Media, republished by Ars Technica, placed a hidden Apple AirTag inside a bulk order of rare books purchased through the marketplace site Biblio and tracked its journey across the country -- from California through Wisconsin and Colorado, finally landing at an Amazon fulfillment complex in Las Vegas known as LAS8. Inside, the shipment fed a unit labeled VGT3, marked with a dinosaur-and-book logo, where workers told the outlet their job is to slice the bindings off books and scan the pages before destroying the physical copies. Pulse previously covered Amazon's expanding AI ambitions with its Nova model family, the effort this training-data pipeline feeds.
Amazon uses the scanned text to train its Nova family of AI models. Printed books that predate roughly 2022 carry particular value for this purpose because they exist largely outside the internet's already-scraped text corpus and are, by definition, free of AI-generated content that has increasingly polluted newer web data -- a problem researchers call model collapse when synthetic text gets fed back into training pipelines.
“Amazon uses the scanned text to train its Nova family of AI models.”
The pattern beyond Amazon
The investigation adds Amazon to a pattern of major AI labs sourcing training data from physical archives rather than relying solely on digitized text, reflecting how scarce genuinely clean, pre-AI training data has become as the large labs compete on data quality rather than just data volume. The physical destruction of the source material -- rather than digitizing and preserving it -- is the detail drawing the sharpest reaction, particularly for books that may be out of print or rare enough that no other digital copy exists.
The legal exposure here is murkier than a straightforward copyright case: buying physical books and destroying them after scanning sidesteps some of the licensing disputes that have dogged AI labs training on text scraped from the open web, since Amazon owns the physical copies it destroys. But the optics -- an AI-training operation shredding books, some rare, to build a proprietary model -- land at a moment when publishers, authors and librarians are already engaged in multiple lawsuits against AI companies over training data sourcing, and this story is likely to become exhibit material in that broader fight regardless of its narrower legal footing.