VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Hidden AirTag Reveals Amazon Shredding Books for AI
Value Add VC/Pulse/AIDEEP DIVE

Hidden AirTag Reveals Amazon Shredding Books for AI

A 404 Media investigation used a hidden AirTag to trace a shipment of rare books to an Amazon warehouse where workers scan and shred the physical copies to train Amazon's Nova AI models.

By the Numbers

1,000
Books in tracked shipment
Amazon LAS8, Las Vegas
Warehouse
VGT3
Internal unit
Amazon Nova
Target model family
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
August 17, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

404 Media hid an Apple AirTag inside a 1,000-book bulk order placed through marketplace site Biblio and tracked it from California through Wisconsin and Colorado to an Amazon complex in Las Vegas called LAS8

2

Workers at a unit inside the warehouse, labeled VGT3, described slicing bindings off books and scanning the pages before destroying the physical copies -- Amazon uses the scanned text to train its Nova AI models

3

Printed books that predate 2022 are especially valuable training data because they exist outside the internet corpus and are free of AI-generated content contamination

4

The investigation adds Amazon to a growing list of AI companies -- including prior reporting on other labs -- accused of destroying physical media to source clean training data at scale

TC

The VC Read · Trace's Take

Trace Cohen

The legal nuance here matters more than the headline suggests -- Amazon owns the physical books it's destroying, which sidesteps the scraping-copyright fights other labs are fighting, but it doesn't sidestep the reputational exposure of 'AI company shreds rare books' becoming a plaintiff's exhibit in every ongoing training-data lawsuit. If I were diligencing any company using bulk-purchased physical media for training data, I'd want the chain-of-custody and rarity screening process documented before it becomes a headline, not after.

Analysis

An investigation by 404 Media, republished by Ars Technica, placed a hidden Apple AirTag inside a bulk order of rare books purchased through the marketplace site Biblio and tracked its journey across the country -- from California through Wisconsin and Colorado, finally landing at an Amazon fulfillment complex in Las Vegas known as LAS8. Inside, the shipment fed a unit labeled VGT3, marked with a dinosaur-and-book logo, where workers told the outlet their job is to slice the bindings off books and scan the pages before destroying the physical copies. Pulse previously covered Amazon's expanding AI ambitions with its Nova model family, the effort this training-data pipeline feeds.

Amazon uses the scanned text to train its Nova family of AI models. Printed books that predate roughly 2022 carry particular value for this purpose because they exist largely outside the internet's already-scraped text corpus and are, by definition, free of AI-generated content that has increasingly polluted newer web data -- a problem researchers call model collapse when synthetic text gets fed back into training pipelines.

“Amazon uses the scanned text to train its Nova family of AI models.”

The pattern beyond Amazon

The investigation adds Amazon to a pattern of major AI labs sourcing training data from physical archives rather than relying solely on digitized text, reflecting how scarce genuinely clean, pre-AI training data has become as the large labs compete on data quality rather than just data volume. The physical destruction of the source material -- rather than digitizing and preserving it -- is the detail drawing the sharpest reaction, particularly for books that may be out of print or rare enough that no other digital copy exists.

The legal exposure here is murkier than a straightforward copyright case: buying physical books and destroying them after scanning sidesteps some of the licensing disputes that have dogged AI labs training on text scraped from the open web, since Amazon owns the physical copies it destroys. But the optics -- an AI-training operation shredding books, some rare, to build a proprietary model -- land at a moment when publishers, authors and librarians are already engaged in multiple lawsuits against AI companies over training data sourcing, and this story is likely to become exhibit material in that broader fight regardless of its narrower legal footing.

ShareXLinkedInEmail

More on

Amazon →

Reported by 404 Media · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 17, 2026

Anthropic's Annualized Revenue Hits $65B in July

Illustration for: Anthropic's Annualized Revenue Hits $65B in July
AI$65B annualized run rate

Anthropic's Annualized Revenue Hits $65B in July

Anthropic told investors its annualized revenue run rate climbed to $65 billion at the end of July, a sevenfold jump from about $9 billion at the end of 2025, as it prepares for an IPO expected this fall.

AI· Aug 17, 2026

Alibaba Answers Meta With Laptop-Ready Qwen Model

Illustration for: Alibaba Answers Meta With Laptop-Ready Qwen Model
AI

Alibaba Answers Meta With Laptop-Ready Qwen Model

Alibaba released Qwen3.8-27B, a laptop-ready open-weight model it says matches the capability of models ten times its size, sharpening its rivalry with Meta for leadership in open-weight AI.

AI· Aug 17, 2026

OpenAI's Brockman Brushes Off Exec Exodus Worries

Illustration for: OpenAI's Brockman Brushes Off Exec Exodus Worries
AI

OpenAI's Brockman Brushes Off Exec Exodus Worries

OpenAI president Greg Brockman told CNBC the wave of executive departures, including at least a dozen senior leaders in 2026, isn't atypical -- just more visible because OpenAI is under constant public scrutiny.

@Trace_Cohen·t@nyvp.com