What this is
404 Media's investigative reporters ran a clever experiment: we secretly tucked an Apple AirTag into a box of roughly 1,000 used books and tracked its logistics chain. The books came from Biblio (a used-book marketplace), purchased by an unnamed buyer placing large orders with no price sensitivity — characteristics the industry has spent the past year suspecting to be AI companies bulk-scanning inventory for model training.
Tracking results pointed clearly: the box ended up at Amazon's LAS8 facility in Las Vegas, Nevada, in the VGT3 area, whose entrance wears a logo of a dinosaur hugging a book (the irony writes itself). Amazon employees confirmed on online forums that the work in that corner is "destructive scanning" of massive book volumes.
This isn't the first time suspicions surfaced. Simon Willison noted that similar bulk book-buying rumors involving Anthropic were reported back in June 2025. But this is the first time a reporter walked the full chain — from order to scan — with physical evidence.
Industry view
What concerns us most here is that the evidence chain has closed. Previously, outsiders could only hypothesize "these anomalous orders are probably from AI companies"; now logistics records have turned hypothesis into fact. Amazon has, as of writing, issued no public response.
The defense side will argue: scanning used books to train models qualifies as fair use, and the data is "anonymized." But the counterargument is equally sharp — publisher and author associations point out that bulk acquisition, without authorization, combined with "destructive scanning" (physical damage to the original books), sits in a gray zone of copyright law and training-data compliance, and may cross the line.
The bigger risk, in our view, sits on the regulatory side: the EU AI Act has already imposed clear requirements on training-data transparency, and physical evidence like this — if introduced into U.S. copyright litigation — would be a key exhibit. AI companies' strategy of "quietly buying data" is breaking down.
Impact on regular people
For enterprise IT: when procuring AI models or services, "training data source compliance" is about to become a mandatory review item — analogous to the data-privacy compliance reviews of years past.
For individual careers: professionals in publishing, copyright, and content creation may soon face copyright-risk inquiries from clients or employers; the legal industry will see a fresh wave of consultation demand.
For the consumer market: the old books you've sold on secondhand platforms, the out-of-print titles you've bought — any of them may already be scanned into some large model's training set. "My book was read by an AI" has shifted from hypothesis to plausible fact.