Amazon destroyed rare out-of-print books for AI training data, GPS investigation finds
A GPS tracker placed inside a rare book led investigators to an Amazon facility in Las Vegas where employees strip the spines off books and scan the pages for AI training data. The investigation, published by 404 Media on August 17, is the first documented trace of Amazon's physical book-scanning operation.
What happened
404 Media reporter Emanuel Maiberg listed a rare book for sale online, inserted a GPS tracker before shipping, and monitored the package as it traveled to Amazon's VGT3 facility in Las Vegas, Nevada. Per the investigation, employees at VGT3 receive large book shipments, cut the bindings off to allow faster flatbed scanning, and discard the physical copies once the pages are digitized. The team operating VGT3 uses a logo depicting a dinosaur holding a book.
Amazon confirmed the core finding in a statement to 404 Media. The company "purchases books through commercial channels to improve the products and services customers use," per the investigation, and declined to elaborate. TechCrunch and Futurism both reported on the story within hours of the investigation publishing on August 17.
Rare and out-of-print books represent a specific category of training data. Titles published before 2022 cannot have been generated by a language model, making them high-confidence human-written text. Labs that have already absorbed most available internet text look for sources that are provably pre-generative. Physical books that were never digitized or went out of print before large-scale crawling fit that requirement exactly.
Why it matters
For authors and publishers, the investigation makes visible a mechanism that had been suspected but not documented: books that went out of print are entering AI training pipelines without licensing agreements or royalty payments. Rare titles do not lose copyright status when they leave print. Under US copyright law, protection runs for the author's life plus 70 years, which means the large majority of titles VGT3 processes remain under active protection.
Amazon's book operations have faced copyright questions before, including around Kindle Unlimited licensing terms with publishers. Training data acquisition operates under a contested fair-use argument currently being tested in litigation involving OpenAI, Meta, and others. If courts reject that argument in those cases, book-scanning operations at any lab face significant exposure. Amazon's confirmed statement does not dispute the scanning practice; it frames it as routine commercial purchasing.
The methodology 404 Media used is significant in its own right. A GPS tracker in a paperback exposed a corporate training-data operation Amazon had not disclosed. It demonstrates that the physical supply chain for AI training data can be traced with consumer hardware, a technique likely to attract imitators at publications covering the same beat.
What to watch next
Whether the Authors Guild or similar organizations file a formal complaint citing this investigation is the first thread to track. Congressional hearings on AI training data have grown more frequent in 2026, and this investigation gives committee members a concrete, documented case to reference. The second question is whether investigative journalists apply the same GPS methodology to book-scanning operations at other labs or data brokers.
Sources
- We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility: 404 Media, Emanuel Maiberg, August 17, 2026 (primary)
- Amazon, once an online bookseller, is destroying rare books to train AI models: TechCrunch, August 17, 2026 (secondary)
- Amazon Caught Destroying Rare Books to Train AI: Futurism, August 17, 2026 (secondary)
