Amazon, once an online bookseller, is destroying rare books to train AI models

| Source: TechCrunch AI

Tags: Amazon, AI training data, LLM training, copyright, data sourcing, 404 Media

Amazon is physically destroying rare books — cutting off spines and scanning them — to generate training data for its LLMs, according to 404 Media, which tracked a book to Amazon's VGT3 facility in Las Vegas.

Details

A 404 Media investigation found that Amazon is purchasing rare and out-of-print books, cutting off their spines, and scanning them for AI training data. The outlet confirmed the practice by placing a GPS tracker inside a rare book, which eventually arrived at an Amazon facility in Las Vegas called VGT3, identified by a dinosaur-holding-book symbol. Amazon confirmed the activity in a statement: it 'purchases books through commercial channels to improve the products and services customers use.' The company declined to elaborate further. Rare books are particularly valuable for LLM training because they predate the era of AI-generated content. Models that ingest AI-generated text risk 'model collapse' — a degradation in output quality over successive training rounds — making pre-2022 human-authored text especially prized. Unlike books available online, out-of-print titles offer genuinely novel data unavailable from web crawls. The practice raises unresolved questions about copyright, rights of original authors and estates, and preservation ethics. Amazon is not alone in this data sourcing pressure: Anthropic has faced legal challenges over book use, and most frontier labs are actively seeking non-internet text sources as online data exhaustion becomes a real constraint.