AI companies are buying used books by the thousands. Some may be destroyed for training
| Source: Fast Company AI
Tags: training data, copyright, data licensing, OpenAI, Meta, used books
AI companies are bulk-purchasing used books from secondhand bookstores to scan as legally acquired training data, giving booksellers a surprise sales boost while raising questions about whether digitizing purchased physical copies infringes copyright.
Details
Multiple AI labs are placing large bulk orders at used bookstores, acquiring physical books to scan and use as training data. The strategy lets companies claim legal ownership of the physical objects — sidestepping publisher licensing negotiations — while the actual copyright status of digitizing and training on those works remains legally ambiguous. Used booksellers report unusually large orders and an unexpected revenue boost from this activity. Some purchased volumes may be destroyed after scanning rather than resold.\n\nThe tactic parallels earlier approaches that led to major lawsuits: The Authors Guild and publishers have sued OpenAI, Meta, and others over training data practices. As courts weigh fair use arguments in those cases, companies appear to be seeking data acquisition paths with cleaner legal optics. Physical purchase is unambiguously legal; whether it confers rights to digitize and train commercially is the open question.\n\nThe scale of the practice suggests it's not limited to one lab. Used book markets may experience sustained demand as AI companies exhaust higher-quality digital sources and face tighter licensing terms from publishers who have grown wary of these arrangements.