Regulation
Amazon Is Destroying Rare Books to Train Its AI Models
404 Media tracked a rare book with an AirTag to an Amazon warehouse, where spines are cut off for scanning and the text is used to train Nova models.

Putting an AirTag inside a rare book turned out to be the most productive piece of reporting this week.
404 Media bought a rare book, hid a tracker in the shipment, and followed it across the country. It came to rest in an Amazon warehouse in Las Vegas, home to a team called VGT3 whose logo is a T-Rex holding a book. The work there involves cutting the spines off books so the pages feed through a scanner quickly. After the scan, the book is no longer a book.
Where the scanned text goes
The extracted text trains Amazon's Nova model family. The company's comment on the practice is brief: it "buys books through commercial channels to improve its products."
Booksellers believe AI companies are working through catalogues systematically by ISBN. Printed text is unusually valuable for two reasons. Much of it exists only physically and has never been online at all. And it predates 2022, which means no AI-generated text is mixed in. Most of the open web lost that second property some time ago.
A court ruling may have encouraged the method
Anthropic ran the same operation under the name Project Panama, buying books on the open market, stripping the spines, and digitizing them. A court found that scanning to be fair use, and part of the reasoning is worth sitting with: the printed originals were destroyed, and therefore not copied and resold.
So the safe harbour runs through the shredder. The same company's $1.5 billion copyright settlement over pirated books, approved in July, explains why that route looks attractive: buy, scan, destroy is the path that survives litigation.
The procurement question this raises
Set the ethics aside for a moment and a concrete business fact remains. Human-written text from before 2022 is now a scarce input, and companies are moving it inside closed systems. When the only surviving copy of a rare title is scanned and then pulped, that knowledge leaves the public record and enters a set of model weights.
For buyers, this belongs in vendor diligence. Asking what a model was trained on used to be a curiosity question; for anyone selling into markets with training-data transparency rules, it is now a compliance question. Does the contract include copyright indemnification? Does the provider publish a summary of its training data? Those lines belong in the RFP, not in a blog post you read afterwards.
The mirror image applies internally. Your own archive of documents is your equivalent of the pre-2022 printed book: material nobody else has, generated by your own operations. Load it into an AI tool by all means, but know exactly what you are granting when you do.
Sources: TechCrunch, The Decoder, 404 Media

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.