Amazon began in 1994 as a website that sold books. Three decades later, it is buying rare ones so it can tear them apart.
According to a TechCrunch report published on August 17, 2026, the company has been destroying rare texts to feed its AI models. The mechanics are blunt. Physical books are acquired, unbound, scanned, and converted into the kind of clean digital text that large language models can ingest. What the machines gain, the shelves lose.
There is a grim symmetry to it. The company that made its name shipping paperbacks to doorsteps now sees printed books as raw feedstock, and the rarer the book, the more it is worth to an algorithm. That inversion says something about where the value in text has migrated, and it is not toward the reader.
Why the obscure book is suddenly the prize
The reason is straightforward once you understand how these models are built. Large language models have already been trained on the open internet, or something close to all of it. Every scraped web page, every digitized public-domain title, every forum thread and Wikipedia edit has been vacuumed up. The well of easily accessible text is running dry.
Rare books are one of the few places genuinely new material still lives. A limited-run monograph, an out-of-print technical manual, a title that never made it past a single small printing, none of that exists in a form a crawler can reach. It sits on paper, in a handful of copies, unindexed and unread by any machine. For a model builder chasing fresh training data, that scarcity is the whole point. The text has never been seen before, which means it can still teach the model something.
So the calculus flips. In the antiquarian trade, a rare book is precious because so few copies survive. In the AI trade, it is precious for the same reason, right up until the moment it is fed through a scanner and its binding is discarded. The value that made it worth acquiring is the value that gets consumed.
Destruction as a feature, not a bug
Why destroy the book at all? Non-destructive scanning exists, but it is slow and expensive. Cutting the spine off a volume and running the loose pages through a sheet feeder is faster and cheaper by a wide margin. When the goal is volume, throughput wins, and the physical object becomes an obstacle to be cleared rather than an artifact to be preserved.
That trade-off is easy to wave through when the book is a mass-market title with thousands of surviving copies. It reads very differently when the book is rare. A scanned file is not the same thing as the object it came from. It carries the words but not the marginalia, the printing quirks, the physical evidence that scholars and collectors actually study. Once the last few copies of something have been sliced up for data, the digital ghost is all that remains, and it remains inside a private model.
A quiet reordering of what text is for
Step back and the story is less about one company and more about a shift in what written material is understood to be. For most of the history of publishing, a book was something you read. Increasingly it is something you train on. The audience for a text is no longer only human, and the most valuable readers may be the ones that never open the cover, only the file.
Amazon is a fitting company to sit at that hinge. It spent years arguing about the future of reading, then built one of the largest cloud and AI operations on the planet. That it would eventually see the two businesses converge, treating the book both as a product and as a substrate, feels less like a contradiction than a logical endpoint.
The uncomfortable part is what it implies about everything still trapped on paper. If the scarce and unscanned is now the frontier, the pressure to convert it will only grow, and conversion here can mean consumption. The open question is whether anyone is drawing a line between the copies that can be spared and the ones that cannot, before the scanners decide it for them.
For more coverage of AI training data, visit Mylistingo.
Source: Original Article







