AI News
  • Home
  • The Hague
  • Moving to NL
  • Tech News
    • AI & Tech
    • Machine Learning
    • Startups
    • Tools & Apps
    • Robotics
    • Future Tech
    • AI in Sport ⚽
    • AI in Health
    • AI in Education
    • AI in Finance
    • AI in Business
    • AI in Law
    • AI in Climate
No Result
View All Result
AI News
  • Home
  • The Hague
  • Moving to NL
  • Tech News
    • AI & Tech
    • Machine Learning
    • Startups
    • Tools & Apps
    • Robotics
    • Future Tech
    • AI in Sport ⚽
    • AI in Health
    • AI in Education
    • AI in Finance
    • AI in Business
    • AI in Law
    • AI in Climate
No Result
View All Result
AI News
No Result
View All Result

Amazon Is Destroying Rare Books to Train Its AI Models

Ramo by Ramo
17 August 2026
in AI & Tech
0
Amazon, which started off selling books, is destroying rare texts to train AI
0
SHARES
2
VIEWS
Summarize with ChatGPTShare to Facebook

Amazon began in 1994 as a website that sold books. Three decades later, it is buying rare ones so it can tear them apart.

According to a TechCrunch report published on August 17, 2026, the company has been destroying rare texts to feed its AI models. The mechanics are blunt. Physical books are acquired, unbound, scanned, and converted into the kind of clean digital text that large language models can ingest. What the machines gain, the shelves lose.

There is a grim symmetry to it. The company that made its name shipping paperbacks to doorsteps now sees printed books as raw feedstock, and the rarer the book, the more it is worth to an algorithm. That inversion says something about where the value in text has migrated, and it is not toward the reader.

📖
RECOMMENDED READ
The Coming Wave: AI, Power, and the Greatest Dilemma of Our Age
Mustafa Suleyman
The definitive book on where AI is heading - written by one of the field founders.
View on Amazon →affiliate link

Why the obscure book is suddenly the prize

The reason is straightforward once you understand how these models are built. Large language models have already been trained on the open internet, or something close to all of it. Every scraped web page, every digitized public-domain title, every forum thread and Wikipedia edit has been vacuumed up. The well of easily accessible text is running dry.

Rare books are one of the few places genuinely new material still lives. A limited-run monograph, an out-of-print technical manual, a title that never made it past a single small printing, none of that exists in a form a crawler can reach. It sits on paper, in a handful of copies, unindexed and unread by any machine. For a model builder chasing fresh training data, that scarcity is the whole point. The text has never been seen before, which means it can still teach the model something.

So the calculus flips. In the antiquarian trade, a rare book is precious because so few copies survive. In the AI trade, it is precious for the same reason, right up until the moment it is fed through a scanner and its binding is discarded. The value that made it worth acquiring is the value that gets consumed.

Destruction as a feature, not a bug

Why destroy the book at all? Non-destructive scanning exists, but it is slow and expensive. Cutting the spine off a volume and running the loose pages through a sheet feeder is faster and cheaper by a wide margin. When the goal is volume, throughput wins, and the physical object becomes an obstacle to be cleared rather than an artifact to be preserved.

That trade-off is easy to wave through when the book is a mass-market title with thousands of surviving copies. It reads very differently when the book is rare. A scanned file is not the same thing as the object it came from. It carries the words but not the marginalia, the printing quirks, the physical evidence that scholars and collectors actually study. Once the last few copies of something have been sliced up for data, the digital ghost is all that remains, and it remains inside a private model.

A quiet reordering of what text is for

Step back and the story is less about one company and more about a shift in what written material is understood to be. For most of the history of publishing, a book was something you read. Increasingly it is something you train on. The audience for a text is no longer only human, and the most valuable readers may be the ones that never open the cover, only the file.

Amazon is a fitting company to sit at that hinge. It spent years arguing about the future of reading, then built one of the largest cloud and AI operations on the planet. That it would eventually see the two businesses converge, treating the book both as a product and as a substrate, feels less like a contradiction than a logical endpoint.

The uncomfortable part is what it implies about everything still trapped on paper. If the scarce and unscanned is now the frontier, the pressure to convert it will only grow, and conversion here can mean consumption. The open question is whether anyone is drawing a line between the copies that can be spared and the ones that cannot, before the scanners decide it for them.

For more coverage of AI training data, visit Mylistingo.

Source: Original Article

SummarizeShare
Ramo

Ramo

Ramo is the editorial voice of Mylistingo — an AI and technology news platform based in The Hague, Netherlands. Covering artificial intelligence, machine learning, robotics, and the future of technology, Ramo delivers accurate, accessible reporting for both general audiences and industry professionals. Every article is fact-checked and written to meet Mylistingo's strict no-fabrication editorial standards.

Related Stories

Clinician reviewing patient data on a laptop with a stethoscope nearby

OpenAI Plugs ChatGPT Into Epic’s Patient Records

by Ramo
4 September 2026
0

OpenAI connected ChatGPT for Healthcare to Epic, the record system behind 325 million patients. Access is read-only, and the stakes are high.

Abstract visualisation of an artificial intelligence network

Anthropic’s Fable 5.1 Cuts AI Agent Costs by Up to 45%

by Ramo
3 September 2026
0

Claude Fable 5.1 keeps base prices flat but cuts cache-read costs 75%, making long-running AI agents up to 45% cheaper. Mythos 5.1 stays gated.

Nvidia’s $3.5B MediaTek bet reveals its plan for tackling Big Tech’s AI chip buildout

Nvidia’s $3.5B MediaTek Bet Against Big Tech AI Chips

by Ramo
31 August 2026
0

A $3.5 billion vote of confidence, and self-defense Nvidia just wrote a $3.5 billion check to a company that doesn't build the chips everyone associates with the AI...

Musk’s faster path to more gas turbines comes with pollution problem

Musk’s Gas Turbine Bet: Faster Power, Dirtier Air

by Ramo
31 August 2026
0

A foundry, a shortcut, and a fuel nobody wants next door Elon Musk has a habit of solving other people's bottlenecks by building the part himself. His latest...

Next Post
Anthropic’s annualized revenue surges to $65B

Anthropic's Revenue Hits $65B Annualized Run Rate

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

The Hague, for internationals

One email a week: what changed for expats in The Hague, what's on this weekend, and one guide worth reading. No spam, unsubscribe any time.

Free European Bank Account
Open a 100% mobile bank account in minutes
Free virtual Mastercard, zero foreign transaction fees, and instant European IBAN setup with no paperwork.
Get Started Free
Sponsored · Advertise

Recommended

Abstract technology and data visualization background

EU Cybersecurity in 2026: NIS2 Enforcement, AI Threats, and the New Cyber Solidarity Network

11 July 2026
AI Medical Diagnosis Breakthroughs Transform Healthcare in 2026

AI Medical Diagnosis Breakthroughs Transform Healthcare in 2026

16 July 2026

Popular Story

  • ml_feat_56193023

    ASML’s Next-Gen High-NA EUV Machines Drive Eindhoven Expansion, Creating 20,000 New Jobs

    0 shares
    Share 0 Tweet 0
  • PixVerse closes $439m series C extension at $2b valuation

    0 shares
    Share 0 Tweet 0
  • Is Your Home Truly Safe The Smart Security Tech You Need in 2025

    0 shares
    Share 0 Tweet 0
  • How to Register at The Hague Municipality (Gemeente Den Haag): A 2026 Step-by-Step Guide

    0 shares
    Share 0 Tweet 0
  • The Rise of Neuromorphic Computing: How Brain-Inspired Chips Are Transforming AI in 2026

    0 shares
    Share 0 Tweet 0
logo ainews

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Recent Posts

  • MIT’s CW-Net Makes Self-Driving AI Explain Itself
  • OpenAI Plugs ChatGPT Into Epic’s Patient Records
  • AI Boom Puts Tech’s Climate Pledges Under Strain

Partner

Free European Bank Account
Open a 100% mobile bank account in minutes
Free virtual Mastercard, zero foreign transaction fees, and instant European IBAN setup with no paperwork.
Get Started Free
Sponsored · Advertise

Categories

  • AI & Tech
  • AI & Tech in the Netherlands
  • AI in Business
  • AI in Climate
  • AI in Education
  • AI in Finance
  • AI in Health
  • AI in Law
  • AI in Sport
  • Economy & Finance
  • Future Tech
  • Machine Learning
  • Moving to the Netherlands
  • Politics & Geopolitics
  • Robotics
  • Social Topics
  • Sport
  • Startups
  • The Hague
  • Tools & Apps
  • Uncategorized

The Hague, for internationals

One email a week: what changed for expats, what's on, one guide worth reading.

  • Home
  • Advertise
  • Latest News
  • Contact Us
  • Data Deletion Instructions
  • Editorial Policy

No Result
View All Result
  • Home
  • The Hague
  • Moving to NL
  • Tech News
    • AI & Tech
    • Machine Learning
    • Startups
    • Tools & Apps
    • Robotics
    • Future Tech
    • AI in Sport ⚽
    • AI in Health
    • AI in Education
    • AI in Finance
    • AI in Business
    • AI in Law
    • AI in Climate