Artificial intelligence companies are purchasing and destroying millions of printed books to extract text for training large language models [1].
This practice highlights a growing desperation for high-quality data as the internet becomes saturated with AI-generated content. Because models can degrade when trained on their own output, firms are returning to physical archives to find "clean" human language.
Reports indicate that these technology firms are acquiring books from libraries and booksellers worldwide [2]. Once the books are purchased, the companies scan the pages to ingest the contents into their databases. After the digitization process is complete, the physical copies are shredded [3].
Industry observers said that millions of printed books have been bought and destroyed through this process [1]. This includes the acquisition of antique and rare books, which are scanned for their unique contents before being destroyed at scale [3].
The shift toward physical media is driven by the need for reliable training sets. As the volume of synthetic text increases online, the value of printed materials—which predates the current AI surge—has risen. These firms are effectively treating the world's physical libraries as raw material for digital intelligence [4].
Critics of the practice said the permanent loss of physical history is a concern. While the text is preserved digitally, the original artifacts are lost forever. The scale of the destruction is reported to be in the millions of volumes [2].
Companies have not detailed the specific lists of titles being targeted, but the goal remains the same: obtaining the highest quality human-written text available [2].
“AI firms are quietly buying and destroying millions of printed books to train their models”
The transition from web-scraping to the physical destruction of books signals a critical turning point in AI development. It suggests that the 'digital gold rush' has exhausted the available high-quality public internet data, forcing companies to monetize and destroy physical archives to avoid model collapse caused by synthetic data loops.



