AI companies are purchasing millions of rare and antique books to scan their contents for model training before destroying the physical copies [1].

This practice highlights a growing desperation for "clean" data. As AI-generated text floods the internet, developers are returning to physical archives to find human-authored material that has not been contaminated by synthetic output.

Reports indicate that firms such as Amazon and Anthropic have been involved in these acquisitions [1, 2, 3]. Some of these books were shipped to an AI-training warehouse operated by Amazon in Las Vegas, Nevada [2]. Once the text is digitized, the original physical copies are discarded [1, 3].

The scale of the operation is vast, with some documents indicating that millions of physical books have been bought for scanning [3]. This effort focuses specifically on text produced before 2022 [1, 3]. Developers believe that material from this era is free of AI-generated content, making it more valuable for training large language models.

There is some conflicting information regarding which specific company led the purchasing efforts. While reports link the shipments to an Amazon facility in Las Vegas [2], other documents from a lawsuit suggest that Anthropic purchased millions of the books [3].

The process involves high-volume scanning to convert physical pages into machine-readable text. This allows AI models to learn from complex human language patterns found in older literature and academic texts, sources that are less likely to be mirrored in the current, AI-saturated digital landscape [1, 3].

AI companies are purchasing millions of rare and antique books to scan their contents for model training before destroying the physical copies.

The shift toward physical book acquisition signals a critical turning point in AI development known as 'model collapse.' When AI models are trained on data generated by other AI, the resulting output degrades in quality and diversity. By destroying physical copies of pre-2022 texts after scanning them, these companies are treating historical archives as raw industrial fuel, prioritizing the immediate needs of algorithmic training over the long-term preservation of physical cultural heritage.