Major artificial-intelligence companies are purchasing rare antique books and destroying the physical copies to digitize the text for training large-language models.
This practice represents a significant loss of physical cultural heritage. By destroying the original volumes, these companies are permanently erasing unique historical artifacts to fuel the growth of digital AI systems.
Reports published this month indicate that firms, including Anthropic and other large AI providers, have used anonymous middle-man services to acquire these materials [1, 4]. The process involves slicing off bindings or shredding the original books after the content has been digitized [1, 2, 3].
According to reports, AI companies have bought millions of physical antique books for this purpose [2]. The scale of these acquisitions suggests a systemic effort to gather vast amounts of textual data that may not be available in digital archives.
Industry observers said the companies use anonymous intermediaries to avoid public backlash over their data-harvesting practices [1, 5]. This method allows the firms to acquire rare texts without alerting historians, librarians, or the general public to the destruction of the physical assets [1].
The focus on antique books stems from the need for high-quality, diverse textual data to improve the reasoning and knowledge capabilities of AI models [5]. As digital data sources become exhausted or restricted by copyright, companies have turned to physical archives to find untapped information [1, 5].
Critics said the destruction of these books is an unnecessary step in the digitization process. While scanning technology can preserve a book without damaging it, the reported shredding and slicing of bindings suggest a streamlined industrial process designed for speed rather than preservation [3].
“AI companies have bought millions of physical antique books for training data”
The destruction of physical books for AI training highlights a growing conflict between the demands of machine learning and the preservation of human history. As AI companies exhaust available internet data, the pursuit of 'clean' or rare training sets may lead to the irreversible loss of physical archives, shifting the value of historical texts from cultural artifacts to raw data points.


