AI developers are purchasing rare and antique books to digitize their text for training large language models before destroying the physical copies.

The practice has sparked backlash from cultural heritage groups and archivists who argue that the permanent loss of these physical artifacts outweighs the technical gains of AI training.

Companies including Anthropic and reportedly Amazon have utilized U.S.-based procurement channels to source these materials [1, 2]. These channels include middle-man services that acquire books from collectors, rare-book dealers, and libraries worldwide [5, 2]. Reports indicate that AI companies have purchased millions of physical books for this purpose [2].

The process involves scanning the books to extract high-quality text corpora needed to improve the performance of large language models [1, 3]. Once the digitization is complete, the original volumes are destroyed. Methods of destruction vary by report; some sources said the books are shredded [4], while others described the use of hydraulic cutting machines to slice the volumes off at the spine [3].

Industry analysts said the destruction of the physical copies is a deliberate move to avoid public backlash that would occur if the companies were seen hoarding rare cultural artifacts [1, 4, 3]. By removing the physical evidence, firms can integrate the data into their models without maintaining a visible archive of the sourced materials.

The scale of the operation came to light following a lawsuit filed earlier this year [1]. The legal proceedings revealed the systematic nature of the procurement and destruction cycle used to feed the growing hunger for high-quality training data [1].

AI companies have purchased millions of physical books for this purpose.

This trend highlights a growing conflict between the data requirements of generative AI and the preservation of human history. As digital archives become the primary source of training, the 'data hunger' of LLMs is shifting from the open web to physical, rare assets. The destruction of these books suggests that AI firms view cultural artifacts primarily as raw data rather than historical objects, potentially creating a permanent gap in the global physical record of literature and thought.