AI companies are purchasing millions of physical books, scanning them for data, and then destroying the original copies [1].

This practice threatens the preservation of human knowledge by permanently removing rare and out-of-print titles from physical circulation to feed large-scale language models.

Reports indicate that firms, including Anthropic and other large AI labs, have targeted titles published before 2022 [2]. These companies seek vast volumes of textual data to improve the capabilities of their AI models. To facilitate this, books are sourced globally through services such as ISBNdb and then processed in facilities owned by the AI firms [3].

Once the text is digitized, the physical books are destroyed. This process includes the destruction of rare editions that may not have other surviving copies, creating a permanent loss of the physical artifact. The acquisition process relies on the legal right of a buyer to destroy a physical object they own, meaning the firms can legally discard the books after scanning them [3].

Industry response has been minimal. For example, ISBNdb removed its bulk-book service nine days after 404 Media reported on the practice [4]. However, this change did not stop the broader trend of AI labs acquiring and destroying physical libraries to secure a competitive edge in training data.

Critics have compared the systematic destruction of these texts to a modern Library of Alexandria event. The drive for data has created a scenario where the physical record of history is sacrificed for the efficiency of a digital model. Because many of these books are rare or non-recoverable, the loss is absolute once the disposal occurs [1].

AI companies are purchasing millions of physical books, scanning them for data, and then destroying the original copies.

The destruction of physical books for AI training highlights a growing tension between the acceleration of machine learning and the preservation of cultural heritage. As AI labs exhaust high-quality digital data, they are turning to the 'analog' world, treating physical books as raw material rather than historical artifacts. This creates a permanent data bottleneck where the only remaining copies of rare texts may exist solely within the proprietary, closed-source weights of a corporate AI model.