AI companies destroy physical books – let's scan rare books before it's too late
Anna's Archive alleges AI companies are destroying physical books after scanning them for training data, creating a knowledge monopoly and erasing cultural heritage. This provocative claim, framed as a desperate call for volunteers to digitize books, ignited intense debate on HN. Commenters questioned the factual basis of the destruction, the role of copyright, and the motivations behind such a sensationalist warning.
The Lowdown
The article, a guest post from Anna's Archive volunteer 'u', makes a stark accusation: AI companies are covertly purchasing, scanning, and then destroying physical books to train their models, creating a 'cultural crime against humanity'. It highlights a race against time to preserve human knowledge.
- AI companies, exemplified by Anthropic's 'Project Panama', are reportedly acquiring vast quantities of secondhand books, scanning them for data, and then destroying them.
- This destruction is said to be driven by a desire for 'untouched by machines' pre-2022 data, preventing competitors from using the content, avoiding legal risks, and being cheaper than lossless scanning.
- The author argues this practice leads to a permanent monopolization of human knowledge on private servers, paradoxically dismantling knowledge carriers while promising accessibility.
- Anna's Archive, as a major 'shadow library', calls for global volunteers to scan and upload books, especially rare and easily lost materials, to combat this trend and preserve knowledge in the public domain.
- The piece warns of a future where AI-generated content dominates the internet, emphasizing the critical role of shadow libraries in preserving human civilization's memory.
The article positions this as an urgent battle against knowledge monopolization by both publishers and AI companies, urging collective action to ensure free access to humanity's wealth for future generations.
The Gossip
Skeptical Scrutiny: Are the Claims Credible?
Many commenters expressed doubt regarding the article's core claim that AI companies are systematically destroying physical books, particularly 'rare' ones. They questioned whether 'rare' was being used loosely, noting that most books have multiple copies, and suggested the narrative might be an exaggeration or a strategic marketing tactic by Anna's Archive. Some even posited that copyright holders might be subtly manipulating the situation to force licensing deals for e-books.
Copyright's Complications: Fueling the Fire?
A significant portion of the discussion centered on the paradoxical role of copyright law. Commenters argued that restrictive copyright, enforced by rights holders who often don't reprint older works, forces AI companies to acquire physical copies and, in some interpretations, contributes to their destruction. This legal landscape is seen as inadvertently pushing companies towards physical media, highlighting the tension between intellectual property rights and knowledge accessibility.
Scanning Scenarios: A Call for Collective Action
While some agreed with the urgency of preserving knowledge, many questioned the feasibility of Anna's Archive's ambitious call for millions of volunteers to scan books. Critics highlighted the immense time and effort required for home scanning, casting doubt on the practicality of such a large-scale, decentralized effort, even though one comment linked to resources for home scanning.