Back
AI companies are destroying physical books, Anna's Archive post says, urging volunteers to scan
SiTech AI Team3 წთ. საკითხავი

AI companies are destroying physical books, Anna's Archive post says, urging volunteers to scan

A guest post on Anna's Archive's blog says several AI companies buy, scan and then destroy secondhand books to get pre-2022 training data, and calls on volunteers to scan rare material before it is lost.

Anna's Archive, which describes itself as the largest shadow library, has published a guest post by one of its volunteers — signed "u" and translated from Chinese — arguing that several AI companies buy large quantities of physical books, scan them and then destroy them, and calling on volunteers worldwide to scan rare material before it is lost.

The post is dated 2026-08-05 on the archive's blog and has circulated widely since.

What the post describes

According to the author, AI companies acquire secondhand books in bulk through intermediaries, scan them and destroy them, in order to obtain training data "untouched by machines" from before 2022. The text points to Anthropic's "Project Panama", which it says was exposed in a $1.5 billion copyright settlement: launched in early 2024 as a highly confidential effort, it involved spending tens of millions of dollars on millions of paper books, scanning them to train Claude, and then destroying them all.

Because the books are destroyed after digitisation, the company that scanned them is left as the only holder of the digital copies. In the author's reading, that locks knowledge permanently inside private corporate servers.

Why destroy the books

The post lists three motives: to prevent competitors from scanning the same books and training on them; to limit legal exposure; and because destroying a book is cheaper than scanning it losslessly. The author calls this legally permissible but, in his words, an extremely serious crime against humanity — a charge aimed at the practice rather than at anyone using the resulting models.

He frames the situation as a paradox: companies that promise to make human knowledge accessible are dismantling the most durable physical carriers of that knowledge, so the public may gain smarter assistants while a large body of knowledge leaves the public domain.

A call for volunteer scanning

Anna's Archive asks volunteers to scan and upload books, journal articles, newspapers, magazines, ancient and rare books from libraries and archives around the world, prioritising material that is easily lost. The arithmetic offered is simple: if each of 10 million volunteers scans one book, 10 million items are preserved. Small contributions are met with recognition and lifetime membership, while for large-scale scanning and uploading the archive says it can help cover scanning fees. The post describes the effort as a race against time and warns that since the beginning of 2025 AI-generated material has made up more than half of newly published internet content. It links to tickets #223 and #187 for further detail.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.