A guest post from Anna's Archive has ignited a huge debate: AI companies are reportedly buying up large quantities of secondhand books through intermediaries, scanning them, and then destroying the physical copies. The goal is training data untouched by machines, mostly from before 2022. Anthropic's Project Panama was exposed during a 1.5 billion dollar copyright settlement, with the company said to have spent tens of millions buying millions of paper books to train Claude before destroying them.
The practice is legally permissible but ethically alarming. Destroying the physical books ensures competitors cannot scan and train on the same copies, locking knowledge inside private corporate servers. Critics argue it erases cultural heritage: rare and out-of-print works are being lost permanently, with no digital copies left behind. The story has drawn over 800 comments on Hacker News, splitting readers between outrage at the waste and debate over the legality of the shadow library response.
In response, Anna's Archive is urgently calling on volunteers worldwide to scan and upload rare books before they disappear. The broader lesson for founders: the AI training data race is reshaping what happens to physical knowledge, and questions about provenance, consent, and preservation are only going to get louder as model makers compete for unique data.
