How AI Companies Are Destroying Rare Books to Train Their Models
AI-generated, human-reviewed.
AI firms are aggressively acquiring and destroying rare physical books to feed their hunger for high-quality training data, according to investigative journalist Emmanuel Maiberg on Tech News Weekly. This behind-the-scenes process involves Amazon and other companies buying these books in bulk, disassembling them for rapid digitization, and ultimately shredding them to comply with copyright rules—all in the name of improving artificial intelligence.
Why Are AI Companies Buying and Destroying Books?
Emmanuel Maiberg explained that the rush to collect rare physical books started after AI companies exhausted accessible online data sources. Modern large language models (LLMs) like ChatGPT have already been trained on the open web, articles, and e-books available online. However, continued improvements require fresh, diverse, and genuinely human-written content, making rare and print-only books highly valuable.
Printed books, especially those published before late 2022, supply unique text not tainted by already AI-generated material—critical for avoiding what researchers call "model collapse." Essentially, if AI starts learning from content written by other AIs, its accuracy and quality degrade over time.
The Process: From Seller to Shredder
Maiberg's investigation detailed the journey of a rare book tracked by an AirTag after being sold. Once purchased, the book passed through multiple warehouses before arriving at an Amazon facility in Las Vegas specializing in “print-on-demand.” However, a subset of the warehouse, run by a team known as VGT-3, had a different focus: cutting the spines off books and rapidly scanning their pages for digital conversion.
The books are destroyed in this process, because removing the binding and feeding single pages to a scanner is far faster (and cheaper) than scanning them gently page by page. This also aligns with recent legal guidance: If a company destroys the print copy after digitization, it can argue in court that only one copy exists—potentially qualifying for fair use protection. The original purpose for cutting bindings and digitizing books predates AI, but the scale and frequency have dramatically increased.
What Are Companies Using This Data For?
Amazon confirmed to Maiberg that these scanned books help “develop and improve products and services,” a vague statement encompassing several possibilities:
- AI Training Data: Feeding rare, high-quality text into LLMs to improve reasoning and reduce errors.
- Print-on-Demand Expansion: Creating digital versions for future printing.
- Searchable Digital Archives: Enabling users to search within the content of books, similar to Google Books.
For AI companies, this move is a way to secure a temporary advantage in producing models with richer, more nuanced knowledge.
Community Reaction: Ethical and Cultural Concerns
Book collectors and rare book sellers initially celebrated historic sales spikes, only to grow uneasy as they realized why so many books were being purchased. Many feel discomfort over contributing to the mass destruction of rare cultural objects—even if these aren’t universally known classics, their historical, linguistic, or niche value might be recognized only in the future. There’s a widespread cultural aversion to the destruction of books, regardless of intent.
Libraries and institutions have occasionally destroyed books to save space, but AI companies are doing it for data rather than preservation or utility. The possibility that irreplaceable knowledge could be lost forever for the sake of technological progress is a growing concern.
Legal Gray Area and the Future of Knowledge
The process takes advantage of a legal loophole: destroying the original print copy after digitization can make the copying defensible under US fair use law. However, Maiberg stressed that mass digitization at the cost of permanent destruction may deprive future generations and historians of unexpected treasures.
As AI companies devour more physical books, fears mount about what lessons, stories, and languages could vanish for good.
Key Takeaways
- AI companies have exhausted most online training data and are buying rare books as the next source.
- Many books are destroyed in the digitization process to speed up scanning and avoid copyright complications.
- Ethical concerns are rising about the destruction of rare and potentially culturally significant works.
- Book sellers experienced a surge in sales, not realizing the books were being purchased to be torn apart for AI.
- Legal gray areas exist, as destroying the original book after scanning may protect companies under fair use.
- Amazon and others confirmed they use this data to improve AI and related services.
- Loss of unique records and writings is possible, potentially erasing resources future historians might need.
- Public discomfort and debate are growing as the practice becomes more widely known.
The Bottom Line
According to Emmanuel Maiberg on Tech News Weekly, the race to create smarter AI has led companies to quietly destroy rare books by the thousands. While this may boost technical progress in the short term, it raises tough questions about cultural preservation, ethics, and the long-term consequences of sacrificing analog knowledge for digital innovation.
Subscribe to Tech News Weekly for expert tech reporting and thoughtful interviews:
https://twit.tv/shows/tech-news-weekly/episodes/451