BOOK WAR: AI Firms Accused of Destroying Books for Data.
A growing backlash is erupting over the artificial intelligence industry’s aggressive effort to acquire physical books for machine-learning data, with Anthropic’s secretive “Project Panama” at the center of the controversy. Newly unsealed court records show that Anthropic pursued a large-scale operation to purchase books, remove their bindings, scan their pages and discard the original physical copies. Internal planning documents described the initiative as an effort to “destructively scan all the books in the world,” raising a startling question: how far should technology companies go to obtain the data needed to build increasingly powerful AI systems?
The operation was not simply about ordinary new releases. Court records and reporting indicate that Anthropic sought huge quantities of high-quality books, including older and out-of-print works, because books provide long-form human-written material that can be valuable for training large language models. The physical copies were processed by stripping or cutting away their bindings so pages could be scanned rapidly, after which the originals were discarded or recycled. Reports have described purchases reaching into the millions of books and spending in the tens of millions of dollars, although the precise scope varied across different stages of the effort.
The revelations have intensified concerns among authors, booksellers, collectors and preservation advocates who fear that an industrial appetite for training data could put culturally significant physical works at risk. Books are not merely containers of text: some surviving copies can have historical, archival or collectible value that cannot be restored once a volume has been cut apart. Booksellers have also reported unusual bulk purchasing patterns, fueling speculation about whether some large orders were ultimately intended for AI-related scanning rather than conventional resale.
The legal picture, however, is more complicated than the dramatic destruction of the books might suggest. In 2025, a federal judge ruled that Anthropic’s use of lawfully purchased books to create digital copies for its library was fair use, describing the format conversion as transformative, while separately finding serious problems with Anthropic’s earlier acquisition and retention of pirated books. Anthropic later agreed to a $1.5 billion settlement with authors over claims involving pirated books, meaning the legal fight over AI training data has continued even as companies increasingly seek books through lawful purchases.
That distinction has done little to quiet the broader debate. The controversy now extends beyond whether scanning purchased books can qualify as fair use and into a much larger question about what happens when technology companies possess the money and computing power to acquire enormous portions of the world’s written culture. As AI developers race to secure high-quality human-created data, critics are asking whether commercially purchased books should be treated as disposable raw material—and whether the drive to build smarter machines is moving faster than society’s ability to protect the books, authors and cultural history being consumed along the way.