Panama should be associated with happy thoughts and shady hats
...not with destroying valuable books.
If you have already heard of Project Panama, skip to the next thing in your feed. You can already imagine how I feel about it.
For those of you who have not heard of it, allow me to introduce you to what began as a good idea but was too quickly enshittified.
First, I am not dogmatically anti-AI. In fact, I coach people in how to use AI ethically, effectively, and wisely. But I share John Warner’s key concerns about AI and am cautious about how I use it in my own life and practice (which is where the “wisely” comes into play). So I understand that LLMs need to be trained, and that they need to be trained on the most accurate content, and the best possible writing, that you can find for them. This is not found in social media. Blogs are better, especially highly-rated and professionally produced blogs (like this one!). Respected magazines and journals are (or should be) another step up the ladder of factuality, reasoning, and quality of prose.
The Gold Standard: Authoritative Books.
But the gold standard is books from respected publishers, especially ones that are recognized as authoritative in their field, or strong contributions to their topics or genres. Anthropic et alia have known this from the beginning. They have tried to train their models on good books from the get-go, claiming it is “fair use” to do so, and therefore not a copyright violation if they scan books without paying for them.1
Setting aside the fair-use debate for now, the next model-training scandal has surfaced.
Project Panama
Borrowing language from its defenders,2
Project Panama is Anthropic’s large-scale book digitization initiative, created to build a high-quality library of published works for training and improving large language models. By acquiring millions of physical books, scanning them into searchable digital text, and converting them into a machine-readable corpus, Project Panama aims to expose AI systems to the breadth and depth of human knowledge and writing. According to court documents, the goal was to create a comprehensive collection of high-quality books that would help language models learn from carefully edited, professionally published prose rather than relying primarily on the uneven quality of the public internet.
Sounds great! As a casual user of LLMs, dictation software, AI search-and-summary, and perhaps agentic AI in the future, I would much rather see them trained on “the breadth and depth of human knowledge and writing” based on “carefully edited, professionally published prose” than commonplace internet dreck and data contamination from other chatbots.3 As one used-book distributor put it, these books are “curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”4
This is becoming a challenge. Since 2022, I don’t know how many modern authors, fiction or nonfiction or in academia, have been using AI-generated text in their books, but it hasn’t been zero. And it isn’t reliably disclosed when it is done. Lena McDonald, KC Crowne, and Rania Faris are just the most public examples of undisclosed and clumsy use of AI in their recently published books. It’s hard to be sure how common it is these days for authors of all sorts to generate text with AI and then use it, unrevised and unedited by humans, in their books… but given the risk of recursive contamination, AI companies really want to avoid this.
So, Project Panama. Let’s focus on all those great books published before generative AI was introduced! Choose books that “dense, edited” from every decade, in every genre and specialty. Buy them, instead of slipping them quietly out of a library or downloading them from a piracy site. Sounds great so far.
So why was Anthropic keeping Project Panama secret?
Because Project Panama is buying up huge quantities of books, including rare books, and destroying every one of them after they have scanned it.
Yes. Every physical book they locate and purchase, they cut off the spine, scan all the pages, and shred & pulp the thing.5
Anthropic internal memo: “Project Panama is our effort to destructively scan all the books in the world.”6
For common books, this is no big deal. Books are lost or accidentally destroyed every year, especially academic and nonfiction books, and publishers don’t mind printing and selling additional copies. Just one less copy of that book in the world, no one will notice. Or so Anthropic hopes. And honestly, it’s a decent justification for commonplace books, especially those still in print (publisher-speak for “still being printed, ready when you want it”).
But what about out-of-print books and books that are hard to find?
Many of these are not valuable because they are famous or have artistic merit. They are simply solid bastions of their niche, well-researched, well-written, respected, and rare… and even if Project Panama only purchases one copy and pulps it, that’s still one less physical copy of a rare book, and it isn’t going to be replaced.
And Project Panama seems like just one of several, possibly many(?), similar secret operations. If even Anthropic, which seems better-behaved than other big AI companies, is destroying the books it scans, this is worse than any book-burning crusade, since it’s destroying high-quality books indiscriminately, even the boring practical books and the ones considered “good books” by whoever might want to burn “bad books” these days.
What’s more, I suspect that Project Panama may be destroying the books they scan in order to make them even more rare, more difficult for other AI companies to find, increasing the value of their (enormous, very-high-quality, human-writing-only) training model. I haven’t seen evidence yet to prove this, but it’s the sort of thing that companies do when they are vying to be a monopoly, and hope to be the last one standing. And ISBNdb, one of the large-volume providers of books to Project Panama, seems to have other similar very-large-volume clients…7
Viscerally, just the thought of millions of books being fed into a destructive scanner and coming out the other side as paper pulp …it sickens me.
As Jonathan Last says,
Yeah, my feelings are not mixed. It is one thing to buy up books to train your LLM. It is another to actively destroy this supply of books to create a moat so that other LLMs can’t get access to them in the future.
This is the digital equivalent of strip-mining, or clear-cutting. It ought to be illegal.
And One More Thing…
It didn’t have to be this way. It still doesn’t. Destructive scanning isn’t the only way to scan a million books per week (or whatever their pace was going to be). There are very-high-speed book scanners that don’t destroy the book as it is scanned. All those books consumed by Project Panama (and similar secret book-destroying dataset-builders) could be resold in batches (at a loss of course) to the same high-volume used book companies that provided them. And Anthropic could license its Project Panama dataset to other AI companies, eliminating the need for competing book-destroyers. Yes, that would cost more money than pulping the books, but pulping and disposing of the pulp and other remnants isn’t cheap either; license fees would go a long way toward paying for the more expensive option of re-homing those books. More than selling the paper pulp to paper manufacturers, at least.
Does this sound unreasonable, Anthropic? Cost-prohibitive? I don’t understand the scale of Project Panama and the enormous number of books they would need to store and the costs of handling & delivery, not to mention the costs of advertising? Why it would be like setting up a rival of Amazon! And who am I to demand that Anthropic share its unique proprietary training model, of unmatched size, quality, and depth?
Points well made. And I must admit that some folks would love to get rid of even one copy of their grandfather’s clever economic treatise, the one he vanity-published back in 1996, boxes of which have been gathering dust in his garage since his death. Small publishers, agents, editors, and used bookstores around the world would be delighted to unload some of their inventory knowing those books will meet their final fate. Fine, destroy the common books (just one of each, remember) that don’t hold their value and are being remaindered anyway.
Just preserve and re-sell, or simply return, the valuable out-of-print books. Especially the rare ones. Don’t pretend they are just as disposable as the commonplace books.
Yes, people will still be upset by your criteria for which ones to save and which ones to pulp, no matter what your criteria is. But far fewer people will be upset by it. And more to the point, valuable and rare books will survive the scanning process. No one will be able to accuse you of building a giant castle of a training model and then digging a moat around it to keep anyone else from doing it (a moat filled with the gooey remains of millions of rare, even one-of-a-kind, books!).
It’s okay, Anthropic, to settle for a slimmer profit margin if you can achieve your business goals in a less-destructive, less-horrifying way. You really don’t need to destroy all those books to build your training model. You can do better. As we’ve seen in ISBNdb’s response to all the bad press about Project Panama (see footnotes), if maximizing every last bit of profit is your only measure of virtue, negative publicity will eventually take a big bite out of those profits.
And if you go too far, no moat of pulp will protect your castle from the mobs with torches and pitchforks.
If you’ll only read one article on this debate, the US Copyright Office has some good thoughts that aren’t “absolutist” either pro or con: https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
Aaron Schaffer, Will Oremus, and Nitasha Tiku, "Inside an AI start-up's plan to scan and dispose of millions of books," The Washington Post, January 27, 2026. —a brief summary based on this excellent news-breaking article, not a direct quote from it. This is the original article that revealed the existence of Project Panama to the world, and it’s worth reading in its entirety.
When an AI model trains on data generated by AI models, the AI model simply “collapses.” So anything generated by AI is scrupulously scrubbed from the training data. If any AI content slips in, the dataset is considered “contaminated.” Sounds like clumsy and undisclosed use of AI in recently-published books… https://www.404media.co/authors-are-accidentally-leaving-ai-prompts-in-their-novels/
Hmm, that page seems to have been taken down. Perhaps this other page on the same site explains why: https://isbndb.com/news
Can we at least hope the pulp is being sold to paper mills and recycled?
Court Exhibit #21, as reported in the Guardian article, “Why is Anthropic destroying books?” (https://www.theguardian.com/commentisfree/2026/aug/05/anthropic-ai-destroying-books) Please note that Anthropic’s goal was to destructively scan one copy of all the books in the world. They never set out to destroy all printed books in the world. Just one representative sample of every book printed before 2022.
Yep, that page has been taken down too. This page does indeed explain why: https://isbndb.com/news


