Confirmed: Anthropic Scanned Books. Unconfirmed: Everyone Else Did Too.

Court records establish that Anthropic destructively scanned millions of books under "Project Panama" and paid $1.5 billion to settle author claims, but a wider panic over an industry-wide book-buying spree…

Two different claims are getting flattened into one story, and the distinction matters for anyone tracking how AI labs source training data. What’s independently established, via unsealed court documents reported by The Washington Post and confirmed in Anthropic’s own settlement: the company bought millions of new and used books in 2024, sliced off their spines, scanned the pages, and discarded the originals under an internal program called Project Panama. A federal judge ruled that training on legally purchased, destructively scanned books is fair use, and Anthropic paid $1.5 billion to resolve related author claims, according to The Next Web. That is documented fact, not rumor.

What’s still just an allegation is that this is happening at industry scale, coordinated through a broker. 404 Media reported that ISBNdb, a book-metadata company, had advertised a bulk-purchasing service pitched explicitly at “LLM training needs.” ISBNdb has since taken the pages down and told Fortune the service “was never launched” and was merely “a test of market interest.” No outlet, including 404 Media itself, has produced evidence tying ISBNdb to an actual completed transaction. Meanwhile, unusual bulk orders are real and documented independently — a Dutch antiquarian bookseller received a spreadsheet of 3,001 ISBNs tied to a firm called “2077AI,” per Fortune’s reporting, and a Houston bookseller told The Atlantic that 95 of his last 100 sales went to one buyer — but none of these orders has been publicly linked to ISBNdb, to Anthropic, or to any named AI lab.

The $1.5 billion Anthropic paid to settle author claims proves book destruction can be legal; it proves nothing about who else is doing it.

The rarity question compounds the confusion. Social media framed this as rare and irreplaceable texts being pulped, but Anthropic drew a distinction to Snopes between “rare and valuable” collectibles and merely “harder-to-find” reference and academic titles — and the booksellers quoted by The Atlantic describe buyers snapping up unglamorous 1970s–90s nonfiction, not first editions. The underlying economic logic, laid out by The Next Web, is straightforward: pre-2022 print runs are guaranteed free of AI-generated text and resistant to data-poisoning attacks, making them valuable as clean training data regardless of literary merit. Snopes rated the broader rumor “mostly true but with undetermined elements,” which is the accurate state of play.

What would move this from rumor to reporting: a buyer’s name attached to the Dutch or Houston orders, verified purchase records from a lab other than Anthropic, or an on-record confirmation from ISBNdb of a completed client engagement rather than a shelved pitch page. Until one of those surfaces, the responsible framing is that one company’s practice is proven and everything else is inference from spreadsheet metadata and anonymous buyers.

AI Companies Bulk-Buy Books, Scan and Destroy for Training Data

조선일보

Read the full story at 조선일보 →

The Data Commenter, in your inbox

Data markets, alt data, and the AI training-data economy. No spam, unsubscribe anytime.

Discussion lives in the inline notes attached to article passages.