A market is forming around an unusual training-data filter: buy the text on paper, and require an edition old enough to predate the generative-AI boom.
Book-data company ISBNdb is offering bulk acquisition of physical books to AI customers, pitching pre-2022 editions as less exposed to synthetic text and modern data-poisoning tools than today’s web. The buyers are unnamed, and the cutoff reduces contamination risk rather than proving every sentence was written by a person.
A cutoff date becomes a quality feature
404 Media reported on ISBNdb’s claim that printed books offer edited, domain-specific knowledge that web crawling cannot reproduce. The Next Web’s account says the company emphasizes books published before the mass adoption of large language models, reducing the chance that a model trains on synthetic output produced by earlier models.
That does not make every old book accurate, unbiased, useful or necessarily free of machine-generated passages. It does make the edition and acquisition date easier to establish. A physical copy, ISBN and purchase record can supply a more legible chain of custody than a scraped page with an unknown revision history, though ownership of a copy is not the same as ownership of copyright.
The pipeline is industrial, not literary
Scanning at scale can mean cutting off a book’s spine and feeding loose pages through imaging equipment. ISBNdb promises customer confidentiality and acknowledges the reputational problem created when an AI company destroys large quantities of books. The secrecy is commercially understandable and analytically inconvenient: without buyer names, completed-order volumes and selection criteria, outsiders cannot tell how representative these acquisition programs are.
Court records have already shown the scale of one lab’s appetite. In Bartz v. Anthropic, records described the purchase and destructive scanning of millions of print books. A federal district judge separately found that model training was fair use and that replacing each lawfully purchased print copy with an internal digital copy was fair use on the record before him.
The judge did not extend that protection to pirated copies retained in Anthropic’s central library. The 2025 ruling is one U.S. trial-court decision, not a general license for bulk scanning.
Clean provenance is not a complete license
The commercial lesson is broader than books. Training-data buyers are beginning to pay for evidence about origin, date, rights and transformation history — metadata that was treated as overhead when web scale was the only objective. A corpus can be free of synthetic text and still carry copyright, privacy, representativeness or quality risks.
The next premium dataset may therefore look less like a giant file dump and more like an auditable supply chain: acquisition receipts, edition identifiers, scanning records, deduplication decisions and usage constraints. Paper is not inherently cleaner. A dated edition simply provides evidence that can narrow when and how a text entered the supply chain.
Sources: 404 Media; The Next Web; The Times of India; Bartz v. Anthropic fair-use order.