Shred to Train: Inside the Legal Loophole Turning Rare Books Into AI Fuel

Column Overview
There is a strange kind of quiet at the center of this story. No protests, no headlines that stick, no viral outrage cycle — just trucks, warehouses, and machines that cut spines off books faster than any human could read them. Somewhere in that quiet, a business model has formed around buying up rare books, scanning their pages, and destroying what's left. A federal judge has already said it's legal. The company doing the buying isn't hiding, exactly, but it isn't advertising either. And the reason any of us know about it at all is that a secondhand bookseller finally said something to a reporter.
The Mechanics of a Very Efficient Erasure
The process itself is almost clinical. High-speed scanning rigs are built to process books by destroying their bindings — the spine gets cut away so pages feed through like loose paper, which is dramatically faster than photographing a book page by page while keeping it intact. What's left afterward isn't a book anymore. It's loose leaves, pulped or discarded, because putting it back together was never the plan.
A service called ISBNdb sits in the middle of this, brokering bulk orders — reportedly up to a million books at a time — while keeping the identity of buyers anonymous. That anonymity is doing a lot of work here. It means a library, a rare bookshop, or a collector selling off a lot of eighteenth-century agricultural texts has no idea their inventory is headed for a shredder rather than a shelf. It also means the public has no clean way to track how many irreplaceable volumes have already gone through the pipeline.
Why the Destruction Is the Point, Not a Side Effect
Here's the part that took me a second read to fully absorb: the destruction isn't incidental to the legal argument, it's the argument. A federal judge ruled this practice falls under fair use specifically because eliminating the original ensures that only one copy of the work exists at any given time. In other words, you can't be sued for effectively duplicating a copyrighted book if you make sure the duplicate and the original never coexist. Scan it, shred it, and legally speaking, nothing was copied — something was merely converted from paper to data.
It's a clever reading of copyright law, and it's easy to see why buyers would want books published before roughly 2022. Anything older is guaranteed to be free of AI-generated text, which matters enormously if you're trying to train a model on authentic human writing rather than accidentally feeding it a loop of machine-generated prose scraped from more recent print runs. Pre-2022 books, in this framing, aren't just historical artifacts. They're clean data.
A Business Built on Not Being a Headline
What makes this different from the last few years of AI training controversies — scraping websites, torrenting shadow libraries, pulling audio off streaming platforms — is that ISBNdb's own marketing seems to understand exactly how bad this looks. Reporting from 404 Media surfaced language on the company's site acknowledging, more or less verbatim, that "AI company destroys two million books" is not a headline that generates sympathy. And yet the business exists anyway, built around NDAs as a standard feature and a preferred vocabulary where "shredding" becomes "digital preservation."
That phrase, digital preservation, is doing something almost impressive in its dishonesty. Preservation usually implies keeping something safe for the future. Here it means the opposite: the physical object is deliberately destroyed, and what survives is a digitized copy that exists purely as training material, likely never to be read by a human being again in any meaningful sense. A bookseller who spoke to 404 Media described feeding in volumes with almost no surviving copies elsewhere — books that had already made it through wars, fires, floods, and a few hundred years of ordinary neglect, only to meet their end in a scanning rig built for speed rather than care.
Scraping Is Reversible. This Isn't.
I've read plenty of stories about AI companies scraping open web pages or pulling from pirated book collections, and there's always been an implicit "well, at least it's still out there somewhere" comfort to it. A website can be re-uploaded. A bestseller can be reprinted indefinitely. Even a torrented archive, ethically messy as it is, doesn't erase the original text from existence — someone, somewhere, still has a copy.
Rare books don't have that safety net. If a particular edition of a botanical survey or a regional history has three surviving copies in the world, and one of those copies gets fed into a scanner and shredded, the number of surviving copies just became two, permanently. There's no reprint, no backup server, no second source. This is the first AI training controversy I've come across where the raw material itself is the casualty, not just the compensation model around it.
It also puts a new kind of pressure on the broader industry conversation about acquiring training data at scale — companies including Anthropic have reportedly brought on senior talent from major digitization efforts, including a former head of partnerships at Google Books, specifically to pursue comprehensive book access for training purposes. The ambition of getting "all the books in the world" into a usable digital form isn't new; Google itself spent close to two decades scanning millions of volumes. What's new is doing it in a way that doesn't leave the physical original behind.
Legal, Efficient, and Worth Sitting With
None of this requires a villain. Nobody is breaking a law. A court looked at the practice and found a reasonable, if unsettling, basis to call it fair use. ISBNdb isn't hiding that it exists. The buyers presumably see themselves as solving a genuine data problem — clean, pre-AI text is a real scarcity, and training models well is a real technical challenge.
But legality and comfort aren't the same thing, and it's worth being honest about which one we're actually discussing. A ruling that treats irreversible destruction as the condition that makes copying acceptable is going to shape behavior for as long as it stands, and the behavior it's shaping is bulk acquisition of physical scarcity — books that can't be replaced once they're gone. That's a different category of loss than a scraped webpage or an out-of-print paperback, and it's the kind of thing that tends to be understood only in hindsight, once specific texts are gone and someone goes looking for a copy that no longer exists.