How Destructive Book Scanning Turns Print Volumes Into AI Training Data
A printed book can become machine-readable data in two very different ways. A library may preserve the binding and photograph each page carefully. An industrial scanning operation may do the opposite: cut off the spine, feed loose pages through high-speed equipment, run optical character recognition, and discard the physical copy.
That second method is known as destructive scanning. It is faster and cheaper because a bound volume does not need to be held open or photographed page by page. The trade-off is permanent. Once the binding is removed, the original object can no longer function as a book, and any historical or material information carried by the binding may be lost.
Why AI developers want physical books
Large language models learn patterns from enormous collections of text. Books are attractive training material because they usually offer longer, more coherent writing than short web pages, along with specialized vocabulary, narrative structure, and sustained arguments.
A company seeking a particular title may buy a used or out-of-print copy, digitize it, and add the resulting text to a searchable archive or training collection. Commercial book databases and logistics firms can help locate titles by ISBN, author, edition, or language. The available evidence does not establish that every book acquired through these channels is rare, or that every shipment is destined for one particular AI model.
The Amazon connection has drawn attention because reporting traced at least one large shipment of books to an Amazon facility in Nevada identified as VGT3. The reporting described workers removing bindings and scanning books, but it did not publicly establish which AI product or outside customer would use the resulting data. That distinction matters: a warehouse’s role in receiving or processing books is not, by itself, proof of ownership of the training dataset.
The legal question is separate from the physical destruction
U.S. copyright law evaluates copying under the four-factor fair-use test: purpose, nature of the work, amount copied, and effect on the market. The U.S. Copyright Office says only a federal court can make the final determination in a particular dispute.
In Bartz v. Anthropic, a federal judge ruled in June 2025 that certain copying of lawfully purchased books for a central digital library and AI training qualified as fair use on the record before the court. The same ruling found that obtaining millions of pirated books and retaining them in a central library was not fair use. The decision shows why “the books were bought” and “the books were copied” do not settle every copyright issue.
What may be lost when a book is cut apart
The ethical concern goes beyond copyright. A scarce edition can contain evidence about paper, typography, annotations, provenance, repairs, and previous ownership. A digital text may preserve the words while erasing the artifact.
Libraries and archives normally scan rare materials non-destructively, even when the process costs more. The choice between those methods reflects a broader question: whether books are merely containers of text, or cultural objects whose physical survival has value of its own.
Sources
- We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
- Bartz v. Anthropic, Order on Fair Use
- U.S. Copyright Office: More Information on Fair Use
- U.S. Copyright Office: Copyright and Artificial Intelligence, Part 3
- Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement