Secondhand booksellers are reporting an unusual surge in large orders, with thousands of books being purchased and shipped to warehouses around the world. The unexpected buying spree has prompted speculation that artificial intelligence companies may be behind at least some of the demand.
Independent booksellers who would normally sell a few thousand books over an entire week have reported receiving individual orders containing unusually large numbers of titles.
At Barter Books in Northumberland, owner Stuart Manley said one recent order from a Canadian company was roughly equivalent to an entire week’s normal sales.
After three decades working in the secondhand book industry, Manley said the scale and nature of the purchases were unlike anything he had previously encountered.
Similar reports have emerged from other booksellers, but exactly where many of these books ultimately end up remains unclear.
One theory attracting attention within the industry is that physical books are being purchased to provide additional training material for artificial intelligence systems.
Modern generative AI systems require enormous quantities of information during development. Books can be particularly valuable because they contain carefully edited writing, specialist knowledge and material that may not be readily available on the public internet.
The speculation has intensified following legal developments in the United States concerning the use of books for AI training.
In 2025, a US judge ruled in a copyright case involving AI company Anthropic that certain uses of legally acquired books for training artificial intelligence could qualify as transformative under US copyright law.
Court documents connected to the case later revealed details about how large quantities of physical books had been acquired and digitised for AI development.
One method involves what is known as “destructive scanning”.
Instead of manually scanning a book page by page while keeping it intact, the spine can be removed so individual pages can be rapidly fed through industrial scanning equipment. Once digitisation is completed, the remaining physical material may be recycled.
The process makes it possible to convert enormous collections of printed material into machine-readable datasets far more quickly.
However, the practice has created unease among some booksellers.
For many dealers, selling books that have remained unsold for years is commercially attractive. At the same time, the possibility that books could immediately be dismantled and destroyed after purchase conflicts with the traditional role of booksellers in preserving printed material.
The issue becomes particularly complicated when rare books are involved.
Not every old or difficult-to-find book is historically significant. Millions of copies may exist of some popular titles, making the loss of an individual copy relatively insignificant.
But specialist booksellers sometimes handle publications where only a handful of copies are known to survive. Destroying one of those books for digitisation could permanently reduce the number of physical copies available to collectors, researchers and libraries.
The unusual range of books reportedly being purchased has added to speculation about AI involvement.
Orders can include everything from specialist academic publications and obscure historical texts to fiction and older popular literature, with little obvious connection between the subjects.
For AI developers, however, that diversity could be useful.
Large language models benefit from exposure to different writing styles, subjects, historical periods and specialist knowledge. Rare material that has never been widely digitised could therefore provide training data that cannot easily be collected from websites.
The situation also highlights differences between copyright rules in different countries.
Legal decisions in the United States have shaped how AI companies can potentially use legally obtained material for model development, while copyright frameworks elsewhere, including the UK, may impose different requirements concerning copying and the permission of rights holders.
These differences are becoming increasingly important as AI companies search globally for high-quality training data.
The rapid development of generative AI has already increased demand for enormous collections of text, images, video and audio. As easily accessible internet data becomes increasingly exhausted or duplicated across training datasets, companies may have greater incentives to seek information from previously untapped sources.
Physical books represent one such source.
For secondhand booksellers, the mysterious orders therefore present both an opportunity and a dilemma.
Books that have remained on shelves or online catalogues for decades may suddenly have buyers, generating unexpected revenue for independent businesses.
But if some of those books are being purchased primarily to be cut apart, scanned and recycled, the AI boom could also create new questions about preservation, copyright and the value society places on physical knowledge.
Whether artificial intelligence companies are responsible for the wider surge in secondhand book purchases remains uncertain.
What is clear is that the AI industry’s appetite for high-quality information is changing the value of data — and even books that struggled to find a reader for decades may suddenly have a new kind of customer.







