How AI is giving forgotten books a new lease of life

A wave of bulk purchases is sweeping through second-hand bookshops, but the lack of transparency is raising eyebrows

Pooja Sreenivasan

How AI is giving forgotten books a new lease of life

In a modest second-hand bookshop in Badalona, near Barcelona, Marcel Font began to notice something he had never seen before. A Canadian company had begun buying books that had languished on shelves for years, many of them non-fiction Catalan-language works. The orders sometimes came only minutes apart.

Hundreds of kilometres away, the German bookseller Michael Strötter encountered a strikingly similar pattern. At 2:53 a.m. on 30 April 2026, the first order arrived. More followed throughout the night. In almost every case, the buyer was the same company: Zoom Books, a Canadian firm that describes its business as selling and reusing second-hand books, while recycling those it cannot sell.

The books being sought were neither rare first editions nor costly manuscripts. They were cookbooks, novels, biographies, and ageing specialist titles. Similar accounts in the UK, Australia, New Zealand, and elsewhere emerged, all describing the same recurring pattern: large orders, at times seemingly automated, spanning an extraordinary range of titles and often requesting only a single copy.

These companies have denied buying the books for AI training purposes, but their refusal to reveal the identity of some clients, citing confidentiality agreements, has fuelled suspicion of ulterior motives.

In a copyright lawsuit against Anthropic, the developer of Claude, court documents revealed the company had purchased millions of printed books and then hired contractors to remove their covers, cut away their pages, scan them, and convert them into digital files before disposing of the physical copies.

The undertaking, known as ‘Project Panama’, frequently relied on ‘destructive scanning’, a process in which a book is dismantled so that its pages can be digitised more rapidly. The concerns voiced by booksellers therefore have a known precedent: acquiring books that are difficult to find in digital form and converting their contents into data that can be used to train language models.

The legal question is considerably more complex than the formula ‘buy the book, then destroy it’. Under US law, purchasing a physical copy and destroying it does not, by itself, make its use in AI training lawful.

REUTERS
Visitors gather at the Anthropic booth at the Moscone Centre during the Dreamforce 2026 tech summit in San Francisco, California

In the Anthropic case, federal judge William Alsup ruled in June 2025 that converting lawfully purchased books into internal digital copies could fall within fair use. His reasoning rested partly on replacing each physical copy with a single digital copy that was not distributed to the public. He also found that using the books to train the model qualified as fair use in the circumstances of the case.

The court nevertheless distinguished those books from the millions of copies Anthropic had obtained from piracy sites. Their unlawful acquisition did not become legitimate simply because the material was subsequently used for training.

Destroying the physical copy was therefore one factor in the court’s legal analysis, rather than a general rule. Nor does the $1.5bn settlement approved in July this year over pirated books confer on AI companies an unrestricted right to use any book merely because they have purchased a copy.

The US Copyright Office has cautioned against reducing ‘fair use’ to a single rule applicable to every instance of AI training. The assessment depends on the nature of the works involved, the manner in which they were obtained, the purpose for which they are used, and the effect of the resulting outputs on the original market

As high-quality human writing is becoming a scarcer resource, tech companies are returning to printed books

Scarcity of quality writing

As high-quality human writing is becoming a scarcer resource, tech companies are returning to printed books. Large language models have consumed enormous quantities of online content in recent years, and as those models have grown, finding fresh, varied, and non-duplicative text has become increasingly difficult.

Researchers at Epoch AI estimate that the effective stock of publicly available text written by humans may amount to roughly 300 trillion tokens (the small pieces of text that AI models process), and that models could consume a substantial share of it between 2026 and 2032. The estimates vary, but one thing is for certain: high-quality human text is becoming harder and more expensive to obtain.

Herein lies the value of older books. A specialist volume that has never been digitised may not appeal to the average reader. But for a company building a language model, it may offer tens of thousands of carefully structured and edited human words, covering fields or periods poorly represented on the modern internet.

These books become more valuable still because many were written before the age of AI-generated content, at a time when the web itself is steadily filling with material produced by models. Old books have consequently become an appealing source of text whose human origin is far more trustworthy.

REUTERS
A man browses books at a used-book market in Lima, Peru

How it works

According to an investigation by The Atlantic, in March ISBNdb offered AI developers the opportunity to source printed books in quantities of up to one million titles per order, while keeping the identity of the client confidential. The company later said the proposal had merely been a test of demand and that the service was never launched. It also denied purchasing, scanning, or destroying books to train AI models.

The proposal nevertheless offers a revealing glimpse of how such a market might operate. An algorithm could use ISBN numbers to compare millions of titles with those already contained in a dataset, then purchase a copy of every title that is missing. This may help explain the peculiar assortment of orders being reported by booksellers.

This does not mean that buyers are hunting for rare treasures. Many of the books being sought are ordinary, ageing volumes, some of which have remained unsold for years. Yet limited commercial value does not mean limited cultural value. The second-hand book trade helps keep out-of-print works within reach of readers and researchers, and by taking substantial numbers of copies out of circulation, it makes them harder to find and more valuable.

Booksellers do not necessarily have an issue with the technology itself, but rather the lack of transparency surrounding the process

Pros and cons

For their part, booksellers have more or less welcomed the opportunity to clear long-dormant stock. Digitisation may preserve the contents of books that no one is willing to buy anymore. In some cases, digitisation may even save texts that might otherwise disappear as physical copies deteriorate or are discarded. Taken together, AI is essentially giving forgotten books a new lease of life.

Booksellers do not necessarily have an issue with the technology itself, but rather the lack of transparency surrounding the process. Knowledge once scattered across bookshops and private homes may get sucked into a proprietary database. Eventually, books could disappear altogether from readers' hands and exist solely in the AI sphere.

That said, it may still be too early to jump to conclusions and sound the alarm. The scale of risk largely depends on which books are acquired, how many copies survive, what happens to the digital version, who owns it, and who is permitted to access it.

font change

Related Articles