UK and Ireland secondhand booksellers suspect AI firms behind 'strange' bulk orders

Published 2026-08-15 · AI Daily — AI-assisted deep research, methodology & disclosure

Following reports that Anthropic spent millions on book scanning for data acquisition, secondhand booksellers in the UK and Ireland have reported a surge in bulk orders from mysterious buyers, fueling speculation that AI companies are mass-acquiring books for training data.

Background and Context

In recent months, a concerning pattern has emerged among secondhand booksellers across the United Kingdom and Ireland. Independent bookstores and major online platforms have reported a surge in unusual bulk orders from anonymous buyers. Unlike typical purchases made by individual collectors or libraries, these orders are characterized by their scale and specificity. Buyers frequently request dozens or even hundreds of books from specific academic disciplines, particular decades, or niche authors. The demand is particularly high for out-of-print titles, specialized journals, and obscure academic works that are difficult to find in public digital libraries.

This trend has intensified following media reports that Anthropic spent millions of dollars on book scanning for data acquisition. The timing of these bulk purchases coincides with the release of new model versions by major AI laboratories, suggesting a direct link between the acquisition of physical texts and the development of new AI capabilities. The buyers often use anonymous accounts or temporary email addresses and refuse to disclose their purpose. Furthermore, shipping addresses show a concentration trend, pointing toward data annotation centers or scanning studios. This organized behavior indicates that AI companies are actively seeking to overcome training data bottlenecks by acquiring physical assets rather than relying solely on digital sources.

The phenomenon represents a shift in how AI firms approach data collection. While the internet provides vast amounts of text, much of it is low-quality, repetitive, or erroneous. Physical books, especially those that have undergone peer review or are established classics, offer higher information density and accuracy. By targeting the secondhand market, AI companies are bypassing the legal barriers associated with scraping copyrighted digital content. This strategy allows them to access high-value data while navigating the complex legal landscape of digital rights management.

Deep Analysis

The underlying logic of this bulk acquisition reflects a critical pivot in large language model training. The industry is moving from a competition of breadth, relying on the sheer volume of public internet data, to a competition of depth, seeking high-value, low-noise datasets. As model scales expand, the marginal utility of public web data is diminishing rapidly. Consequently, AI developers are turning to physical books to enhance reasoning capabilities and reduce hallucinations. These texts provide the structured and accurate information necessary for advanced model performance.

Acquiring books through the secondhand market offers a cost-effective and legally ambiguous pathway. Directly purchasing digital copyrights is expensive and subject to strict licensing agreements. In contrast, buying legally circulated physical books does not inherently violate copyright law, even if the subsequent scanning and usage involve complex legal questions. This approach allows AI companies to build a proprietary data moat. By securing exclusive access to specific corpora, they can prevent competitors from using the same training materials, thereby establishing a short-term performance advantage. The focus of competition is shifting from algorithmic architecture to the control of high-quality textual resources.

This strategy also exploits the relative leniency of the secondhand book market regarding data compliance. Booksellers are generally focused on the sale of physical goods and may not scrutinize the end-use of their inventory. This gap in oversight allows AI firms to operate in a gray area, where the act of purchase is legal, but the application of the data may be contentious. The result is a new form of data acquisition that relies on the physical supply chain rather than digital infrastructure.

Industry Impact

This trend has profound implications for the traditional publishing industry and authors. When AI companies acquire books through the secondhand market, they are leveraging existing cultural assets to generate new commercial value without compensating copyright holders. This form of value extraction undermines the incentive for creators and weakens the innovation dynamics of the publishing sector. Authors and publishers are effectively being bypassed in the monetization of their intellectual property, raising serious ethical and economic concerns.

The secondhand book market itself is facing structural disruption. Small and medium-sized booksellers, who traditionally serve collectors and readers, are being forced into the AI data supply chain. This shift alters their customer base and forces a reevaluation of inventory strategies. Titles previously considered slow-moving are now in high demand, driving up prices and making them less accessible to general readers. Major online platforms such as AbeBooks and ThriftBooks are also under increasing compliance pressure. Regulators may require these platforms to monitor bulk transactions more closely, potentially restricting the resale or export of certain book categories. This would increase operational costs and legal risks for these intermediaries.

For consumers, the potential monopolization of high-quality training data by a few AI companies could lead to homogenized AI-generated content. If a small number of firms control the best available datasets, their models will dominate in terms of knowledge accuracy and style. This concentration of data power exacerbates the digital divide, trapping users in a data loop where they are exposed to limited perspectives and information sources. The diversity of human knowledge, as reflected in the physical book market, risks being flattened into a narrow set of AI outputs.

Outlook

This situation may serve as a turning point for AI data ethics and regulation. Current legal frameworks, such as the EU AI Act and US copyright lawsuits, primarily focus on the scraping of public data. However, the legal boundaries for data acquired through physical transactions remain unclear. In the coming months, copyright collective management organizations, such as PRS for Music in the UK or ASCAP in the US, may intervene. They could demand that AI companies disclose their data sources and push for the establishment of data usage transparency mechanisms.

The secondhand bookseller association is likely to respond by developing industry self-regulation guidelines. These may include refusing service for bulk orders with undisclosed purposes or requiring buyers to sign data usage commitments. For AI companies, the risks of the current gray-market strategy are increasing. They may need to shift toward formal data licensing agreements with publishers, paying reasonable royalties in exchange for legal access to high-quality data. This transition from gray acquisition to white licensing will be a key indicator of the industry's maturity.

Key signals to watch include whether major AI laboratories begin to publish the composition of their training data, whether copyright organizations initiate new collective lawsuits, and whether secondhand book platforms introduce restrictions on bulk purchases. These developments will determine the future path of AI data acquisition and how traditional cultural assets are respected and compensated in the digital age. The resolution of this issue will require a balance between technological innovation and the protection of intellectual property rights.

Sources