UK and Ireland Secondhand Booksellers Suspect AI Firms Behind 'Strange' Bulk Orders

Published · AI Daily — AI-assisted deep research, methodology & disclosure

Following reports that Anthropic spent millions scanning books for data acquisition, secondhand booksellers in the UK and Ireland are reporting a surge in bulk orders from mysterious buyers, with speculation that AI companies are purchasing the tomes for training data.

Background and Context

The secondhand book trade in the United Kingdom and Ireland is currently experiencing a significant shift in demand patterns that has alarmed industry veterans. Since the first half of the year, numerous independent booksellers and library clearance departments have reported a surge in bulk purchase requests from anonymous buyers. These transactions are characterized by their scale, with buyers acquiring hundreds or even thousands of volumes in single orders. Unlike traditional collectors or institutional archivists, these purchasers exhibit a distinct lack of concern for the physical condition of the books. They frequently accept copies with yellowed pages, damaged spines, or other cosmetic flaws, provided the textual content remains intact and legible. This indifference to the material object in favor of the digital text is a hallmark of the current trend.

The opacity surrounding these buyers has fueled widespread speculation within the trade. The purchasers consistently refuse to disclose their institutional affiliations or the specific purpose of their acquisitions, often citing vague justifications such as "private research" or "archival preservation." This behavior has coincided with recent reports that Anthropic spent millions of dollars scanning books to acquire training data for its large language models. The timing of these bulk purchases aligns closely with the release of new models or updates to data strategies by major AI laboratories, leading many in the industry to suspect that these "strange" orders are part of a coordinated effort by tech giants to secretly expand their proprietary corpora.

For secondhand booksellers who rely on long-tail inventory for their survival, this sudden influx of capital has disrupted established operational rhythms. While the immediate effect has been a boost in cash flow, it has also led to the rapid depletion of scarce titles. Ordinary readers and collectors are finding it increasingly difficult to acquire specific books, as they are being absorbed by these anonymous entities. This dynamic marks a significant intrusion of AI data acquisition strategies into the traditional publishing supply chain, moving beyond digital scraping to physical market manipulation.

Deep Analysis

The underlying driver of this phenomenon is the AI industry's insatiable demand for high-quality, long-tail textual data. In the current data landscape, publicly available internet text has been extensively cleaned, deduplicated, and exhausted by major AI companies, leading to diminishing marginal returns. To break through performance bottlenecks, these firms are seeking data sources that offer greater diversity, professionalism, and depth. The secondhand book market fills this gap effectively. Unlike web text, books have undergone editorial review, possess rigorous logical structures, and contain specialized knowledge in fields such as historical archives, local gazetteers, professional journals, and out-of-print literature that have not been widely digitized or crawled.

From a strategic perspective, purchasing physical books and using Optical Character Recognition (OCR) to scan them presents a cost-effective alternative to complex copyright licensing negotiations. This approach allows AI companies to bypass lengthy legal processes and high licensing fees, leveraging the price disparity in the secondhand market to acquire massive amounts of text at a fraction of the cost of formal authorization. This "data arbitrage" exploits information asymmetry and market segmentation to quickly build a private data moat. However, this method is not without technical hurdles. Scan quality is limited by print clarity, paper material, and binding methods, requiring significant engineering resources for text cleaning, denoising, and structuring. Ensuring accuracy and avoiding OCR errors that introduce noise remains a critical technical challenge for these firms.

Despite these challenges, the legal and temporal advantages of direct acquisition are compelling. In many jurisdictions, the purchase of a physical book and the subsequent scanning of its content do not directly constitute copyright infringement, particularly given the legal ambiguities surrounding "fair use" and the "first sale doctrine." This legal gray area allows AI companies to operate with lower legal risk compared to negotiating licenses with publishers. The result is a shift in the value proposition of books, where their worth is increasingly determined by the scarcity and completeness of their textual data rather than their literary or academic merit alone.

Industry Impact

The entry of AI companies as primary buyers is distorting the market pricing mechanisms for books. A "data premium" is emerging, where the cost of specific titles is artificially inflated based on their utility as training data rather than their cultural value. This price distortion creates a barrier for ordinary readers and collectors, who face higher costs or complete unavailability of certain works. The traditional equilibrium of the secondhand market, which balanced supply and demand for cultural consumption, is being disrupted by a new class of industrial buyers whose primary interest is data extraction.

This trend has intensified anxiety within the publishing industry. Traditional publishers rely on copyright licensing and royalty income to sustain their operations. By acquiring secondhand books to bypass authorization, AI companies are effectively circumventing the copyright system. While legally defensible in some contexts, this practice raises significant ethical and commercial concerns. The publishing sector argues that this approach deprives authors and publishers of their rightful share of revenue from data usage, potentially undermining creative incentives and damaging the diversity of the cultural ecosystem in the long term.

Secondhand booksellers, acting as intermediaries, face a complex identity crisis. Historically positioned as custodians of cultural heritage and knowledge sharing, they now find themselves integrated into the AI data supply chain. This shift impacts their social reputation and forces a reevaluation of their business models. Some booksellers have begun to actively screen orders, refusing to sell sensitive or high-value books to suspected AI buyers to maintain industry ethics. However, this resistance is fragile, particularly for small independent stores, where refusing large orders could lead to existential financial crises. The economic pressure to sell often outweighs ethical considerations, leading to a fragmented response across the trade.

Outlook

The AI data acquisition war is transitioning from "public data crawling" to "private data acquisition." This shift implies that data barriers will become increasingly high, with companies possessing more high-quality private data gaining a competitive advantage in model development. For smaller AI startups, insufficient financial resources may prevent them from participating in this large-scale data acquisition race, potentially leading to further industry consolidation and a strengthened head effect. This dynamic is pushing the publishing and AI industries to explore new cooperation models. Some publishers are beginning to negotiate direct data usage licenses with AI companies, aiming to establish transparent revenue-sharing mechanisms. If this model becomes widespread, it could reshape the publishing industry's revenue structure, shifting it from a reliance on sales to a dependence on data services.

However, the complexity of negotiations and the fairness of revenue distribution remain significant obstacles. Looking ahead, this trend may trigger stricter regulatory intervention. Governments in various countries may begin to scrutinize the legality of AI data acquisition, potentially enacting specific regulations on large-scale text scanning and data usage. These regulations could include requirements for AI companies to disclose data sources or set compensation standards for specific types of data usage. The industry is watching for key signals, such as whether AI companies will publicly acknowledge their data acquisition practices, whether the publishing sector will form a unified response or collective litigation, and whether the secondhand book market will develop specialized trading channels for AI data needs.

The role of books as knowledge carriers is being redefined. They are no longer just objects for reading but have become fuel for algorithmic training. This transformation involves deep adjustments in technology, culture, ethics, and economic structure. The relationship between AI and the publishing industry will likely be determined by these upcoming developments, deciding whether the two sectors move toward confrontation or symbiosis. The ongoing tension highlights the need for continuous societal attention and reflection on the implications of this shift.

Sources

FAQ

What has been happening at secondhand bookshops in the UK and Ireland?

Since early this year, anonymous buyers have placed bulk orders of hundreds to thousands of books, accepting damaged copies as long as the text is intact.

Why do these bulk purchases matter for the publishing industry?

The buys inflate prices for niche titles, creating a data premium that distorts pricing, while bypassing copyright licensing and sparking ethical debate.

What should the industry watch for next?

Watch whether AI firms admit the purchases, whether publishers unite for legal action, and whether dedicated channels emerge for AI data trading in books.