Somewhere between a dusty poetry shelf and a warehouse loading bay, AI training data has become a procurement problem. The Guardian reports that secondhand booksellers in the UK and Ireland are seeing unusual bulk orders from mystery buyers, with speculation that AI companies are acquiring books for data. The same report notes this follows Anthropic having spent millions on books to scan for data acquisition, which is the kind of sentence that makes web scraping sound quaint, like dial-up with a tote bag. This is not proof that every suspicious cart full of paperbacks is secretly Claude wearing a trench coat. It is, however, a useful signal: high-quality text is scarce enough that the industry may be treating physical books like strategic inputs. Everyone loves talking about parameter counts. Nobody wants to talk about the supply chain for nouns. ## The Guardian found a procurement pattern hiding in plain sight The Guardian says UK and Ireland bookshops have received orders from buyers in the US, Canada, continental Europe, and the UK, while sellers describe the pattern as strange enough to raise questions. Its report centers on speculation among booksellers, not a confirmed buyer list, which matters because attribution is the difference between reporting and cosplay detective work. Still, the alleged workflow is legible: acquire books, digitize text, and feed cleaner material into model training. That is not magic. That is ETL with paper cuts. The Guardian also highlights Stuart and Mary Manley of Barter Books in Alnwick, Northumberland, where the captioned theory is blunt: ‘My theory is that it’s masses and masses of AI money’. That is not a lab benchmark, but it is market behavior from people who know what normal book buying looks like. In AI, the weirdest telemetry often comes from the humans nearest the bottleneck. ## The BBC shows why sellers noticed The BBC reports that Stuart Manley of Barter Books would typically sell two or three thousand books in a week, but one single bulk order from a Canadian company equalled what he would expect to move in seven days. Manley told the BBC he had "never seen the like of this after 30 years in the second-hand book trade". That is the kind of anomaly detection you do without a dashboard, just a till, a van, and the creeping sense that someone trained a model on your inventory spreadsheet. The BBC also says booksellers do not know the final destination for the books, while suspicion has turned toward AI because models have a voracious appetite for new information. The important word is suspicion. Builders should resist turning every unknown buyer into a villainous GPU cluster, but they should pay attention to why books are appealing: edited prose, long-form structure, and less of the synthetic mush now floating around the open web like protein powder in a jacuzzi. ## Chosun connects the rumor to scanning economics Chosun reported that a request arrived last May at Kenneth’s Bookshop in Galway, Ireland, to purchase 5,000 books at once. The outlet framed the broader trend as AI companies buying large quantities of used books for scanning to secure original text materials for LLM training. It also referenced Anthropic launching Project Panama, though the snippet does not provide enough detail to responsibly turn that into a spy novel. Annoying, yes, but accuracy is the only seasoning we are allowed to use here. The scanning angle matters because it shifts data acquisition from passive collection to active procurement. Instead of hoovering public web pages and hoping the tokenizer can survive the comment section, a company can buy physical copies, digitize them, and build a corpus with known inputs. That does not automatically solve copyright, licensing, or author compensation. It does make the pipeline more auditable than the old internet buffet, where provenance was often a shrug wearing an API key. ## Why builders should treat data like inventory The BBC reports that in 2025, a US judge ruled that using books purchased in this way to train AI software was not a violation of US copyright law. That does not mean the same answer applies everywhere, or that every procurement path is reputationally clean. It means data strategy is now legal strategy, vendor strategy, and operations strategy. Congratulations, your model card needs a supply chain tab. For product teams, the lesson is boring in the best possible way: document where data comes from, what rights attach to it, and what happens after ingestion. For researchers, the interesting question is whether higher provenance corpora become a competitive moat as public web data gets noisier and more synthetic. For publishers and booksellers, this could create new licensing markets if everyone resists the urge to solve culture with a shredder. Watch the next moves around disclosed data partnerships, procurement audits, and whether AI labs start treating text sources with the same seriousness they give chips and cloud contracts. If the model race is now reaching into secondhand shelves, the limiting reagent may not be compute. It may be clean sentences, properly sourced, and preferably not pulped after a robot has had a snack. ## Sources - Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders
- AI Companies Bulk-Buy Books, Scan and Destroy for ...
- Secondhand book sales are booming. Is it because of AI?
Sources
- Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders
- AI Companies Bulk-Buy Books, Scan and Destroy for ...
- The Guardian - The spate of orders has fuelled speculation...
- Secondhand book sales are booming. Is it because of AI?
- Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders
- AI Companies Bulk-Buy Books, Scan and Destroy for ...
- Secondhand book sales are booming. Is it because of AI?
- What's Really Going On With Bulk Purchase Of Used and ...