Ai Engineering 4 min read

AI Companies Are Buying Japan's Used Books by the Ton, Then Shredding Them

Japanese used bookstores report surging bulk sales, including a 50-ton shipment to the US for scan-and-destroy processing attributed to Anthropic, per an NTV Japan report. The buyers target philosophy, history, medicine, and law, and Japanese law lets them skip permission.

Japan’s used bookstores are having a strange boom, and the reporting behind it reads like a supply chain investigation. Per an NTV Japan report covered by Tom’s Hardware on September 24, some stores are selling hundreds of books a day through online marketplaces, with certain days seeing five times normal sales. One recorded consignment: 50 tons of Japanese books shipped to the United States for what the report calls destructive scanning, attributed to Anthropic, whose scanning process the books are suspected of feeding before being pulped. A bookstore in western Japan received an inquiry from a foreign company for tens of thousands of volumes. Multiple bulk buyers, all shipping to the same logistics center in Okayama Prefecture, which declined to comment.

The Genre Pattern Gives Away the Purpose

The tell is in what is being bought. Novels and comics, the staples of the used-book trade, are not moving in bulk. The orders target philosophy, history, political history, medicine, law, and niche topics like Edo-period life and culture, which is to say, exactly the long-tail specialized knowledge that is underrepresented in web-scraped training corpora and expensive to license at scale. Physical scanning of obscure printed material is one of the few remaining ways to acquire text that does not exist online, and the destruction step is simply logistics: once scanned, the physical book is dead weight, and pulp is cheaper than storage. This follows earlier reports of similar scan-and-shred schemes in the US and Europe, with tech companies allegedly outsourcing acquisition to middlemen specifically so that no employee asks where the books came from.

The Legal Structure Is the Story Inside the Story

Here is the part that makes Japan the venue: under Japanese copyright law, bulk scanning of copyrighted works for AI training is legal unless it unfairly harms the interests of the copyright holders, and as the report notes, nobody in this pipeline is likely to ask permission from anyone. Compare the litigation environment elsewhere: the Seattle Times and Newsday are suing OpenAI and Microsoft over copyright, and American publishers have spent two years testing whether training is fair use through the courts. Japan’s copyright exception was written to be permissive toward text and data mining, and AI data operations are now running straight through it at industrial scale. A law that contemplated researchers processing a few corpora is absorbing 50-ton consignments, and whether that counts as “unfairly harmful” is a question Japanese courts have not answered.

What Is Actually at Risk

Two things are being consumed, and only one of them regrows. The books themselves are the smaller loss: used copies of specialized academic and historical works exist in finite supply, and once a ton of them is pulped in Ohio, the secondhand copies are gone from the market permanently. Bookstores report fearing exactly this, a short-term boom that creates long-term supply droughts and permanently removes culturally significant works from circulation. The larger consumption is normative: the world’s physical archives are becoming raw material for model training, acquired through intermediaries, processed destructively, in a legal gray zone that no legislature has reviewed. Whatever one thinks about training on books, the acquisition pattern, bulk, anonymous, destructively processed, and structured to avoid asking anyone anything, is worth scrutinizing on process grounds alone.

What to Watch

Three things. First, Anthropic’s response, if any: the company’s name is attached to the most concrete allegation in the report, and every AI lab has learned this month that silence on a data-provenance story does not age well. Second, whether Japanese publishers and the Book Wholesaler Association push a test case on the “unfairly harmful” clause, which would force the first judicial look at bulk scan-and-shred operations anywhere. Third, whether other governments treat physical-data extraction differently from web scraping: a bookstore receipt is a document of provenance in a way that a scraped page never is, and regulators looking for a tractable copyright enforcement target will notice that the paper trail here is unusually legible.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How to Automate Workflows with Claude Code Routines

Learn how to use Claude Code's new routines to schedule tasks, trigger API workflows, and automate GitHub PR reviews on cloud infrastructure.

Ai Engineering

Seattle Times and Newsday Sue OpenAI and Microsoft Over AI Training on Journalism

The Seattle Times and Newsday filed a federal copyright lawsuit against OpenAI and Microsoft on Friday, alleging their journalism was copied to train AI models without permission or compensation.

Ai Engineering

Sony Music and Warner Chappell Sue Anthropic Over Training Data

Sony Music Publishing, Warner Chappell, and other publishers filed a new copyright lawsuit against Anthropic and its co-founders on August 29, alleging systematic torrenting and scraping of copyrighted works to train Claude.

Ai Engineering

Claude Agents Discovered a Novel CRISPR-like Enzyme System

Anthropic announced on September 23 that roughly 950 Claude agents, running 21 hours on 210 million tokens, found a previously uncharacterized bacteriophage enzyme system with CRISPR-like repeat arrays, now partly validated in the lab and published as a pre-print.

Ai Engineering

Claude Opus 5.5 Matches Fable 5.1 While Cutting Running Costs 40%

Anthropic released Claude Opus 5.5 on September 22, a flagship that matches or beats Fable 5.1 on most benchmarks while costing 40% less to run than Opus 5, as the frontier price war that started with Grok 4.7 claims its biggest scalp.