Ai Engineering 3 min read

Sony Music and Warner Chappell Sue Anthropic Over Training Data

Sony Music Publishing, Warner Chappell, and other publishers filed a new copyright lawsuit against Anthropic and its co-founders on August 29, alleging systematic torrenting and scraping of copyrighted works to train Claude.

On Friday, August 29, 2026, Sony Music Publishing and Warner Chappell filed a copyright lawsuit against Anthropic in the U.S. District Court for the Northern District of California, joined by numerous other music publishers. The complaint alleges a “brazen campaign of illegally torrenting, scraping, and downloading copyrighted works” to train Claude models, and names Anthropic co-founders Dario Amodei and Benjamin Mann as individual defendants. TechCrunch and Music Business Worldwide both characterize the potential damages as reaching into the billions, with statutory damages of up to $150,000 per infringed work on the table.

The Piracy Argument, Refined

The publishers’ legal theory has been sharpened by precedent. In Bartz v. Anthropic, Judge Alsup ruled that training AI models on copyrighted works was transformative and legal, but that acquiring those works through piracy was not. That distinction led directly to Anthropic’s landmark $1.5 billion settlement with authors, approved in July 2026. The same legal team behind the authors’ case, and behind a January 2026 suit from Universal Music Group and Concord, is now applying the framework to song lyrics and sheet music: the allegation is not that Claude learned from copyrighted texts, but that Anthropic obtained millions of books containing them through flagrant torrenting.

That framing matters because it narrows the fight to acquisition methods rather than training itself. The January publisher suit covered around 20,000 works; the new complaint describes a broader pattern, and tech sites report “thousands of copyrighted works” at issue in this filing alone.

Anthropic’s Position

Anthropic’s response so far is a single line from a spokesperson: “We disagree with the publishers’ claims and we intend to defend ourselves robustly in court.” Notably, the company did not settle the January publisher suit before this filing. After paying $1.5 billion to authors in July, the calculus on a second large settlement has evidently changed.

Why Developers Should Care

This case will help define where the “legally obtain your training data” line sits in practice, and the answer propagates into every AI product built on foundation models. Teams fine-tuning on scraped datasets, or building retrieval products over copyrighted corpora, are one court ruling away from inheriting the same exposure. The Bartz precedent suggested a safe harbor for training; this suit tests whether the harbor survives when acquisition looks like piracy at scale. Given the July settlement’s size, the economics of “ask forgiveness later” are getting worse quickly, and any organization sourcing training data should be tracking this docket in the Northern District of California.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading