Ai Engineering 2 min read

Seattle Times and Newsday Sue OpenAI and Microsoft Over AI Training on Journalism

The Seattle Times and Newsday filed a federal copyright lawsuit against OpenAI and Microsoft on Friday, alleging their journalism was copied to train AI models without permission or compensation.

Two more US newspapers took AI’s largest players to court on Friday. Per Reuters, the Seattle Times and Newsday filed a federal copyright lawsuit against OpenAI and Microsoft on September 4, alleging the companies copied their journalism to train large language models without permission or compensation. GeekWire’s report notes the suit lands in OpenAI’s backyard and adds both plaintiffs to a docket that already includes the New York Times’ landmark case against the same two defendants.

Local News Joins the National Fight

The plaintiffs matter as much as the claims. Most AI copyright litigation so far has been brought by national giants (the New York Times, the Authors Guild, and the music publishers behind the Anthropic suit that produced a $1.5 billion settlement and inspired this week’s new filing against Anthropic), organizations with the budget for a decade of litigation. The Seattle Times is a regional metro daily, and Newsday is a suburban Long Island paper; if outlets of this size can sustain copyright claims, the addressable plaintiff pool for AI training disputes stops being a handful of media conglomerates and becomes thousands of local publishers whose archives are exactly the clean, dated, high-quality text corpora models are trained on.

The Overlap With the Anthropic Precedent

The legal theory tracks the post-Bartz framework that has shaped every case since: training may be transformative, but acquisition through infringement is not, the distinction that produced Anthropic’s $1.5 billion author settlement and this week’s follow-on publisher suit. What remains unsettled is whether the same split applies to LLM training corpora assembled from web scraping, which is the question the New York Times case is positioned to answer first. For engineering teams, the practical note stands from earlier in this litigation wave: dataset provenance records, licensing trails, and takedown responsiveness are now product features, because the plaintiffs’ discovery requests in these cases increasingly start with your data pipeline documentation.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading