Anthropic Researcher Quits the AI Industry, and His Lab's Safety Lead Agrees With Him
Pretraining researcher Jacob Coxon resigned from Anthropic citing a race toward self-improving AI, and alignment science lead Evan Hubinger publicly backed him with a greater-than-10% estimate that AI kills all humans within a decade.
The AI safety debate stopped being abstract on September 9. Per The Wall Street Journal, Jacob Coxon, a 27-year-old pretraining researcher with three years across OpenAI and Anthropic, resigned from Anthropic and announced he is leaving the AI industry entirely, warning that labs are “gambling with our lives” by racing toward self-improving systems they will not be able to control. He called the current phase “the endgame,” and notably judged Anthropic’s safety efforts the best of the labs he worked at; he is leaving anyway.
The Safety Lead’s Response Is the Bigger Story
What elevates this beyond a standard departure is what happened next. Evan Hubinger, Anthropic’s alignment science lead, publicly agreed with Coxon on X, stating his personal estimate that there is a greater than 10% chance AI could kill all humans within the next decade, per CNBC’s coverage. Read that carefully: the person who runs alignment science at a frontier lab endorsed a double-digit existential-risk estimate from the inside, on the same day a researcher quit over exactly that risk. The resignation and the estimate together drew calls for congressional action from multiple lawmakers, and the story was picked up across the financial and mainstream press within hours.
What It Means for the Labs and the Market
The departure is the latest in a string of safety-driven exits from frontier labs, but the combination here is new: a resignation that the lab’s own safety leadership substantively validated in public. For the labs, it sharpens a recruitment and retention problem that money does not fix; the people most qualified to evaluate frontier risk are the people most likely to leave over it, and their public statements now move markets and lawmakers. For enterprises choosing vendors, this week’s sequence (the chief scientist conceding alignment is unsolved, the METR investigation restrictions, and now this) is a governance signal to weigh alongside benchmarks: ask vendors not just what their models can do, but what their own safety staff say on the record.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Deploy Claude Code Auto Mode in Production
Learn how to configure Claude Code's auto mode to run unattended agent workflows, set up defense-in-depth tool guards, and manage the safety classifier.
OpenAI's Chief Scientist Says No Lab Has Solved Alignment, Calls for Slowdowns
In an essay titled An Alien Mind, OpenAI chief scientist Jakub Pachocki describes AI as grown rather than designed, admits no lab has solved alignment, and calls for voluntary slowdowns and third-party audited safety frameworks.
Anthropic Withheld Mythos 5.1 From the UK's Pre-Release Safety Testing
The Financial Times reports Anthropic declined to submit Claude Mythos 5.1 to the UK AI Security Institute before launch, the first time it excluded the UK from pre-release access, restricting the model to vetted US organizations.
Fable 5 Update Drops Actionable Biological Intelligence by 74%
Anthropic deployed new biology safeguards for Fable 5, cutting the model's Actionable Biological Intelligence score by 74 percent.
Identity Checks Mandatory for Claude Fable 5 After US Ban
Anthropic has restored access to Claude Fable 5 with mandatory identity verification and stricter safety classifiers following a temporary US export ban.