Researchers Welcome Embedded Safety Evaluators, Then Ask the Obvious Question
Safety researchers welcomed Anthropic and OpenAI's commitment to embed independent evaluators inside their labs as unprecedented access, while TechCrunch's coverage asks whether the evaluators will really be independent.
The embedded-evaluator commitments from Anthropic’s pacing essay and OpenAI’s matching pledge have drawn their first serious scrutiny, and the verdict from the safety community is a conditional yes. Per TechCrunch’s coverage, researchers welcome the unprecedented access, with several calling it a genuine step forward, while the same coverage asks the question that decides everything: will the evaluators really be independent?
What Was Actually Promised, and What Was Not
The commitments are specific about access and silent about authority. Anthropic has committed to permanent, employee-level access for outside evaluators and publication of findings without editorial control, with only narrow redaction rights. OpenAI matched with a pledge of employee-like access for independent evaluators. What neither company has specified is whether evaluators can choose their own scope, whether findings require pre-approval before publication, how disagreements between evaluator and lab get resolved, and whether the arrangement survives a change of leadership. The Frontier Act’s version answers some of this by statute (outside audits, incident reporting, an Under Secretary of Commerce for AI Security), which is why the researchers’ welcome came with a call for lawmakers to enshrine the commitments quickly.
The Precedent That Should Worry Everyone
The week’s reporting provides the test case for why independence language matters. When OpenAI limited METR’s investigation to a single week of the Hugging Face breach, the restriction was defensible operationally and corrosive procedurally: an outside evaluator with scoped access cannot contradict the company that scopes it. Embedded evaluators with contractual publication rights are the structural fix, because their leverage comes from the contract rather than the relationship. The details to read in the coming contracts are three: who selects the evaluators, whether redaction is limited to enumerated categories, and whether evaluators can unilaterally escalate to regulators when redaction disputes deadlock.
For teams building AI products, the debate is a preview of their own audits: enterprise customers are already asking for the same embedded-evaluator guarantees for deployed agents, and the contract language hashed out between METR and the frontier labs will become the template vendors get handed.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Deploy Claude Code Auto Mode in Production
Learn how to configure Claude Code's auto mode to run unattended agent workflows, set up defense-in-depth tool guards, and manage the safety classifier.
Altman, Musk, and Hassabis Back Amodei's Plan: All Four Frontier Labs Agree to Pace
Sam Altman pledged OpenAI will adopt independent evaluators with employee-like access, Elon Musk said Dario is right, and Demis Hassabis endorsed the slowdown, marking the first time all four frontier lab chiefs have publicly aligned on pacing.
OpenAI's Chief Scientist Says No Lab Has Solved Alignment, Calls for Slowdowns
In an essay titled An Alien Mind, OpenAI chief scientist Jakub Pachocki describes AI as grown rather than designed, admits no lab has solved alignment, and calls for voluntary slowdowns and third-party audited safety frameworks.
Zuckerberg Breaks From the AI Pacing Consensus: Every Lab Paces Itself, No Coordination
Mark Zuckerberg rejected coordinated AI slowdown calls, arguing every lab has the responsibility and incentive to pace itself safely, splitting the frontier-lab consensus that OpenAI, Anthropic, xAI, and Google DeepMind built this week.
Anthropic Withheld Mythos 5.1 From the UK's Pre-Release Safety Testing
The Financial Times reports Anthropic declined to submit Claude Mythos 5.1 to the UK AI Security Institute before launch, the first time it excluded the UK from pre-release access, restricting the model to vetted US organizations.