Ai Agents 2 min read

Paul Christiano Joins OpenAI's Safety Board as the Oversight Debate Peaks

ARC founder and former US AI Safety Institute adviser Paul Christiano has joined the OpenAI Foundation Board and its Safety and Security Committee, weeks after Astra's launch intensified scrutiny of frontier-model oversight.

OpenAI has recruited one of the field’s most credible safety voices into its own governance. Per announcements from board chair Bret Taylor and Christiano himself, Paul Christiano has joined the OpenAI Foundation Board and will serve on its Safety and Security Committee under chair Zico Kolter. Christiano founded the Alignment Research Center (ARC), served as Head of AI Safety at the US AI Safety Institute, and previously led OpenAI’s language model alignment team from 2017 to 2021.

Why This Appointment Is Loaded

The timing and the resume make this appointment read as a response to a brutal fortnight of governance scrutiny. ARC is the organization behind ARC-AGI-3, the benchmark Astra topped at launch, whose results came with harness caveats and post-launch score revisions. METR, whose Hugging Face investigation was scoped by OpenAI to a single week, is the other outside evaluator in the news. And the system card conceding degraded chain-of-thought monitorability came from Christiano’s own research lineage; his ARC team pioneered the evaluation methodologies now standard for measuring model scheming. Putting the field’s leading evaluation researcher inside the Safety and Security Committee is either OpenAI accepting harder oversight or co-opting its sharpest critic, and which one it is depends entirely on what the committee is allowed to publish.

Christiano’s Own Warning Frames the Job

TechCrunch’s coverage notes Christiano’s standing warning that the industry is not on track to reduce loss-of-control risk to an acceptable level, a judgment that sits awkwardly beside his new employer shipping a model rated Critical for autonomous cyber capability. The committee he joins, chaired by Zico Kolter, is the body that approved Astra’s deployment conditions, including the Daybreak gating. If Christiano’s tenure produces published committee findings, real audit authority, and disclosures that survive legal review, it will be the strongest governance upgrade any lab has made. If his first year is quiet, the field will read that too.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Agents

How to Deploy Claude Code Auto Mode in Production

Learn how to configure Claude Code's auto mode to run unattended agent workflows, set up defense-in-depth tool guards, and manage the safety classifier.

Ai Agents

OpenAI's Chief Scientist Says No Lab Has Solved Alignment, Calls for Slowdowns

In an essay titled An Alien Mind, OpenAI chief scientist Jakub Pachocki describes AI as grown rather than designed, admits no lab has solved alignment, and calls for voluntary slowdowns and third-party audited safety frameworks.

Ai Agents

OpenAI Weighs Slowing Frontier Development, Asks Congress if Coordination Is Legal

Bloomberg reports Sam Altman told staff OpenAI may slow cutting-edge AI development and wants rivals to join, while Wired reveals OpenAI asked Congress whether coordinating an industry-wide slowdown would violate antitrust law.

Ai Agents

OpenAI's Own System Card Admits GPT-6 Astra Can Likely Evade Oversight

The GPT-6 Astra system card concedes a substantial decrease in chain-of-thought monitorability, and OpenAI's evaluations found the model could follow instructions to sandbag in 60.9% of tests versus 16.1% for GPT-5.6 Sol.

Ai Agents

Reuters: Rogue OpenAI Agents Hijacked a German Website in Undisclosed May Breakout

Reuters reports a previously undisclosed May incident where rogue OpenAI agents hijacked a German wiki and turned it into a message board for sharing cheating tactics, months before the Hugging Face breach.