Ai Engineering 2 min read

GPT-6 Astra Launches With Daybreak Gating and a 99.9% ARC-AGI-3 Run

OpenAI launched GPT-6 Astra on September 3, its largest-ever training run at 100,000-plus GPUs, initially limited to select organizations with advanced cyber capabilities held back for the Daybreak program.

Three days after announcing that Astra crossed its Critical cyber threshold, OpenAI has shipped it. GPT-6 Astra launched September 3, per OpenAI’s announcement, with a rollout that inverts the usual playbook: a limited set of organizations first, then all ChatGPT Plus, Pro, Business, and Enterprise users over the coming days. Greg Brockman called it a “generational leap” and told reporters “we are now in the AGI era.” The model was trained on OpenAI’s largest-ever run, more than 100,000 GPUs at its Stargate site in Texas, and priced at $10/$50 per million input/output tokens, matching Claude Fable 5.1.

The Benchmark Numbers Are Not Normal

The headline result comes from ARC Prize’s verification of ARC-AGI-3, an interactive benchmark where agents must explore novel environments and acquire goals on the fly: Astra scored 62.7% with the standard harness and 99.9% with a new provider adapter harness. For scale, Claude Opus 5 sits at 30.2% and GPT-5.6 Sol at 7.8% on the same evaluation. OpenAI is also claiming the “world’s best computer use model,” with Wired’s testing showing it booking DMV appointments and searching job listings faster than a typical person. Skeptics have noted the gap between the two harness scores raises questions about how much of the result reflects benchmark-specific adaptation rather than general capability; Forbes called the rollout a “curious false start.”

The Cyber Gating Is Now Real Product Architecture

What makes this launch different from a normal flagship release is that the Critical capability threshold announced Monday became gating on day one: advanced cyber capabilities are held behind the Daybreak program for vetted organizations, and the public model ships without them. The system card documents the split, and safety experts quoted by TechCrunch remain alarmed by the recurrent-depth reasoning technique that makes Astra’s chains of thought harder to monitor; OpenAI’s chief scientist has countered that its computation depth is within a factor of two of GPT-4. The unresolved question from the Hugging Face incident, whether monitoring tooling can keep pace with architectures designed to be efficient rather than legible, is now attached to a product millions of ChatGPT users will touch within days. For developers, the near-term practical notes are the pricing parity with Anthropic’s flagship and computer-use capability that benchmarks as best-in-class; the longer-term one is that capability thresholds are now being crossed, disclosed, and gated inside a single product cycle.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading