Ai Engineering 4 min read

Gemini 4 Argon Lands With 1M-Token Output and a Safety-Gated Rollout

Google announced Gemini 4 Argon on September 30 with a 1 million token output limit, a DeepSWE state of the art at 77.9%, and defensive-cyber focus, but most users cannot touch it yet: rollout runs through a trusted-defenders program first.

Google announced Gemini 4 Argon on September 30, closing out the strangest release month in memory, and its two headline numbers point in opposite directions. The first is capability: a 1 million token output limit (up from 64K), a DeepSWE v1.1 state of the art at 77.9%, a 51.3% top score on Zapier’s AutomationBench, and the lead on Wiz’s black-box penetration-testing benchmark. The second is access: Argon is currently rolling out to trusted cyber defenders through the Fairwind Program, with paid API customers, Ultra subscribers, and consumers later, an unusual sequencing Google attributes to its Frontier Safety Framework, since the model is explicitly designed for autonomous offensive-and-defensive security work. Introductory pricing is $2/$10 per million tokens, rising to $4/$20 after the intro period, which lands it exactly on the GPT-6.1 Sol and Sonnet 5.5 price line the market settled at this week.

The 1M Output Limit Is the Underappreciated Number

Every frontier lab now ships million-token context windows; Argon ships a million-token output, which is a different and harder engineering claim. Sixty-four thousand tokens of output was the practical ceiling that forced agents to checkpoint, summarize, and lose state mid-task. A million tokens of continuous generation means an agent can produce an entire codebase migration, a full audit, or a complete formal proof in one uninterrupted run, which is precisely what Google says it used Argon for internally: C-to-Rust migrations of up to 800,000 lines including the Fuchsia Zircon kernel, and a libgav1 port that replaced 32K lines of SIMD code with a memory-safe decoder running 2.7 times faster. For agent builders, output length has been the quiet constraint that shaped every architecture; Argon makes it a budget problem instead.

The Cyber Focus Explains the Gated Rollout

Argon’s positioning as a defensive-cyber model explains why Google is doing something no lab has done with a flagship: shipping it to security professionals before developers. The model leads CWE-bench v1 at 68% on vulnerability remediation, discovered a critical healthcare-software flaw through Wiz’s Scan for Good program, and is designed to refuse cyber and CBRN misuse per Google’s Frontier Safety Framework, with chain-of-thought monitoring that can halt execution on detected misalignment and hard sandboxes for high-risk runs. Google also engaged the US government’s voluntary pre-release access process. The sequencing inverts the usual launch: the riskiest capability set ships to the most vetted users first, and the general public gets it after the incident track record exists. It is the model-level counterpart to Nvidia’s network-silicon containment: trust earned incrementally rather than assumed at launch.

A Month of Releases, Read Together

September’s releases now form a coherent spectrum of access philosophies. OpenAI shipped GPT-6.1 Sol and Dots to paying customers immediately; Anthropic shipped Opus and Sonnet 5.5 broadly with safeguards built in; Google ships its strongest model to vetted defenders first and everyone else later. The Argon bet is that regulatory and enterprise buyers will pay for a model whose deployment story includes gatekeeping, which matters the same week the FTC opened an investigation into AI product risks and Australia’s agent-breach review is still running. Pricing tells the same story: $2/$10 introductory is competitive, but the post-intro $4/$20 is a premium for the governance package, not just the model.

What to Watch

Three things. First, whether the Fairwind Program produces a public incident record (vulnerabilities found, misuse attempts refused) that justifies the phased rollout, because that record is the product Google is really building. Second, DeepSWE 77.9% under independent verification, since a seven-point jump over the prior SWE state of the art inside a month would be remarkable. Third, the intro-period clock: when Argon’s price doubles to $4/$20, its competitiveness against Sol 6.1 and Sonnet 5.5 at $2/$10 becomes the first real test of whether safety-gated access carries a price premium in this market.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How to Use Symbolic Execution for Automated BPF Analysis

Learn how Cloudflare uses the Z3 theorem prover to instantly generate magic packets and reverse-engineer BPF bytecode for security research.

Ai Engineering

GPT-6.1 Sol Arrives With Near-Astra Performance at a Fifth of the Price

Announced at DevDay on September 29, GPT-6.1 Sol claims near-GPT-6 Astra performance at $2/$10 per million tokens, with a 1.05 million token context window and a new Ultrafast tier at 300 tokens per second, one week after GPT-6 Sol launched.

Ai Engineering

Claude Sonnet 5.5 Matches Opus-Class Workloads at Half the Price

Anthropic released Claude Sonnet 5.5 on September 28: Terminal-Bench 70.6%, GDPval 1844 against Opus 5.5's 1846, at $2/$10 per million tokens, with first-for-Sonnet cyber safeguards and a fallback to Sonnet 5 when tasks get risky.

Ai Engineering

Claude Opus 5.5 Matches Fable 5.1 While Cutting Running Costs 40%

Anthropic released Claude Opus 5.5 on September 22, a flagship that matches or beats Fable 5.1 on most benchmarks while costing 40% less to run than Opus 5, as the frontier price war that started with Grok 4.7 claims its biggest scalp.

Ai Engineering

Grok 4.7 Launches at Half the Price of GPT-5.6 Sol and a Fifth of Claude Fable 5.1

SpaceXAI released Grok 4.7 on September 21, claiming major coding and agent gains over Grok 4.6 while pricing at $2 per million input tokens, undercutting GPT-5.6 Sol and Claude Fable 5.1 by 2x to 5x on input and far more on output.