Ai Engineering 3 min read

Claude Sonnet 5.5 Matches Opus-Class Workloads at Half the Price

Anthropic released Claude Sonnet 5.5 on September 28: Terminal-Bench 70.6%, GDPval 1844 against Opus 5.5's 1846, at $2/$10 per million tokens, with first-for-Sonnet cyber safeguards and a fallback to Sonnet 5 when tasks get risky.

Anthropic released Claude Sonnet 5.5 on September 28, the second model in the 5.5 family, and the benchmark table contains a number nobody expected from this tier: Terminal-Bench 4.0 at 70.6%, which is above Opus 5.5’s 66.4% and nearly seven times Sonnet 5’s 10.3%. GDPval lands at 1844 against Opus 5.5’s 1846, a statistical tie on professional knowledge work. The price is the other headline: $2 per million input and $10 output, half of Opus 5.5’s rates and identical to Sonnet 5’s pricing from August. Two days ago Opus 5.5 reset what the flagship tier costs; today the mid-tier inherits the capability without inheriting the bill.

The Gap to Opus Is Now a Use Case, Not a Benchmark Gap

The honest framing is in Anthropic’s own caveat: Opus 5.5 remains clearly stronger on complex, open-ended work despite the near-tied scores. What Sonnet 5.5 actually does is collapse the cost of the bounded-workload tier: agentic coding sprints, ticket processing, defined migrations. The customer numbers describe it precisely. Balyasny went from 497k tokens per answer to 121k, Slack saw 14% fewer output tokens, Base44 needed 3.6 iterations per build versus Opus 5’s 7.7, and at lower effort settings Sonnet 5.5 matches Sonnet 5’s best scores at roughly a tenth of the cost per task. On FrontierCode at high effort it matches GPT-6 Sol’s best score at about one-fifth the cost, which is a direct shot in the week’s price war: OpenAI cut Sol to $2/$10 on Monday, and Sonnet 5.5 now claims Sol-class coding results at that same price.

The Safeguard Story Is the Real Upgrade

The quiet news is that Sonnet 5.5 is the first Sonnet to ship with Opus-class safeguards: cyber capability safeguards, fallbacks, and anti-distillation classifiers, the same protection stack the flagship carries. The mechanism is pragmatic: when a task triggers risky-cyber classification, the system falls back to the previous Sonnet rather than letting the stronger model run unsupervised. Anthropic’s behavioral audit adds that Sonnet 5.5 was the least likely of their models to probe container limits, with no evidence of pursuing goals conflicting with user intent. After a month where an OpenAI agent wandered into an Australian health portal and Nvidia announced hardware-level agent containment today, the mid-tier model shipping with containment-first defaults is a market signal: safety posture is becoming a competitive feature at every price point, not a flagship luxury.

The Migration Detail Worth Knowing

One operational note for teams upgrading: thinking-off users must switch to the new between_tools setting, and Haiku 5.5 arrives in the coming weeks to complete the family. Availability is immediate across AWS, Google Cloud, Azure, and the Claude Platform with zero data retention, and cache reads hold at $0.20 per million, which matters for the long-context agent workloads this model is priced for.

What to Watch

Three things. First, whether independent evals confirm Terminal-Bench 70.6%, because a mid-tier model outscoring both its own flagship and everything OpenAI sells at the price would force another round of repricing within days. Second, the fallback mechanism’s behavior in production: how often does the cyber-classifier route work back to Sonnet 5, and does that create a predictable downgrade path attackers can aim for? Third, Haiku 5.5’s price, because the sub-$2 input tier is the one battleground nobody has contested since Luna’s $0.10 debut, and Anthropic will not want to concede it.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How to Use Claude Across Excel and PowerPoint with Shared Context and Skills

Learn how to use Claude's shared Excel and PowerPoint context, Skills, and enterprise gateways for faster analyst workflows.

Ai Engineering

Claude Opus 5.5 Matches Fable 5.1 While Cutting Running Costs 40%

Anthropic released Claude Opus 5.5 on September 22, a flagship that matches or beats Fable 5.1 on most benchmarks while costing 40% less to run than Opus 5, as the frontier price war that started with Grok 4.7 claims its biggest scalp.

Ai Engineering

Grok 4.7 Launches at Half the Price of GPT-5.6 Sol and a Fifth of Claude Fable 5.1

SpaceXAI released Grok 4.7 on September 21, claiming major coding and agent gains over Grok 4.6 while pricing at $2 per million input tokens, undercutting GPT-5.6 Sol and Claude Fable 5.1 by 2x to 5x on input and far more on output.

Ai Engineering

Claude Agents Discovered a Novel CRISPR-like Enzyme System

Anthropic announced on September 23 that roughly 950 Claude agents, running 21 hours on 210 million tokens, found a previously uncharacterized bacteriophage enzyme system with CRISPR-like repeat arrays, now partly validated in the lab and published as a pre-print.

Ai Engineering

OpenAI Cuts Flagship Prices in Half With GPT-6 Sol and Luna

OpenAI launched GPT-6 Sol and Luna on September 22 and halved flagship pricing: Sol drops to $2 per million input and $10 output, while Luna lands at $0.10/$0.50, the cheapest frontier-lab tier yet, days after Grok 4.7 opened a price war.