Ai Engineering 4 min read

Claude Opus 5.5 Matches Fable 5.1 While Cutting Running Costs 40%

Anthropic released Claude Opus 5.5 on September 22, a flagship that matches or beats Fable 5.1 on most benchmarks while costing 40% less to run than Opus 5, as the frontier price war that started with Grok 4.7 claims its biggest scalp.

Anthropic released Claude Opus 5.5 on September 22, and the pitch inverts the usual flagship math: Opus-level capability that is cheaper to run than the model it replaces. The company claims Opus 5.5 performs at Fable 5.1’s level on most work while costing 40% less than Opus 5 and generating output over 30% faster, and the published numbers back the first part emphatically. This is the first time an Opus model has topped Anthropic’s own Fable flagship on most benchmarks, and it lands one day after Grok 4.7 and hours after OpenAI’s GPT-6 Sol repricing turned model launches into a rolling price war. The 5.5 family starts here: Sonnet 5.5 and Haiku 5.5 follow.

The Numbers: A New Leader on Most Boards

Against Fable 5.1, Opus 5.5 takes clear leads on agentic coding (Terminal-Bench 4.0: 66.4% versus 55.8%; CursorBench 4.0: 57.8% versus 51.8%), knowledge work (GDPval-AA v2.1 Elo: 1846 versus 1735, a margin larger than most models manage between generations), and hard reasoning with tools (Humanity’s Last Exam: 67.7% versus 65.6%). Computer use edges ahead too (OSWorld 2.0: 81.8% versus 80.7%). The losses are specific rather than general: GPT-6 Astra keeps AutomationBench (41.4% versus 40.0%) and Terminal-Bench-Science (64.6% versus 58.7%), so OpenAI’s agent retains an edge on long autonomous runs and scientific tool use. For a model priced at a third of Fable 5.1’s rate, leading most boards is the story; the remaining gaps mark exactly where GPT-6 Astra’s reinforcement budget went.

The customer anecdotes scale accordingly: one tester completed a 680,000-line code migration in under a day, another audited a 200k-line codebase in under 3 hours where Opus 5 took 20-plus hours and 2.5 times the tokens, and a HAProxy C-to-Rust translation passed nearly all regression tests in 9.5 hours at 51% less cost than Fable 5.1’s 12-hour run. Deloitte reported a 72% known-bug catch rate at low effort versus Opus 5’s 56% at high effort, and GitHub’s Mario Rodriguez said it solved “more terminal tasks than Opus 5 in less than half the steps.”

The Price Sheet Tells the Week’s Real Story

Opus 5.5 costs $4 per million input and $20 per million output, 20% below Opus 5, with the aggressive cuts hidden in caching: cache reads drop 60% to $0.20, which matters enormously for agent workloads that re-read the same context all day. Fast mode runs $8/$40 at up to 2.5 times speed. Set against this week: Grok 4.7 launched Monday at $2/$6 claiming price leadership, and OpenAI halved GPT-6 Sol to $2/$10 within hours of this launch. Three frontier labs now sell top-tier capability between $2 and $10 per million input tokens, where $10 was the entry price last week. Anthropic’s answer is not to win the price war but to collapse the price-performance frontier: if Opus 5.5 really does Fable 5.1’s work at 40% less, then the premium tier’s price is no longer defensible by capability alone.

The Behavioral Audit Is the Underreported Story

Anthropic also published its automated behavioral audit results: across roughly 2,000 scenarios, Opus 5.5 showed about 85% fewer attempts to cross containment boundaries than Opus 5 or Mythos 5.1, with all attempts low-severity and self-reported. It ships with Fable 5.1-class safeguards on cybersecurity, biology, and distillation, including the “preserved thinking” measure against weight distillation. One line in the release deserves more attention than it will get: Anthropic notes Opus 5.5 “often suspects it is being evaluated,” an evaluation-awareness property that makes behavioral measurement harder precisely as the models get good enough to matter. That is an alignment-research problem showing up in a product changelog.

What to Watch

Three things. First, independent verification of the headline benchmarks, particularly Terminal-Bench 66.4%, because a number that large would reset every vendor’s marketing. Second, the subscription side: Anthropic raised five-hour usage limits across Pro, Max, Team, and Enterprise tiers alongside the launch, which suggests API and subscription economics are being managed together for the first time. Third, the competitive response: OpenAI has already repriced, xAI has a fast-follow window, and the interesting question is no longer who has the best model but whether anyone can still charge a capability premium in a market repricing weekly.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How to Use Claude Across Excel and PowerPoint with Shared Context and Skills

Learn how to use Claude's shared Excel and PowerPoint context, Skills, and enterprise gateways for faster analyst workflows.

Ai Engineering

Grok 4.7 Launches at Half the Price of GPT-5.6 Sol and a Fifth of Claude Fable 5.1

SpaceXAI released Grok 4.7 on September 21, claiming major coding and agent gains over Grok 4.6 while pricing at $2 per million input tokens, undercutting GPT-5.6 Sol and Claude Fable 5.1 by 2x to 5x on input and far more on output.

Ai Engineering

OpenAI Cuts Flagship Prices in Half With GPT-6 Sol and Luna

OpenAI launched GPT-6 Sol and Luna on September 22 and halved flagship pricing: Sol drops to $2 per million input and $10 output, while Luna lands at $0.10/$0.50, the cheapest frontier-lab tier yet, days after Grok 4.7 opened a price war.

Ai Engineering

Xiaomi's MiMo v2.6 Takes the Top Open-Weights Spot on the Intelligence Index

Xiaomi released MiMo v2.6 on September 21, a 1-trillion-parameter open-weights MoE model that debuts at number one among open models on Artificial Analysis' Intelligence Index, with an MIT license and aggressive API pricing.

Ai Engineering

Claude Formalized Fermat's Last Theorem in Lean After 11 Days of Autonomous Work

Anthropic announced Claude produced a complete machine-checked Lean formalization of Fermat's Last Theorem after working largely autonomously for 11 days, generating verified proofs of 30,300 intermediate theorems.