Claude Haiku 5.5 Collapses Small-Model Prices to $0.10 Per Million Input Tokens
Anthropic released Claude Haiku 5.5 on October 7 with GDPval 1620 against Haiku 4.5's 735, an adjustable effort setting, and pricing of $0.10/$0.50 per million tokens for prompts under 100k, roughly 90% cheaper than its predecessor.
Anthropic released Claude Haiku 5.5 on October 7, completing the 5.5 family, and the number that ends the month’s pricing war is the small one: $0.10 per million input tokens and $0.50 output for prompts under 100k, roughly 90% cheaper than Haiku 4.5, with cache reads at a penny. When Sonnet 5.5 launched two weeks ago, the open question was who would contest the sub-$2 input tier that OpenAI’s Luna claimed at $0.10. Anthropic’s answer arrives a dollar below the line, with the benchmark profile of a model a tier above.
The Benchmarks: A Category Redefinition
Against Haiku 4.5, the generation gap is the widest Anthropic has published: GDPval-AA knowledge work at 1620 versus 735, OSWorld computer use at 72.4% versus 15.7%, Terminal-Bench 4.0 at 39.2% versus literally zero, and Humanity’s Last Exam at 57.4% with tools versus 18.7%. The comparisons that matter commercially are against the tiers above: Haiku 5.5’s 46.4% on FrontierCode beats GPT-6 Luna’s 42.4%, its HLE score approaches Sonnet 5.5’s, and HubSpot scored it 92.8% on a CRM task suite while Cognition runs it as Devin Fusion’s sidekick at a FrontierCode 66.2. Anthropic is explicit that Opus 5.5 remains stronger on complex open-ended work, but the tier that used to mean “weak but cheap” now means “cheap, and strong enough that the flagship’s remaining advantages need long-horizon tasks to show.”
The Effort Dial Is the Architectural Story
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, Low through Max, trading cost against intelligence per call. Combined with pricing that halves again above 100k prompt tokens (input $0.50, output $2.50), Anthropic is pricing the same weights along a cost-intelligence curve instead of shipping separate model sizes per tier. That is a structural shift in how small models are sold: one checkpoint, dialed per workload, with the dial exposed in the API rather than chosen by product managers. For high-volume agent builders, the practical effect is that summarization, compaction, classification, and subagent calls, the workloads Anthropic names explicitly, can run at Low effort while escalations dial up, all inside one model ID with one integration.
The Ecosystem Moves With It
The release drags the rest of Anthropic’s pricing along: Sonnet 5.5 cache reads halved to $0.10 per million tokens, making agentic work about 20% cheaper retroactively, and Max and Team subscribers now get monthly API credits ($100 at 5x, $200 at 20x, up to $500 pooled for Team), which quietly merges the subscription and API businesses. Early customer numbers describe the latency story: Asana reported over 30% latency reduction and up to 2.5 times faster inference per agent turn, and Box reported an 11-point quality improvement at roughly half the latency versus Haiku 4.5. Safety posture scaled as designed: cyber safeguards are stricter than Haiku 4.5 but looser than Sonnet 5.5, still blocking penetration testing, with biology safeguards matching the flagship tier.
What to Watch
Three things. First, OpenAI’s response on Luna: Luna’s $0.10/$0.50 was the cheapest frontier-lab tier; Haiku 5.5 matching it with GDPval 1620 makes the next move OpenAI’s, and the week’s pattern says it will come fast. Second, subagent economics: Cognition’s Devin Fusion result implies agent frameworks will re-point their sidekick calls at Haiku 5.5, which is the volume market, and watching which frameworks switch tells you where the real workloads are. Third, the family-complete question: with Opus, Sonnet, and Haiku 5.5 all shipped and monthly API credits binding subscriptions to the API, Anthropic’s next release cycle starts from a unified price ladder, and the first 5.x price change will show whether this ladder is stable or another front in the price war.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Use Claude Across Excel and PowerPoint with Shared Context and Skills
Learn how to use Claude's shared Excel and PowerPoint context, Skills, and enterprise gateways for faster analyst workflows.
Claude Sonnet 5.5 Matches Opus-Class Workloads at Half the Price
Anthropic released Claude Sonnet 5.5 on September 28: Terminal-Bench 70.6%, GDPval 1844 against Opus 5.5's 1846, at $2/$10 per million tokens, with first-for-Sonnet cyber safeguards and a fallback to Sonnet 5 when tasks get risky.
Claude Opus 5.5 Matches Fable 5.1 While Cutting Running Costs 40%
Anthropic released Claude Opus 5.5 on September 22, a flagship that matches or beats Fable 5.1 on most benchmarks while costing 40% less to run than Opus 5, as the frontier price war that started with Grok 4.7 claims its biggest scalp.
Gemini 4 Argon Lands With 1M-Token Output and a Safety-Gated Rollout
Google announced Gemini 4 Argon on September 30 with a 1 million token output limit, a DeepSWE state of the art at 77.9%, and defensive-cyber focus, but most users cannot touch it yet: rollout runs through a trusted-defenders program first.
GPT-6.1 Sol Arrives With Near-Astra Performance at a Fifth of the Price
Announced at DevDay on September 29, GPT-6.1 Sol claims near-GPT-6 Astra performance at $2/$10 per million tokens, with a 1.05 million token context window and a new Ultrafast tier at 300 tokens per second, one week after GPT-6 Sol launched.