Grok 4.7 Launches at Half the Price of GPT-5.6 Sol and a Fifth of Claude Fable 5.1
SpaceXAI released Grok 4.7 on September 21, claiming major coding and agent gains over Grok 4.6 while pricing at $2 per million input tokens, undercutting GPT-5.6 Sol and Claude Fable 5.1 by 2x to 5x on input and far more on output.
SpaceXAI announced Grok 4.7 on September 21, calling it the company’s most powerful model for coding and knowledge work, and the announcement’s sharpest edge is not a benchmark, it is the price list: $2 per million input tokens and $6 per million output. For comparison, xAI’s own table lists GPT-5.6 Sol at $4/$20 and Anthropic’s Fable 5.1 at $10/$50. That makes Grok 4.7’s output tokens 3x cheaper than OpenAI’s comparable tier and 8x cheaper than Anthropic’s, at claimed quality within striking distance of both. The model is live today in Cursor, Grok Build, the Grok API, and third-party harnesses, routers, and cloud platforms.
The Benchmarks: Big Steps, No Crossover
Against Grok 4.6, the gains are large across the agentic coding suite: CursorBench 4.0 jumps from 40.4% to 46.3%, Terminal-Bench 4.0 nearly doubles from 20.3% to 38.0%, EEBench goes from 53.0% to 64.0%, and DeepSWE v1.1 rises from 65.2% to 71.0% at high effort. Vertical benchmarks moved too: Harvey’s legal agent benchmark from 15.8% to 19.6% and HealthBench Professional from 48.5% to 56.7%. The architecture story behind the numbers is conventional but credible: a larger base model with a longer reinforcement learning run weighted toward multi-hour problems, better self-verification, and training that natively targets the Grok Bot agent harness rather than adapting a chat model to it after the fact.
Against the competition the picture is more honest than the marketing. On CursorBench 4.0, Grok 4.7’s 46.3% clears GPT-5.6 Sol Max (41.7%) but stays well behind Claude Fable 5.1 Max (51.8%), and on GDPval its 1695 Elo (xhigh) sits just under Fable 5.1’s 1735. The claim that survives scrutiny is not “best model” but “best value”: it beats one frontier rival outright on coding-agent tasks and prices far below both.
The Price War Math
The pricing deserves more attention than the benchmarks, because it compounds. Grok 4.5 introduced the $2 per million input tier back when it was a value play against weaker models; 4.7 holds that price while moving quality into the frontier conversation. For agent workloads that burn millions of tokens per task, the gap between $6 and $50 per million output tokens is the difference between a marginable product and a loss leader, and it explains why Grok has been showing up in enterprise share discussions alongside models from labs with far larger revenue bases. If coding agents are the next great consumer of inference (and Cursor, Devin-class products, and Grok Build all say they are), xAI is pricing to be the default substrate for that consumption, not the premium option.
The fast variant doubles output speed at double price, which is aimed squarely at interactive coding loops where time-to-token is the perceived quality.
The Safeguard Stack Is Part of the Pitch
Unusually for a launch focused on price-performance, xAI led with safety engineering claims too: an entirely new safeguard stack with jailbreak resistance and calibrated refusals, quantified as only 3.3% of risky dual-use prompts passing through in HackerBench v0.3, and a top score of 62.4% on LatchBio’s biosafety evaluation. Red-team access for selected cybersecurity partners is invite-only at launch. The context matters here: Grok Build was caught uploading full git repositories and encrypted prompts exposed Grok chat data in 40% of tests in recent incidents, so the security section reads as a direct response to the company’s own track record rather than abstract virtue. Whether the new stack holds under the community’s adversarial testing, which is historically how Grok safety claims have been stress-tested, is the practical question.
What to Watch
Three threads to follow. First, whether the price-performance claim holds up in independent evaluations the way the 4.6-to-4.7 deltas suggest, particularly on long-horizon agent tasks where self-verification claims are hardest to verify. Second, enterprise share: GPT-6 Astra has been reported at roughly 13% enterprise share against Fable’s 8%, and a frontier-capable model at a fifth of the output price is designed to move those numbers. Third, whether Anthropic and OpenAI answer on price, because the current spread (8x on output tokens versus Fable 5.1) is not a equilibrium either incumbent can enjoy for long if Grok’s quality claims survive contact with production workloads.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
Agent Skills vs Cursor Rules: When to Use Each
Cursor has both rules and skills for customizing the AI agent. They overlap, but they're not the same. Here's when to use each and how they interact.
Xiaomi's MiMo v2.6 Takes the Top Open-Weights Spot on the Intelligence Index
Xiaomi released MiMo v2.6 on September 21, a 1-trillion-parameter open-weights MoE model that debuts at number one among open models on Artificial Analysis' Intelligence Index, with an MIT license and aggressive API pricing.
Qwen 3.6-Plus Debuts With 1M-Token Context Window
Alibaba's Qwen 3.6-Plus introduces a 1-million-token context window and advanced agentic coding capabilities to challenge Claude 4.5 Opus.
Meta's Muse Spark 1.3 Hits 61 on the AI Index With a Catch: Its Best Mode Is Gated
Meta released Muse Spark 1.3 on September 2, its fourth release in five months, scoring 61 on the Artificial Analysis Intelligence Index, though the top results come from a max-reasoning variant in limited preview.
Pentagon Puts ChatGPT Mil and Grok for Government in Front of 3M Troops
The Pentagon's GenAI.mil platform added ChatGPT Mil and Starshield AI's Grok for Government on August 31, bringing frontier chatbots to roughly 3 million civilian and military personnel.