Meta's Muse Spark 1.3 Hits 61 on the AI Index With a Catch: Its Best Mode Is Gated
Meta released Muse Spark 1.3 on September 2, its fourth release in five months, scoring 61 on the Artificial Analysis Intelligence Index, though the top results come from a max-reasoning variant in limited preview.
Meta released Muse Spark 1.3 on September 2, and the release cadence is now the story in itself: four Muse Spark updates in five months, per Axios, each one climbing the leaderboard. Version 1.3 scores 61 on the Artificial Analysis Intelligence Index at $0.55 per task, edging Gemini 3.8’s 59 at $0.58, and continues a steep progression from 43 for the original Muse Spark, 53 for 1.1, and 57 for 1.2. Meta’s own research post positions it around “max reasoning” for agentic tasks and improved real-world usability.
The Numbers Behind the Update
The measurable gains are broad. Independent analysis reports 75.4% on DeepSWE 1.1 for software engineering tasks, 98.5% on long-context MRCR, a 1M token context window, and a 25% reduction in tokens consumed per task, which compounds with pricing to cut real agentic workload costs meaningfully. Meta’s personal-agent focus is the strategic thread: this is the model line that powers Meta’s consumer AI assistant ambitions, and the release landed as Zuckerberg separately teased open-weight Muse Spark builds “coming soon.”
The Gated Max Mode Is the Honest Caveat
The caveat matters enough to lead with it: VentureBeat’s analysis notes Meta’s best published results come from a max-reasoning variant that developers cannot broadly use yet, so the headline score is not quite the model most teams will call. That pattern, frontier results achieved on a restricted configuration while the general-availability model scores lower, has become the industry’s quiet standard, and it makes third-party indexes like Artificial Analysis more valuable precisely because they publish the configuration they measured. For teams benchmarking against Muse Spark 1.3, run your own evaluation on the generally available variant before budgeting around the headline number.
The competitive read: Meta is iterating fastest on cost-per-task for personal agents, a different race than OpenAI’s capability-at-any-price Astra launch this week. Those are two different products winning two different budgets, and 1.3’s token-efficiency gains suggest Meta knows exactly which one it is running.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Secure AI Agents With Google ADK
Learn how to secure your autonomous workflows and prevent unauthorized actions using Google ADK's hardware-backed tool binding and execution logs.
Arm Launches First In-House AGI CPU
Arm unveiled its first production silicon, a 136-core data center CPU for agentic AI workloads, with Meta as lead partner.
Meta Prices 1M-Token Muse Spark 1.1 at $1.25 Per Million Input
Meta's Superintelligence Labs has launched Muse Spark 1.1, a multimodal reasoning model for agentic workloads, alongside its first metered developer API.
Qwen 3.6-Plus Debuts With 1M-Token Context Window
Alibaba's Qwen 3.6-Plus introduces a 1-million-token context window and advanced agentic coding capabilities to challenge Claude 4.5 Opus.
QAH Pushes 4-Bit Hypernova-60B Past bfloat16 Source
Multiverse Computing’s QAH technique produces a 4-bit, 60B Hypernova-60B model that beats its bfloat16 compressed source on 7 of 9 benchmarks.