Ai Engineering 2 min read

Meta's Muse Spark 1.3 Hits 61 on the AI Index With a Catch: Its Best Mode Is Gated

Meta released Muse Spark 1.3 on September 2, its fourth release in five months, scoring 61 on the Artificial Analysis Intelligence Index, though the top results come from a max-reasoning variant in limited preview.

Meta released Muse Spark 1.3 on September 2, and the release cadence is now the story in itself: four Muse Spark updates in five months, per Axios, each one climbing the leaderboard. Version 1.3 scores 61 on the Artificial Analysis Intelligence Index at $0.55 per task, edging Gemini 3.8’s 59 at $0.58, and continues a steep progression from 43 for the original Muse Spark, 53 for 1.1, and 57 for 1.2. Meta’s own research post positions it around “max reasoning” for agentic tasks and improved real-world usability.

The Numbers Behind the Update

The measurable gains are broad. Independent analysis reports 75.4% on DeepSWE 1.1 for software engineering tasks, 98.5% on long-context MRCR, a 1M token context window, and a 25% reduction in tokens consumed per task, which compounds with pricing to cut real agentic workload costs meaningfully. Meta’s personal-agent focus is the strategic thread: this is the model line that powers Meta’s consumer AI assistant ambitions, and the release landed as Zuckerberg separately teased open-weight Muse Spark builds “coming soon.”

The Gated Max Mode Is the Honest Caveat

The caveat matters enough to lead with it: VentureBeat’s analysis notes Meta’s best published results come from a max-reasoning variant that developers cannot broadly use yet, so the headline score is not quite the model most teams will call. That pattern, frontier results achieved on a restricted configuration while the general-availability model scores lower, has become the industry’s quiet standard, and it makes third-party indexes like Artificial Analysis more valuable precisely because they publish the configuration they measured. For teams benchmarking against Muse Spark 1.3, run your own evaluation on the generally available variant before budgeting around the headline number.

The competitive read: Meta is iterating fastest on cost-per-task for personal agents, a different race than OpenAI’s capability-at-any-price Astra launch this week. Those are two different products winning two different budgets, and 1.3’s token-efficiency gains suggest Meta knows exactly which one it is running.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading