Nvidia's Vera Rubin NVL72 Debuts in MLPerf With Up to 3.7x Gains Over Blackwell
Nvidia's next-generation Vera Rubin NVL72 made its first MLPerf Inference submission, delivering up to 3.7x the throughput of GB300 NVL72 with 99% scaling efficiency, as AMD and Intel benchmarked competing systems in the same round.
Nvidia’s next-generation data center platform has its first public benchmark. Per Nvidia’s announcement, the Vera Rubin NVL72 debuted in MLPerf Inference v6.1 with up to 3.7x the throughput of the previous-generation GB300 NVL72, and achieved 99% scaling efficiency across its 72-GPU systems. The submission marks the production arrival of the Vera Rubin architecture, the successor to Blackwell that Nvidia has projected will drive $1 trillion in business through 2027.
The Numbers Behind the Debut
The detail that matters for buyers is per-GPU serving throughput: cloud provider Nebius posted the leading server result at 16,427 tokens per second per GPU, with verified server throughput ranging from 55,869 to 63,909 tokens/s per system. SemiAnalysis’s independent analysis estimated roughly 3x better performance per megawatt versus Blackwell on trillion-parameter models, the metric that actually determines inference economics at data center scale. The round also drew competition: AMD benchmarked a 512-GPU MI355X cluster and Intel submitted Arc Pro and Xeon results, but Nvidia retained the lead in the categories that matter for large-model serving.
Why Inference Benchmarks Move Markets Now
The industry’s center of gravity has shifted from training to serving, and MLPerf Inference has become the procurement reference for that shift. Every point of tokens-per-second-per-watt translates directly into margin for the neoclouds and enterprises buying racks, which is why Dell’s $95 billion AI backlog and the data-center financing wave both price off these platforms. A 3.7x generational leap also pressures the inference-chip challengers (the Dutch startup Euclyd’s $231 million raise this week among them) whose entire value proposition is competing against Nvidia’s pace.
What to Watch
Two caveats keep the celebration honest. First, harness effects are real: this month’s Astra benchmark revisions showed how configurations move scores, and MLPerf submissions are vendor-configured within the rules. Second, the 3.7x claim compares against Nvidia’s own previous generation, not against the best AMD result; the cross-vendor gaps in this round are narrower than the generational claim suggests. The deployment signal to watch is Nebius and the other neoclouds: when Vera Rubin systems appear in their public pricing, the benchmark numbers have become buyable capacity.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
Why Local AI Belongs in Your Personal Tech Stack
Cloud AI access is conditional. A local model gives you a private, offline capability that remains available when networks and providers fail.
Apple Plans a Return to Servers: M8 Ultra AI Inference Machines by 2029
The Information reports Apple is developing enterprise AI inference servers pairing two or four M8 Ultra chips, has talked with Nvidia about NVLink Fusion networking, and targets a 2029 launch.
Anthropic's IPO Targets a $2 Trillion Valuation With Nvidia in Talks to Anchor
Investors expect Anthropic to go public in October at a $2 trillion or higher valuation in what would be the largest IPO in history, with Nvidia reportedly in talks to invest up to $10 billion as the anchor investor.
OpenAI Quietly Raised Astra's ARC-AGI-3 Score From 98.6% to 99.99% After Launch
Fortune reports OpenAI revised GPT-6 Astra's published ARC-AGI-3 metrics upward after launch, amid an unusual delay in the announcement blog post, deepening scrutiny of benchmark disclosure practices.
Meta's Muse Spark 1.3 Hits 61 on the AI Index With a Catch: Its Best Mode Is Gated
Meta released Muse Spark 1.3 on September 2, its fourth release in five months, scoring 61 on the Artificial Analysis Intelligence Index, though the top results come from a max-reasoning variant in limited preview.