Ai Engineering 2 min read

Nvidia's Vera Rubin NVL72 Debuts in MLPerf With Up to 3.7x Gains Over Blackwell

Nvidia's next-generation Vera Rubin NVL72 made its first MLPerf Inference submission, delivering up to 3.7x the throughput of GB300 NVL72 with 99% scaling efficiency, as AMD and Intel benchmarked competing systems in the same round.

Nvidia’s next-generation data center platform has its first public benchmark. Per Nvidia’s announcement, the Vera Rubin NVL72 debuted in MLPerf Inference v6.1 with up to 3.7x the throughput of the previous-generation GB300 NVL72, and achieved 99% scaling efficiency across its 72-GPU systems. The submission marks the production arrival of the Vera Rubin architecture, the successor to Blackwell that Nvidia has projected will drive $1 trillion in business through 2027.

The Numbers Behind the Debut

The detail that matters for buyers is per-GPU serving throughput: cloud provider Nebius posted the leading server result at 16,427 tokens per second per GPU, with verified server throughput ranging from 55,869 to 63,909 tokens/s per system. SemiAnalysis’s independent analysis estimated roughly 3x better performance per megawatt versus Blackwell on trillion-parameter models, the metric that actually determines inference economics at data center scale. The round also drew competition: AMD benchmarked a 512-GPU MI355X cluster and Intel submitted Arc Pro and Xeon results, but Nvidia retained the lead in the categories that matter for large-model serving.

Why Inference Benchmarks Move Markets Now

The industry’s center of gravity has shifted from training to serving, and MLPerf Inference has become the procurement reference for that shift. Every point of tokens-per-second-per-watt translates directly into margin for the neoclouds and enterprises buying racks, which is why Dell’s $95 billion AI backlog and the data-center financing wave both price off these platforms. A 3.7x generational leap also pressures the inference-chip challengers (the Dutch startup Euclyd’s $231 million raise this week among them) whose entire value proposition is competing against Nvidia’s pace.

What to Watch

Two caveats keep the celebration honest. First, harness effects are real: this month’s Astra benchmark revisions showed how configurations move scores, and MLPerf submissions are vendor-configured within the rules. Second, the 3.7x claim compares against Nvidia’s own previous generation, not against the best AMD result; the cross-vendor gaps in this round are narrower than the generational claim suggests. The deployment signal to watch is Nebius and the other neoclouds: when Vera Rubin systems appear in their public pricing, the benchmark numbers have become buyable capacity.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading