NVIDIA Publishes Vera Rubin Preview With Up to 3.7x Higher MLPerf Throughput

The debut benchmark results offer a sizable claimed capacity gain, but they compare two models using different serving stacks—not a universal measure of system speed.

By 3 min read
NVIDIA Publishes Vera Rubin Preview With Up to 3.7x Higher MLPerf Throughput
NVIDIA Publishes Vera Rubin Preview With Up to 3.7x Higher MLPerf Throughput

Listen to this story

The audio brief

About 1:42
0:001:42
Read transcript
NVIDIA’s first Vera Rubin NVL72 preview on MLPerf Inference version 6.1 reports up to 3.7 times the throughput of GB300 NVL72 on Qwen3-VL. That is a substantial result, but it is not a universal speed rating for Vera Rubin. The comparison covers specific models, serving software, and test modes. For Qwen3-VL, NVIDIA used vLLM with NVIDIA Dynamo across offline, server, and interactive scenarios. For DeepSeek-R1, it reported up to 2.5 times the throughput of GB300 NVL72, this time using TensorRT-LLM. That difference in software matters, because the benchmark is measuring a complete hardware-and-serving configuration, not just the chips. NVIDIA attributes the Vera Rubin results to several pieces working together: NVFP4 precision, disaggregated serving that separates prompt processing from token generation, expert parallelism for mixture-of-experts models, and sixth-generation NVLink. NVIDIA says NVLink and NVLink Switch deliver higher packet rates and lower latency than off-the-shelf Ethernet inside the NVL72 scale-up domain. The same announcement also highlighted a separate GB300 result: a 288-GPU DeepSeek-R1 submission across four racks achieved 99 percent scaling efficiency offline. Nebius submitted Vera Rubin preview results too, but no figures were provided. The key constraint is simple: these gains support NVIDIA’s rack-scale hardware-software strategy, but the real comparison will depend on the workload and serving stack.

Story brief

3 key points

NVIDIA’s MLPerf Inference v6.1 preview positions Vera Rubin NVL72 as a substantial step over GB300 NVL72, but the headline gains depend on the model, serving stack and test mode. Qwen3-VL reached up to 3.7x the comparison throughput, while DeepSeek-R1 reached up to 2.5x. NVIDIA also reported 99% four-rack scaling efficiency for a separate 288-GPU GB300 submission. For buyers and infrastructure teams, the results...

  1. 01

    Vera Rubin’s Qwen3-VL comparison used vLLM with NVIDIA Dynamo across offline, server and interactive scenarios.

  2. 02

    The DeepSeek-R1 comparison used TensorRT-LLM, underscoring that software configuration materially affects the reported gains.

  3. 03

    NVIDIA attributes Vera Rubin’s results to NVFP4, disaggregated serving, expert parallelism and sixth-generation NVLink.

NVIDIA’s first Vera Rubin NVL72 preview submission to MLPerf Inference v6.1 makes a major claim: up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL. The comparison is narrower than a general performance verdict, covering two named models, different serving software and specific benchmark scenarios.

Two workloads, two reported gains

On Qwen3-VL, NVIDIA reported its largest Vera Rubin advantage across offline, server and interactive scenarios. The submission used vLLM with NVIDIA Dynamo. On DeepSeek-R1, NVIDIA reported up to 2.5x higher throughput than GB300 NVL72 using TensorRT-LLM.

That variation is central to the announcement. The figures are results for particular hardware-and-software configurations, rather than one multiplier that applies to every AI workload. The disclosed software stack is therefore part of what the benchmark measures.

NVIDIA’s reported MLPerf snapshot
Up to 3.7xQwen3-VL throughput

NVIDIA reported this maximum Vera Rubin NVL72 throughput gain across offline, server and interactive scenarios.

Up to 2.5xDeepSeek-R1 throughput

NVIDIA reported this Vera Rubin NVL72 comparison using TensorRT-LLM.

99%GB300 scaling efficiency

A separate GB300 NVL72 DeepSeek-R1 offline submission achieved 99% scaling efficiency across four racks, according to NVIDIA.

A rack-level argument, not just a chip comparison

NVIDIA credits the Vera Rubin results to hardware-software co-design. The approach combines NVFP4 precision, disaggregated serving that separates prompt processing from token generation, and expert parallelism for mixture-of-experts models. NVIDIA also says sixth-generation NVLink and NVLink Switch deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet within the NVL72 scale-up domain.

NVIDIA chart comparing reported Vera Rubin NVL72 and GB300 NVL72 throughput in MLPerf Inference v6.1.
NVIDIA’s chart presents its preview MLPerf throughput comparisons for Vera Rubin NVL72 and GB300 NVL72 on Qwen3-VL and DeepSeek-R1. Source: blogs.nvidia.com.

GB300 provides the scale reference

Alongside the Vera Rubin debut, NVIDIA highlighted a separate GB300 result: a 288-GPU DeepSeek-R1 submission spanning four racks achieved 99% scaling efficiency in the offline scenario. That result addresses whether adding racks produces nearly proportional throughput, a different question from the Vera Rubin comparison.

NVIDIA also reported that GB300’s Qwen3-VL performance improved by up to 1.6x from MLPerf v6.0 to v6.1. It attributed those GB300 gains to lower key-value-cache precision, kernel fusion, improved kernels and disaggregated serving. Nebius also submitted Vera Rubin NVL72 preview results, according to NVIDIA.

Editorial analysis

Our Read

NVIDIA’s result is a pitch for the rack as a coordinated product, not merely a newer processor. The preview ties its strongest Qwen3-VL result to vLLM and Dynamo, while its DeepSeek-R1 result uses TensorRT-LLM. NVIDIA also highlights the network inside the NVL72 domain and near-linear GB300 scaling. The next useful comparison will be a broader set of like-for-like configurations that separates what comes from the new platform from what comes from serving software and workload design. That distinction matters as NVIDIA increasingly frames infrastructure competition around whole-system orchestration and power use.

Sources

  1. blogs.nvidia.comNVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

Loading discussion...

NVIDIA Publishes Vera Rubin Preview With Up to 3.7x Higher MLPerf Throughput | Superpower Daily