NVIDIA Publishes Vera Rubin Preview With Up to 3.7x Higher MLPerf Throughput
The debut benchmark results offer a sizable claimed capacity gain, but they compare two models using different serving stacks—not a universal measure of system speed.
Listen to this story
The audio brief
Story brief
3 key pointsNVIDIA’s MLPerf Inference v6.1 preview positions Vera Rubin NVL72 as a substantial step over GB300 NVL72, but the headline gains depend on the model, serving stack and test mode. Qwen3-VL reached up to 3.7x the comparison throughput, while DeepSeek-R1 reached up to 2.5x. NVIDIA also reported 99% four-rack scaling efficiency for a separate 288-GPU GB300 submission. For buyers and infrastructure teams, the results...
- 01
Vera Rubin’s Qwen3-VL comparison used vLLM with NVIDIA Dynamo across offline, server and interactive scenarios.
- 02
The DeepSeek-R1 comparison used TensorRT-LLM, underscoring that software configuration materially affects the reported gains.
- 03
NVIDIA attributes Vera Rubin’s results to NVFP4, disaggregated serving, expert parallelism and sixth-generation NVLink.
NVIDIA’s first Vera Rubin NVL72 preview submission to MLPerf Inference v6.1 makes a major claim: up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL. The comparison is narrower than a general performance verdict, covering two named models, different serving software and specific benchmark scenarios.
Two workloads, two reported gains
On Qwen3-VL, NVIDIA reported its largest Vera Rubin advantage across offline, server and interactive scenarios. The submission used vLLM with NVIDIA Dynamo. On DeepSeek-R1, NVIDIA reported up to 2.5x higher throughput than GB300 NVL72 using TensorRT-LLM.
That variation is central to the announcement. The figures are results for particular hardware-and-software configurations, rather than one multiplier that applies to every AI workload. The disclosed software stack is therefore part of what the benchmark measures.
NVIDIA reported this maximum Vera Rubin NVL72 throughput gain across offline, server and interactive scenarios.
NVIDIA reported this Vera Rubin NVL72 comparison using TensorRT-LLM.
A separate GB300 NVL72 DeepSeek-R1 offline submission achieved 99% scaling efficiency across four racks, according to NVIDIA.
A rack-level argument, not just a chip comparison
NVIDIA credits the Vera Rubin results to hardware-software co-design. The approach combines NVFP4 precision, disaggregated serving that separates prompt processing from token generation, and expert parallelism for mixture-of-experts models. NVIDIA also says sixth-generation NVLink and NVLink Switch deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet within the NVL72 scale-up domain.
GB300 provides the scale reference
Alongside the Vera Rubin debut, NVIDIA highlighted a separate GB300 result: a 288-GPU DeepSeek-R1 submission spanning four racks achieved 99% scaling efficiency in the offline scenario. That result addresses whether adding racks produces nearly proportional throughput, a different question from the Vera Rubin comparison.
NVIDIA also reported that GB300’s Qwen3-VL performance improved by up to 1.6x from MLPerf v6.0 to v6.1. It attributed those GB300 gains to lower key-value-cache precision, kernel fusion, improved kernels and disaggregated serving. Nebius also submitted Vera Rubin NVL72 preview results, according to NVIDIA.
Editorial analysis
Our Read
NVIDIA’s result is a pitch for the rack as a coordinated product, not merely a newer processor. The preview ties its strongest Qwen3-VL result to vLLM and Dynamo, while its DeepSeek-R1 result uses TensorRT-LLM. NVIDIA also highlights the network inside the NVL72 domain and near-linear GB300 scaling. The next useful comparison will be a broader set of like-for-like configurations that separates what comes from the new platform from what comes from serving software and workload design. That distinction matters as NVIDIA increasingly frames infrastructure competition around whole-system orchestration and power use.
Sources
- blogs.nvidia.comNVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.