NVIDIA的NVL72在MLPerf推理v6.1中表现出卓越的性能
NVIDIA的NVL72系统性能、高效基础设施扩展以及持续的软件优化是AI推理经济性的关键,平台可互换性使得在各种工作负载中实现高利用率成为可能。
NVIDIA的NVL72系统性能、高效基础设施扩展以及持续的软件优化是AI推理经济性的关键,平台可互换性使得在各种工作负载中实现高利用率成为可能。
Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0.
NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.
A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency
GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video — 9x higher throughput and 7.5x lower latency than a single node.
NVIDIA’s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario.
NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72.
NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite
scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs)
Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B.
Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios
On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72.
NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074.
The NVL72 scale-up domain — powered by sixth-generation NVIDIA NVLink and NVLink Switch to deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet
MLPerf Inference v6.1, Closed Division.
NVIDIA Vera Rubin NVL72 system debuts with leading performance
Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing.
each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack, while lowering cost per token.
NVIDIA delivers this with high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks and efficient request orchestration across nodes.
Nebius also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance.
In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results.