NVIDIAのNVL72はMLPerf Inference v6.1において最優れた性能を実現
NVIDIAのNVL72システム性能、効率的なインフラスケーリング、および継続的なソフトウェア最適化が、AIインフェレンス経済学において重要であり、プラットフォームの汎用性によりさまざまなワークロードにおける高利用率を実現しています。
NVIDIAのNVL72システム性能、効率的なインフラスケーリング、および継続的なソフトウェア最適化が、AIインフェレンス経済学において重要であり、プラットフォームの汎用性によりさまざまなワークロードにおける高利用率を実現しています。
Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0.
NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.
A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency
GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video — 9x higher throughput and 7.5x lower latency than a single node.
NVIDIA’s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario.
NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72.
NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite
scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs)
Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B.
Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios
On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72.
NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074.
The NVL72 scale-up domain — powered by sixth-generation NVIDIA NVLink and NVLink Switch to deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet
MLPerf Inference v6.1, Closed Division.
NVIDIA Vera Rubin NVL72 system debuts with leading performance
Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing.
each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack, while lowering cost per token.
NVIDIA delivers this with high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks and efficient request orchestration across nodes.
Nebius also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance.
In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results.