Skip to content
StableTechnologyReported 2026-09-16 23:00

NVIDIA's NVL72 Delivers Leading Performance in MLPerf Inference v6.1

NVIDIA's NVL72 system performance, efficient infrastructure scaling, and continuous software optimization are key to AI inference economics, with platform fungibility enabling high utilization across various workloads.

01

Who it touches

  1. 1NVIDIA NVL72
  2. enables →Fact
02

Evidence

  • NNVIDIA BlogCompany2026-09-16 23:00
    Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video — 9x higher throughput and 7.5x lower latency than a single node.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA’s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs)
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    The NVL72 scale-up domain — powered by sixth-generation NVIDIA NVLink and NVLink Switch to deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    MLPerf Inference v6.1, Closed Division.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA Vera Rubin NVL72 system debuts with leading performance
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack, while lowering cost per token.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    NVIDIA delivers this with high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks and efficient request orchestration across nodes.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    Nebius also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance.
    View source
  • NNVIDIA BlogCompany2026-09-16 23:00
    In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results.
    View source