Nvidia’s next rack scales to 99 percent efficiency in MLPerf debut

Nvidia's Vera Rubin NVL72 posted its first MLPerf Inference results, running nearly four times the throughput of a GB300 rack.

ChipNews Staff
2 Min Read

The first benchmark numbers for Nvidia’s next-generation rack are in, and they land well ahead of the hardware it replaces. Vera Rubin NVL72 submitted preview results in MLPerf Inference v6.1.

The suite’s heaviest tests carried the debut. DeepSeek-R1 was one of them, run with Nvidia’s TensorRT-LLM library; the reported throughput there came in 2.5x above a GB300 rack. Qwen3-VL was the other, where the company used vLLM with its Dynamo framework and measured gains up to 3.7x across offline, server and interactive scenarios.

Scaling got its own highlight. A 288-GPU GB300 submission spread over four racks held 99 percent efficiency, with throughput climbing almost linearly from a single-rack baseline. Software work alone added up to 1.6x over the v6.0 round.

These are preview figures rather than audited submissions, so they carry less weight than a full run. Even so, they sketch the case Nvidia is making to buyers weighing inference economics: more tokens per rack, and fewer racks to power.

That case grows louder as memory prices climb and grid connections tighten. Nvidia is separately backing an alliance with Google and Emerald AI to make data center demand more flexible for utilities, the power-side version of the same argument.

Share This Article