NVIDIA Vera Rubin Platform: First Commercial Shipments and 10x Power Efficiency Benchmarks
As the artificial intelligence infrastructure landscape transitions from Blackwell to next-generation architectures, a major milestone has been reached with the first commercial shipments and live-silicon benchmarking of the NVIDIA Vera Rubin NVL72 platform. In June and July 2026, Dell Technologies and CoreWeave completed the industry's first physical delivery, bring-up, and performance evaluation of the unified Rubin system, demonstrating historic efficiency gains that directly target the power constraints of modern AI data centers.
Physical Delivery and Bring-Up
In June 2026, Dell Technologies became the first hardware manufacturer to ship rack-scale systems engineered on the NVIDIA Vera Rubin platform. Dell's PowerRack systems, integrated with PowerEdge XE9812 servers housing Vera Rubin NVL72, were delivered to AI cloud provider CoreWeave.
By early June 2026, CoreWeave successfully completed the industry's first end-to-end bring-up and validation of the system, verifying power, liquid cooling, networking, and compute at production scale. This physical milestone marks the end of the initial delay phase for next-generation hardware deployments and signals a rapid commercial ramp. According to Nvidia, production of the Vera Rubin NVL72 is actively ramping up, with initial racks running across elite cloud partners including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius Group.
Historic 10x Power Efficiency Gains on DeepSeek R1
On July 21, 2026, CoreWeave and Nvidia published the first-ever measured silicon performance numbers for the Vera Rubin NVL72 platform, running the advanced reasoning model DeepSeek R1. The results reveal a massive leap in power efficiency:
- 10x Tokens Per Megawatt: At matched interactivity targets (tokens-per-second per user), the Vera Rubin NVL72 generated 10 times more tokens-per-second per megawatt compared to the NVIDIA Grace Blackwell (GB200) NVL72 on the same DeepSeek R1 workload.
- Reduced Footprint and Cost: The Rubin platform delivers up to 10x better performance-per-watt, requires up to 25% fewer GPUs, and offers a one-tenth cost per million tokens compared to Blackwell.
Extreme Codesign for the Agentic AI Era
The Vera Rubin NVL72's performance leap is driven by extreme codesign across seven chips and five rack trays engineered as a single, unified system rather than assembled from off-the-shelf parts.1 Each rack unifies:
- 72 NVIDIA Rubin GPUs
- 36 NVIDIA Vera CPUs
- A 260 TB/s all-to-all NVLink 6th-generation fabric
- Native NVFP4 precision support built into the Transformer Engine
- ConnectX-9 SuperNICs and BlueField-4 DPUs
This architecture is purpose-built for "Agentic AI" workloads. Because reasoning models like DeepSeek R1 plan, check their work, and generate extensive internal tokens before outputting a final response, they place a heavy burden on memory bandwidth and interconnect fabrics. By sharing a massive NVLink domain and liquid-cooled switching, the Vera Rubin NVL72 dramatically reduces the latency and energy cost of multi-step agentic inference, providing a critical physical solution for power-constrained data centers.2
-
An instance of AI hardware dominance requires owning the entire stack from training to agentic orchestration. — Achieving performance gains for agentic orchestration requires owning and co-designing the entire stack from silicon to system architecture. ↩︎
-
An instance of The transition to agentic intelligence forces an immediate and continuous hardware replacement cycle. — The transition to multi-step agentic reasoning forces operators to deploy next-generation architectures like Vera Rubin that are specifically optimized to handle the heavy memory and power demands of agentic inference. ↩︎