HARDWARE & SERVERS

NVIDIA Vera Rubin NVL72: Rack-Scale Agentic AI Supercomputer

NVIDIA's Vera Rubin NVL72 rack-scale platform unifies 72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, and BlueField-4 DPUs into a single coherent AI supercomputer — delivering 10x agent throughput, 3,600 PFLOPS of NVFP4 inference, and the foundation for million-GPU AI factories.

NVIDIA Vera Rubin rack-scale AI supercomputer Hardware & Servers

NVIDIA announced that the Vera Rubin platform is ramping into full production, with Taiwan's top server makers and global supply chain leaders manufacturing Vera Rubin-based systems at scale. The platform marks the third generation of NVIDIA MGX rack-scale systems and is designed to power agentic AI factories worldwide — delivering 10x agent throughput at scale compared with the previous-generation Grace Blackwell platform.

At the heart of the platform is the Vera Rubin NVL72 rack, which unifies 72 NVIDIA Rubin GPUs, 36 NVIDIA Vera CPUs, ConnectX-9 SuperNICs, and BlueField-4 DPUs. The Rubin GPUs feature HBM4 memory with 20.7 TB of total GPU memory and 1,580 TB/s of bandwidth, paired with a fifth-generation NVFP4 Transformer Engine delivering 3,600 PFLOPS of NVFP4 inference and 2,520 PFLOPS of NVFP4 training performance. The system delivers AI training with one-fourth the GPUs and AI inference at one-tenth the cost per million tokens versus the previous Blackwell generation.

The networking architecture is equally ambitious. NVLink 6 switches provide 260 TB/s of all-to-all scale-up bandwidth per rack, while ConnectX-9 SuperNICs deliver 1.6 Tb/s of per-GPU bandwidth with programmable RDMA for low-latency GPU-direct networking at massive scale. For scale-out, the platform supports both NVIDIA Quantum-X800 InfiniBand and Spectrum-X Ethernet with co-packaged optics — a new generation of switching technology built on CPO that delivers 5x better power efficiency, 5x longer AI uptime, and 1.3x faster deployment than networks using traditional transceivers.

The platform extends beyond the NVL72 rack. The Vera CPU rack delivers dense, liquid-cooled CPU infrastructure purpose-built for reinforcement learning and agentic AI, with each rack integrating 256 Vera CPUs and supporting more than 22,500 concurrent sandbox environments. The Groq 3 LPX inference accelerator, co-designed with Vera Rubin NVL72, features 256 LPUs with 128 GB SRAM, 40 PB/s memory bandwidth, and 640 TB/s scale-up bandwidth per rack — delivering 35x inference performance per watt and up to 10x more revenue opportunity for trillion-parameter models relative to Blackwell.

The Vera Rubin platform also integrates BlueField-4 DPUs with software-defined networking at speeds up to 800 Gb/s and built-in multi-tenant isolation. The NVIDIA DOCA software platform delivers advanced security across every rack and layer, enabling multi-tenant network isolation, zero-trust policy enforcement, runtime threat detection, and end-to-end encryption — all without taxing host CPU resources. Full-stack NVIDIA Confidential Computing provides hardware-level attestation to ensure the system is tamper-proof across high-speed interconnects.

Production shipments begin this fall, with major system builders including Dell, HPE, Lenovo, and Supermicro adopting NVIDIA DSX to accelerate AI factory deployment. CoreWeave, Lambda, and Oracle Cloud Infrastructure are among the first adopters of the Spectrum-X Ethernet Photonics fabric. The platform was designed with an open-source MGX design, with hundreds of supply chain ecosystem partners across 350+ factories and 30 countries ramping Vera Rubin — making it the most widely manufactured AI infrastructure platform in NVIDIA's history.

Hardware • AI Infrastructure • NVIDIA • Data Centers • Agentic AI