NVIDIA announced on July 21, 2026, that production of its Vera Rubin NVL72 platform is ramping across more than 350 manufacturing sites in 30 countries. Functional racks are operating in early deployments at five cloud providers: CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius. Rather than delivering standalone accelerator cards, NVIDIA is rolling out the platform as single factory-integrated blocks designed to function as unified 72-GPU supercomputers.
Codesigned seven-chip, five-tray system architecture
NVIDIA describes the broader platform as seven chips and five rack tray types. Its components include:
- Compute and CPU: The compute trays house the Rubin GPUs alongside the NVIDIA Vera CPU, which features custom Olympus cores that NVIDIA claims offer twice the single-threaded performance, three times the core-to-core bandwidth, and 40% lower memory latency than competing chiplets.
- Dedicated acceleration: The Groq 3 LPX tray handles specialized acceleration tasks.
- Scale-out and scale-up networking: The Spectrum-6 SPX tray runs on 102.4T switching silicon alongside ConnectX-9 SuperNICs.
- System infrastructure: The Vera BlueField-4 STX tray manages data processing, offload, and isolation tasks.
The compute trays eliminate internal cables, fans, and fluid hoses, which NVIDIA says cuts compute tray assembly time from hours to one minute. Scale-up communication depends on a 260 TB/s sixth-generation NVLink fabric connecting all 72 GPUs, while scale-out networking relies on Spectrum-6 switches and optional co-packaged optics switches that NVIDIA claims yield 5x lower transceiver power.
Deployment status across initial partners
Availability remains focused on early deployments and private testing clusters rather than immediate broad-market access:
- CoreWeave: Validated live hardware and deployed liquid-cooled Spectrum-X Ethernet SN6600-LD switches.
- Google Cloud: Hosted early access on bare-metal
A5Xinstances for startup Ineffable Intelligence, combining ConnectX-9 SuperNICs with Google Virgo scale-out fabrics. - Nebius: Installed its first NVL72 unit at its Finland facility to verify the stack with Spectrum-6 switches before broad customer use.
- Microsoft Azure: Deployed systems across European infrastructure in partnership with Mistral Compute for sovereign AI and enterprise deployments.
- Oracle Cloud Infrastructure: Operating production racks and adopting co-packaged optics switches for scale-out fabrics.
Reported efficiency and cooling design
NVIDIA and its early testing partners framed the platform around power-efficiency metrics, though reported performance relies on vendor-run benchmarks rather than standardized independent testing. Running the DeepSeek-R1 mixture-of-experts model on live NVL72 hardware, CoreWeave reported achieving 10x higher token throughput per megawatt than Grace Blackwell NVL72 systems. Google Cloud similarly claimed up to 10x lower inference cost per token and 10x higher token throughput per megawatt on its bare-metal A5X instances.
Operating the systems requires compatible facility infrastructure. NVIDIA designed the liquid cooling loop around a 45-degree Celsius inlet temperature, which it states enables chiller-free dry-cooler operation and can save millions of gallons of water per megawatt annually in new facilities. The water-savings estimate refers to new facilities using the described dry-cooling and closed-loop design.
