NVIDIA AI Updates: July 22, 2026
1. NVIDIA Ramps Vera Rubin NVL72 Into Production Across Major Clouds
NVIDIA. NVIDIA said its Vera Rubin NVL72 rack-scale platform is entering production and is being deployed by cloud providers including Google Cloud, Microsoft Azure, and CoreWeave. The company cited a CoreWeave benchmark showing 10x more throughput per megawatt than the Grace Blackwell NVL72 generation, and highlighted the new Vera CPU tuned for agentic workloads such as tool calling and reinforcement learning. NVIDIA framed performance per watt and cost per token as the metrics that decide whether large-scale AI infrastructure can scale profitably. Source
2. NVIDIA Launches Spectrum-6 Ethernet Switch for Gigascale AI Factories
NVIDIA. NVIDIA introduced Spectrum-6, a 102.4-terabit-per-second Ethernet switch system built for the Vera Rubin platform that doubles the capacity of the prior generation. The company said it delivers up to 1.6x higher AI networking performance than off-the-shelf systems and sustains up to 95% network efficiency across deployments exceeding 100,000 GPUs. Early adopters include CoreWeave, Microsoft, Nebius, and Tesla, with the switch aimed at faster training and better economics at very large scale. Source
3. NVIDIA Details Rubin GPU Architecture for Agentic Inference
NVIDIA. In a technical deep dive, NVIDIA described the Rubin GPU as a 336-billion-transistor chip with 224 streaming multiprocessors, 896 Tensor Cores, 288 GB of HBM4, and 22 TB/s of memory bandwidth, delivering up to 50 petaflops of NVFP4 inference. The company said Rubin reaches roughly 10x better agentic throughput per unit of energy than Blackwell through features such as inline descriptor updates for mixture-of-experts models, activation sparsity for attention, and fine-grained dependent kernel triggering. The post targets practitioners deploying large reasoning and agentic models under fixed power budgets. Source
4. NVIDIA Reports MoE Pre-Training Record on GB300 NVL72
NVIDIA. NVIDIA reported a pre-training throughput record of 1,648 TFLOPs per GPU while training the 671-billion-parameter DeepSeek-V3 mixture-of-experts model on 256 GB300 NVL72 GPUs. That figure is roughly triple the 606 TFLOPs per GPU the company measured on the prior GB200 generation. NVIDIA attributed the gain to coordinated hardware and software optimization for MoE models, where communication bandwidth matters as much as raw compute. Source