Daily News · 2 min read

NVIDIA AI Updates: August 25, 2026

1. NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

NVIDIA. NVIDIA announced that its Vera Rubin NVL72 system delivers up to 30x higher throughput per megawatt and 35x lower token costs than the previous GB300 NVL72 generation on agentic AI workloads. NVIDIA measured performance using the SemiAnalysis AgentX workload, which reflects real agentic coding sessions with growing context and tool calls. The company notes agentic workloads consume roughly 15x more tokens than standard chat requests, making efficiency central to AI data center economics. Source

2. Groq 3 LPX Enters Full Production to Extend Vera Rubin Inference

NVIDIA. NVIDIA said its Groq 3 LPX inference accelerator is in full production, extending the Vera Rubin platform for reasoning-heavy AI agents alongside Vera CPUs and Spectrum-X Multiplane networking. NVIDIA reported the Groq 3 LPX reaches 3,400 output tokens per second on 100,000-token long-context use cases, which it described as 4x faster than the nearest alternative when tested on Gemma 4 31B. Nebius said it would be the first AI cloud to adopt the accelerator, while CoreWeave and SpaceXAI announced plans to deploy the surrounding Vera Rubin technologies. Source

NVIDIA. NVIDIA detailed NVLink Fusion, which lets companies pair custom XPU designs with NVIDIA networking, rack architecture, cooling, power, and software such as NCCL, Dynamo, and Mission Control. The company said sixth-generation NVLink offers 3x lower latency and 10x higher packet rates than Ethernet alternatives, and NVLink-C2C provides 6x better energy efficiency than PCIe for CPU connections. Intel, Quanta Computer, MediaTek, GUC, and Amazon’s Annapurna Labs voiced support for semi-custom AI factories. Source

4. BlueField-4 Powers Scale-In Network Infrastructure for Agentic AI Factories

NVIDIA. NVIDIA introduced BlueField-4 as the basis for a new scale-in network infrastructure aimed at agentic AI factories, offloading and accelerating networking, storage, and security functions as inference workloads expand into multi-step agent pipelines. It sits within the broader Vera Rubin generation of data center technologies the company detailed the same day. The platform targets the growing share of data center work spent orchestrating fleets of AI agents rather than serving single chat requests. Source