Daily News · 3 min read

NVIDIA AI Updates: September 16, 2026

1. NVIDIA Introduced DSX, a Data Center Platform Optimized for Tokens per Watt

NVIDIA. DSX codesigns compute, networking, power, cooling, and operations around token output per watt rather than per rack. The stack includes DSX MaxLPS for dynamically reallocating stranded power inside a fixed budget, DSX Flex for shifting workload priority in response to utility demand signals, an open-source DSX OS for lifecycle and health management, DSX Sim for pre-deployment validation, and an 800V DC distribution architecture that cuts conversion stages. Lambda validated it on HGX B200 servers, going from 4M to 5M tokens per second cluster-wide (a 24% gain) with 23% better performance per watt, and cut power 40% in under a minute during an automated demand response event. Source

2. Vera Rubin NVL72 Claims 30x Throughput per Megawatt Over GB300

NVIDIA. At AI Infra Summit, NVIDIA put Vera Rubin NVL72 at up to 30x higher throughput per megawatt than GB300 NVL72 on DeepSeek V4 Pro, and up to 45x lower cost per million tokens on agentic workloads, with DSX power management freeing enough headroom for roughly 40% more GPU capacity in the same megawatt budget. Amazon’s Annapurna Labs is working with NVIDIA on NVHBM custom high-bandwidth memory, and d-Matrix is pairing Vera CPUs with its Raptor XPUs over NVLink Fusion for low-latency inference. Startup benchmarks on Vera included Redpanda at 5.5x lower latency and 73% higher throughput, Starburst at 3x faster query throughput, and Kinetica at 2.7x. Vera Rubin systems are in full production. Source

3. Salesforce Built Its First Reasoning Model, Koa, on Nemotron 3 Super

NVIDIA. Koa is a CRM reasoning model post-trained by Salesforce and NVIDIA on 27 years of Salesforce enterprise data, using supervised fine-tuning and reinforcement learning through NVIDIA NeMo, with synthetic scenarios spanning more than 14 industries. NVIDIA says it matches leading competitors while making 3x fewer errors on CRM tasks, and it runs entirely inside Salesforce infrastructure with no customer data used in training or inference. It is already driving internal Slack agents, with customer pilots starting in October at Formula 1 and UChicago Medicine among others, and general availability targeted for winter 2026. Source

4. CHOP Cut Cardiac Model Generation From Four Hours to Seconds With MONAI

NVIDIA. Children’s Hospital of Philadelphia is building patient-specific 3D cardiac models for congenital heart disease using an open-source NVIDIA stack: MONAI and MONAI Label for medical imaging, Auto3DSeg for segmentation, the Newton physics engine and NVIDIA Warp for simulation, and Omniverse with OpenUSD for the 3D pipeline. Modeling time dropped from about four hours to seconds, and CHOP expects roughly 200 modeled cases this year. More than 20 U.S. children’s hospitals now run cardiac modeling programs, with Boston Children’s using it on over half its cardiac surgeries, about 500 cases annually. Source