Daily News · 2 min read

NVIDIA AI Updates: September 4, 2026

1. NVIDIA Agrees to Acquire Hugging Face for $12.9 Billion

NVIDIA. NVIDIA confirmed it will buy Hugging Face for roughly $12.93 billion, made up of an $11.9 billion purchase price plus a retention package of up to $1 billion. NVIDIA says the platform will stay independent and framework-neutral, keeping its branding and continuing to host open-source and open-weight models across all clouds and accelerators, with no requirement that developers use NVIDIA compute. Hugging Face currently serves over 18 million developers hosting 3 million models and 500,000 datasets, with more than 200,000 companies on the platform. Source

2. NVIDIA Pushes Local Inference at IFA With PAIR and RTX Spark

NVIDIA. At IFA 2026 NVIDIA announced one-click local model setup for three agent platforms: Hermes Agent from Nous Research across RTX and DGX systems on Windows, Perplexity Portable Computer on RTX GPUs with 24GB or more of VRAM, and OpenClaw’s Windows app on the same class of hardware. It also released NVIDIA PAIR, a free open-source personal AI router that spreads inference across the PCs on a local network rather than queuing everything behind one GPU, in beta for Windows, macOS, and Linux on RTX 20 Series and newer. Runtime optimizations bring up to 1.9x higher llama.cpp throughput on a GeForce RTX 5090 and up to 1.4x on vLLM across dual DGX Spark clusters. RTX Spark laptops from Lenovo and Acer, with a 1 petaflop Blackwell GPU, up to 128GB unified memory, and a 20-core Grace CPU, ship in October. Source

3. NVIDIA Publishes a Centralized Identity Gateway Pattern for Federated AI Platforms

NVIDIA. A developer blog post describes how NVIDIA solved repeated login prompts across multi-cluster, multi-region AI platforms by separating session ownership from request enforcement. One central gateway runs the OIDC flow, stores the session in Redis with a TTL, and sets an HTTP-only cookie; regional gateways stay stateless and call a shared /gateway/userinfo endpoint for trusted claims instead of running their own auth flows. Token refresh and platform-wide logout are coordinated centrally, and service-to-service calls use mutual TLS or workload identity. NVIDIA reports a 55% reduction in repeated login events across its internal developer platforms. Source