NVIDIA AI Updates: October 7, 2026
1. CUDA Green Contexts Get Runtime API Support for In-Process GPU Partitioning
NVIDIA. NVIDIA detailed green contexts, a CUDA feature that assigns dedicated SMs and workqueue resources to specific workloads inside a single process, now exposed through the runtime API in CUDA 13.1 (cudaGreenCtxCreate(), cudaExecutionCtxStreamCreate()) after first shipping in the driver API in CUDA 12.4. On a 148-SM Blackwell GPU, a latency-sensitive kernel sharing the device with bulk work ran in 0.007 ms under a green-context partition, versus 0.140 ms with stream priorities and 3.727 ms with equal priority. The feature is opt-in, so teams overlapping communication with GEMMs or running mixed real-time and batch pipelines can adopt it incrementally. Source
2. NVIDIA AI Cluster Runtime Reaches v1.0 with Stable APIs
NVIDIA. AI Cluster Runtime (AICR), NVIDIA’s open-source project for version-locked, validated recipes covering GPU drivers, container runtimes, operators, and frameworks on Kubernetes, hit v1.0 with stability contracts for its CLI, REST/OpenAPI interface, Go SDK (pkg/client/v1), and bundle schemas. The release adds a live validation dashboard, signed test evidence, and integrations with Pulumi and Mirantis k0rdent, and covers Rubin, Blackwell, Hopper, Ampere, and Ada GPUs across Ubuntu, COS, Oracle Linux, and Talos. Platform teams can use its snapshot, recipe, bundle, and validation steps to pin known-good stacks for Kubeflow, Slurm, Dynamo, and NIM workloads instead of hand-resolving version conflicts. Source
3. DOCA GPUNetIO Becomes the Shared Layer for GPU-Initiated Networking
NVIDIA. NVIDIA consolidated its GPU-initiated networking (GDA-KI) implementations onto DOCA GPUNetIO, which lets CUDA kernels drive Ethernet and RDMA Verbs objects directly. NCCL 2.27+ uses open-source GPUNetIO Verbs as a GIN backend, NVSHMEM 3.7 adds a GPUNetIO-based transport, and Holoscan Sensor Bridge 2.7.0 reports a 2.6 microsecond minimum round-trip latency on IGX Thor with ConnectX-7. For distributed training and inference engineers, this means one code path behind NCCL and NVSHMEM collectives, with NVIDIA reporting the GDA-KI backend outperforming CPU-initiated communication for Reduce-Scatter across 4 KB to 16 GB messages. Source
4. NVIDIA Outlines How Telecom Operators Are Adopting Open Models
NVIDIA. NVIDIA described how SoftBank, AT&T, and Indosat Ooredoo Hutchison are building on open models, citing its State of AI in Telecommunications report finding that 89 percent of respondents consider open source models important to their AI strategy. Examples include the 30-billion-parameter Nemotron 3 Large Telco Model fine-tuned by AdaptKey on telecom data and Indosat’s Indonesian-language Sahabat-AI family, built with NeMo and NVIDIA AI Enterprise. The post is mostly positioning, but it shows the pattern operators are following: open-weight base models fine-tuned on domain data and deployed on infrastructure they control. Source