Daily News · 2 min read

NVIDIA AI Updates: September 30, 2026

1. VSS Blueprint 3.3 Cuts Video Summarization Tokens by 80 Percent With Adaptive Frame Pruning

NVIDIA. Version 3.3 of the Video Search and Summarization (VSS) Blueprint adds Adaptive Efficient Video Sampling, which prunes unchanged visual patches in each frame before they reach the vision-language model. On an RTX PRO 6000 Blackwell running Cosmos 3 Super in FP8, NVIDIA reports 80 percent fewer VLM input tokens for a 60-minute video summary, completed in about half the time, 46 percent more concurrent real-time streams (13 to 19), and alert contextualization latency down 17 percent (1,021 ms to 844 ms). A new vss-build-vision-ai agent skill composes deployments from a natural language prompt starting from one of four validated profiles; in a bottling-line demo, search, alert verification, and shift reporting were running in under 30 minutes on two RTX PRO 6000 GPUs. Source

2. NVIDIA Says Validation, Not Code, Is the Bottleneck in Agent-Built TensorRT Model Connect

NVIDIA. NVIDIA published lessons from building TensorRT Model Connect, an open-source, public-preview set of C++ reference implementations that turn Hugging Face or local checkpoints into versioned .bundle artifacts with task-oriented APIs for text, vision, audio, diffusion, segmentation, and embedding models. The project covered 128 model families at its July public release, all tested on GB300. The team credits keeping model families isolated so agents can build them in parallel without cascading failures, giving agents outcomes and validation criteria instead of implementation instructions, and moving engineers toward acceptance criteria and reproducible CI, since generated code is now cheap and evidence is the constraint. Source

3. NVIDIA Released Kumo Tabular, an Open Foundation Model That Predicts on Tables Without Training

NVIDIA. Kumo Tabular is a family of in-context learning transformers, from 28M to 215M parameters, that classify or regress new rows from labeled context rows with no training, tuning, or feature engineering. It combines column, row, and in-context attention, and was pretrained only on synthetic tables sampled from structural causal models, scaling to 60,000 rows and 100 columns in its final stage. NVIDIA reports first place on TabArena (Elo 1950, 17x faster than LimiX-2), BeyondArena, and TALENT. Weights are on Hugging Face under the commercially usable OpenMDW-1.1 license, with code in NVIDIA/structured-data-models. Source