Hugging Face AI Updates: October 2, 2026
1. Ai2 Releases Olmo-core 3 for Training Trillion-Parameter MoE Models
Hugging Face. Ai2 published Olmo-core 3 on the Hugging Face blog, an open-source training framework rebuilt for large mixture-of-experts models that combines expert parallelism, pipeline parallelism, and a distributed optimizer, with rowwise expert parallelism, GPU-resident routing, and grouped GEMMs to cut communication overhead. Ai2 reports runs up to 1.2 trillion total parameters (58.36B active) on 512 GPUs at 858 TFLOP/s per GPU, and 2.7x the throughput of its earlier FSDP setup at 52,000 tokens per second per GPU. MXFP8 training adds 21 percent throughput over BF16 while trimming peak memory from 103 GiB to 95 GiB, and scaling from 8 to 128 experts at about 3.2B active parameters cost under 5 percent throughput. The framework will underpin Ai2’s next Olmo MoE model, which has not yet been released. Source
2. ServiceNow Details AutoSynthData for Targeted Enterprise Agent Training Data
Hugging Face. ServiceNow’s CoreAI team described AutoSynthData on the Hugging Face blog, a pipeline that compares a target model against a stronger teacher on diagnostic tasks, writes a specification of the capability gaps, and generates tasks that vary entities, environment state, workflow composition, and difficulty. Each candidate passes positive and negative verification, with repair loops and batch-level diversity review. Fine-tuning Gemma-4-26B-A4B-it on 2,000 synthetic samples (about 18 hours to generate) raised Pass@1 by 7.2 points in a hybrid domain, closing 59 percent of the gap to the reference model, while an ITSM run lifted Pass@1 from 18.77 to 27.18 percent; the EnterpriseOps-Gym dataset is available on the Hub. Source