AI News: September 14, 2026
1. AllSpark Released Iris-mini and Iris-pro, the Strongest Open-Weight Search Agents in Their Size Classes
AllSpark. The lab published two open-weight search agents on Hugging Face: Iris-mini, a 35B model built on Qwen3.6-35B-A3B, and Iris-pro, a 397B model built on Qwen3.5-397B-A17B. With context management enabled, Iris-pro scores 88.6 on BrowseComp, 85.1 on BrowseComp-ZH, 92.9 on DeepSearchQA, and 56.4 on Humanity’s Last Exam; Iris-mini scores 82.2, 84.8, 86.9, and 52.3 respectively, leading its size class on three of the four and beating the next-best model of its size on BrowseComp by 3.4 points. Only weights and inference code are out so far, with the data construction and training pipelines promised later. Source
2. GPT-6 Astra Tripled Claude Fable 5.1’s Take on Vending-Bench and Cleared Every Drone-Bench Task
Andon Labs. In the lab’s vending machine business simulation, Astra averaged $15,515 across six runs against $5,422 for Claude Fable 5.1, with every Astra run beating every Fable run. The gap came from negotiation discipline: Astra held supplier pricing steady where Fable’s costs drifted, and Astra avoided defunct suppliers that cost Fable $14,331. On Drone-Bench, Astra is the first model whose best submissions beat the human baseline on all five tasks including reconstruction, and it autonomously flew a drone to identify and track a specific person through an office. Reliability is the caveat: chaining all five subtasks in sequence succeeded only 2.8% of the time, with individual task success rates between one and four runs in ten. Source
3. Musk and Hassabis Joined Altman in Backing Independent Evaluators Inside AI Labs
Industry. Following Dario Amodei’s proposal for embedded third-party evaluators, Elon Musk endorsed independent evaluation requirements and Demis Hassabis partially backed the oversight portion of the plan, putting the heads of four frontier labs on record supporting some form of external audit. Altman separately told Fortune that OpenAI will not go public in 2026, a decision he says was communicated internally in June. Google researcher Peyman Milanfar pushed back on the premise, arguing that recursive self-improvement is inherently unstable and self-limiting, which would make imposed speed limits unnecessary. Source
4. A Two-Year Randomized Trial Found the AI-Free Student Group Finished Last Both Years
Vrije Universiteit Amsterdam. Thibault Schrepel ran a two-year trial splitting students into three conditions, no AI, unguided AI, and structured AI training, on legal revision tasks covering EU AI Act provisions, scored on substance, clarity, proportionality, and innovation. The cohorts covered 66 students in 2024 and 164 in 2025. The AI-free group came last in both years and hit idea exhaustion after 10 to 15 minutes; the untrained AI group accepted suggestions uncritically but still outperformed it; the trained group led in 2024, though the gap nearly closed by 2025 as baseline AI familiarity rose. Schrepel, who started out favoring restrictions, wrote “I was wrong” and now recommends faculty training over blanket bans. He flags the small, self-selected sample and the inability to verify AI use on take-home exams. Source
5. Obama Told Democrats to Make AI a Central Agenda Item
US politics. Speaking at a Democratic fundraising event in an interview with House Minority Leader Hakeem Jeffries, Barack Obama said the party needs “a very clear plan” for AI’s economic and safety effects and should build a public framework for the debate if it regains the House. He characterized the technology as moving very fast in private hands and potentially dangerous without oversight, while pointing to accelerated drug development as an upside. No specific legislative proposals accompanied the remarks. Source
6. Meta Is Rebuilding the Management Layer It Cut From Applied AI
Meta. After moving roughly 7,000 employees into the Applied AI organization and reassigning many former managers as individual contributors, Meta is now asking workers in that division to volunteer to return to management. Reporting attributes the reversal to depleted teams, lower productivity, and a 40% increase in technical incidents under the flatter structure. The change appears limited to Applied AI rather than a company-wide retreat from Zuckerberg’s manager-light push, and lands as Meta tracks toward more than $130 billion in AI chip and infrastructure spending this year. Source
7. A Satirical Post on Frontier Pacing Topped Hacker News With 757 Points
Xe Iaso. The post announces that the industry must halt all frontier model research so that a fictional lab, Techaro’s Lygma AGI lab, can catch up and dominate the world with its Intelliga models. It reads the current round of pacing proposals as competitive positioning wearing safety language, and it reached 757 points and 439 comments on Hacker News, making it the most-discussed item of the week on the topic among developers. Source