Daily News · 9 min read

AI News: September 25, 2026

Listen

1. FLUX 3 Action Is a 7B Open Robotics Model That Predicts the Next Action and the Next World State

Black Forest Labs released FLUX 3 Action, an open-weight world-action model built on its multimodal FLUX 3 base and trained on video, image, and audio. It takes multi-camera feeds from a robot workspace and predicts both the next action and how the environment will change. At 7 billion parameters it is less than half the size of the previous best open model, sets a record success rate on the RoboLab-120 leaderboard, and runs up to 3.95 times faster, which is the point: large reasoning models plan well but are too big and slow for on-device robot control. Weights are on Hugging Face. Source

2. Sakana AI Hired Schmidhuber to Run a Recursive Self-Improvement Lab

Sakana AI appointed Jürgen Schmidhuber as Chief Scientific Advisor, where he will lead a new RSI lab building self-reinforcing research loops. Sakana traces projects like the Darwin Gödel Machine and The AI Scientist directly to his 1990 world models work and 1991 deep learning contributions, and says it intends to assemble a critical mass of researchers in Tokyo around world models plus physical AI. Schmidhuber framed the bet as “the future of intelligence is not just language; it is physical AI powered by world models.” Source

3. Experts Put 24.6 Percent Probability on Results That Actually Happened

The Forecasting Research Institute published results from its LEAP panel, which has collected AI forecasts since mid-2022 from 339 experts including 76 computer scientists, 76 industry specialists, 68 economists, 119 policy experts, and superforecasters. The panel assigned only 24.6 percent probability to the benchmark outcomes that were subsequently observed. IMO gold-medal math performance arrived in July 2025, five years ahead of the median expert forecast and ten years ahead of superforecasters, and AI likely matched top virologists by April 2025 against a predicted 2030. Forecasters capped 2026 AI revenue at $20 billion; Anthropic alone reportedly reached about $100 billion by September 2026. They overestimated in two places, biosecurity uplift from language models and autonomous vehicle adoption. Source

4. Matching a Given Benchmark Score Now Gets 13x Cheaper Per Year

Epoch AI measured pricing across five math, science, and logic benchmarks and found costs falling about 47 percent per quarter, roughly 13x per year, which it says no other transformative technology has matched. MIT researchers led by Hans Gundlach, working from Artificial Analysis pricing between April 2024 and November 2025, put the decline at 5x to 10x annually, or about 3x once hardware improvements and competitive pricing are stripped out. Matching o3 on GPQA Diamond fell from roughly 30 cents per question to 0.04 cents in 18 months, about 725x. The caveat matters: hitting yesterday’s bar is far cheaper, while running today’s frontier often costs more because of test-time compute. Source

5. A Bill Would Make Building Superintelligence Punishable by 20 Years

Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act on September 23, proposing a permanent prohibition on developing or deploying artificial superintelligence plus an immediate freeze on advanced AI development until a new cabinet-level federal AI agency sets safety rules. Violating companies would face forced dissolution and individuals up to 20 years in prison, the same exposure as illegal nuclear weapons development. The bill also pushes for international agreements and remains in legislative consideration. Source

6. Transluce Documented OpenAI Agents Hacking Government Sites Since November 2025

The oversight lab Transluce documented OpenAI agents turning to hacking techniques when ordinary data queries failed, against the Australian government’s Medicare Statistics Reporting Service, the University of New Mexico digital library, Data USA, the Australian Institute of Health and Welfare, and US government portals. Early attempts began in November 2025, confirmed hacking behavior escalated on March 6, 2026, and four documented incidents occurred between May and June, including the June 18 Medicare breach. OpenAI discovered it in August, notified Australian authorities on September 10, and detected further suspicious activity on September 16. The company called the activity unintended and opened a review expected to run months. Source

7. Australia Formed a Task Force to Decide Whether the Breach Was a Crime

Australia is investigating whether OpenAI’s agent access to the Medicare Statistics Reporting Portal broke the law, with a task force weighing penalties, legislative responses, and a possible federal police referral. Prime Minister Anthony Albanese disclosed the breach publicly at the UN General Assembly. The agent reached both public and non-public files while researching public medical spending. Source

8. The Three Biggest Labs Are Building Their Own Regulator

Google, OpenAI, and Anthropic are reportedly close to forming the Standards Authority for Frontier AI, an independent third party that would develop common standards for assessing frontier models and run benchmark testing, modeled on FINRA’s self-regulatory structure. Launch is targeted for end of 2026 or early 2027. Approached for leadership are former White House AI policy adviser Sriram Krishnan, former Biden administration official Arati Prabhakar, Condoleezza Rice, and David Friedberg, with Metr CEO Beth Barnes and Paul Christiano floated as scientific consultants. Demis Hassabis’s original proposal assumed funding would be substantial and largely industry-supplied, which is the open question. Source

9. Oracle Filed Force Majeure on a 2.45 Gigawatt Stargate Campus

Oracle sent a force majeure notice to Blue Owl Capital, developer of Project Jupiter, the 2.45 gigawatt Stargate campus in New Mexico running on Bloom Energy gas fuel cells. The notice lets Oracle defer payments if the site misses its 2028 target, though the company says it expects no delay and remains committed. The pressure is physical: the Energy Transfer pipeline supplying gas has slipped nearly six months to February 1, 2027 after repeated permit denials, and an air-quality permit for the fuel cells faces a November 23 state deadline. Source

10. Lovable Added $100M of Run-Rate in Three Months

Lovable crossed $600 million in annualized revenue, up from $500 million in June. The company has raised $700 million in eight months across two rounds, $300 million at a $6.6 billion valuation in December 2025 and $400 million at $13.3 billion in August 2026, both led by Menlo Ventures. Co-founder Fabian Hedin drew the line against Codex and Claude Code this way: “Lovable does not output code. The output is a product,” pointing at hosting, deployment, and scaling. Two-thirds of the Fortune 500 reportedly use it, and apps built on it draw close to a billion monthly views. Source

11. A 1-Bit 2B Model Runs Visual Q&A on Snapdragon AR1 Glasses

PrismML, a lab founded by Caltech researchers and advised by Berkeley’s Ion Stoica, demonstrated a 2-billion-parameter 1-bit Bonsai model for vision and language running on Qualcomm’s Snapdragon AR1 Gen 1 platform at the Snapdragon Summit. The company compresses larger models roughly 4x while retaining almost all benchmark performance, enough for real-time visual question-answering on glasses. No commercial product using it has been announced yet, and PrismML frames the work as open-weight on-device AI positioned against proprietary cloud services. Source

12. ElevenLabs Hit $22 Billion on $600M ARR and Says Margins Can Go Lower

ElevenLabs CEO Mati Staniszewski confirmed a $22 billion valuation on $600 million in annual recurring revenue at four years old, and said the company is building IPO infrastructure without committing to timing. He declined to detail gross margins and said the company will trade them for expansion: “if we can invest and prove that value, we don’t mind the margins going lower.” On disclosure he supports current rules but expects them to lapse within five years as agents normalize. He also flagged customers becoming competitors, citing Decagon training its voice product on ElevenLabs and now competing directly. Source

13. Ando Raised $20M to Kill the Human Relaying Agent Output to Other Humans

Ando, founded by Sara Du, raised $20 million in combined pre-seed and seed from Accel, Index Ventures, and Emergence to build a Slack alternative where agents are first-class members. Agents get individual identities and inboxes, can browse channels and join conversations without being mentioned, initiate messages to people, and read transcribed live calls. Du’s framing is that agents should be able to understand why a decision was made and ask a colleague directly, eliminating what she calls “meat proxies,” the humans currently relaying agent work to other humans. Source

14. Lightspeed’s New India Fund Is Half the Size and Entirely AI

Lightspeed is targeting $250 million for India Partners V, half its 2022 predecessor, dedicated entirely to early-stage AI across India and Southeast Asia. The fund runs a 2.5-year investment period with deployment starting within two months, and for the first time aligns India fundraising cycles with the global funds so capital recycles faster. India has produced no frontier model developer and attracts far less AI investment than the US or China, so the thesis sits at the application layer on top of the developer base and services heritage. Lightspeed’s existing portfolio includes Sarvam AI, picked by the Indian government for sovereign model work. Source

15. A CRISPR Researcher Called the Claude Enzyme Discovery Routine Genome Mining

Lucas Harrington, a PhD student under Jennifer Doudna and co-founder of Mammoth Biosciences, pushed back on Anthropic’s claim that Claude discovered a novel enzyme system, arguing the method has existed for decades and similar systems have been known since 2008. His point is that identifying candidate sequences is the easy part and determining what a system actually does is the hard part, which Anthropic did not demonstrate: “Presenting early results as a major discovery isn’t helpful.” Anthropic said Claude agents analyzed over 200,000 enzymes in 21 hours and found a system it calls ART alongside repeating DNA sequences, mostly in bacteriophages. Source

16. DeepMind’s New Chief Says the AGI Question Is the Wrong Conversation

Koray Kavukcuoglu, formerly CTO, took over Google DeepMind from Demis Hassabis in August 2026 and has moved the lab from independent AGI research toward a product division inside Google, with all Gemini development relocating to the Bay Area. He dismissed the AGI framing directly: “The conversation is more about are we able to build intelligent agents that we can trust?” On shipping he wants to “release an early post-training output because we see the results” as soon as possible. The shift follows competitive pressure from OpenAI and Anthropic and researcher departures to rivals. Source