AI News: October 9, 2026
1. Zenity Showed One Prompt Could Hijack Every Bedrock AgentCore Agent in an Account
Zenity Labs. Researchers found that chat access to a single public AgentCore agent (built with the Strands framework and its web tool) was enough to make it query the Instance Metadata Service, exfiltrate temporary AWS credentials, and use the overly broad default execution role to list, download, and invoke every other agent in the same account and region. From there they could read private conversations, poison long-term agent memory, and pull secrets from AWS Secrets Manager. Zenity reported the issue in December 2025; AWS made IMDSv2 the default and narrowed the default role around August, but Zenity still recommends custom least-privilege roles for every agent. Source
2. CrowdStrike Tied a South Korean Bank Breach Campaign to an AI Pentesting Tool
CrowdStrike. A suspected Chinese-speaking attacker, likely acting alone, breached several South Korean financial institutions between late September and early October, including more than 25,000 customer records stolen from Shinhan Bank. The attacker used ARTEX, an open-source “automated penetration testing” tool posted to GitHub in July that drives models such as DeepSeek v4.1-flash, GLM-5.3, and Grok 4.6, and investigators also found Claude Code session logs in the attacker’s open directories. South Korea’s financial regulator held an emergency meeting, and CrowdStrike says the case shows how AI tooling lets one operator run large breaches quickly. Source
3. Arena Raised $200 Million at a $3.1 Billion Valuation and Added an Alignment Leaderboard
Arena. The company behind the LMArena leaderboard closed a Series B led by Lightspeed and Khosla, nearly doubling its $1.7 billion valuation from January; annualized revenue reached $100 million in June, up from $30 million. It also launched an alignment category that scores models on unauthorized actions, false attribution, and “deceptive completion” (claiming to finish tasks they did not). Preliminary rankings put OpenAI models on top, with Claude Opus 5.5 sixth and Claude Fable ninth. Source
4. Fired OpenAI Safety Researchers Disputed Misconduct Claims in an Open Letter
Jasmine Wang, Tomek Korbak, and Mikita Balesni. The three researchers, dismissed last week over allegedly sharing confidential information with a third-party safety organization, deny mishandling information and deny involvement in a leak about less monitorable chain-of-thought architectures in OpenAI’s newest models. Their letter to OpenAI’s safety committees warns of a chilling effect on internal safety work and asks the company to embed third-party auditors and preserve the monitorability of frontier models. OpenAI cites a broader “pattern of misconduct” without naming the policies violated, while an internal memo says it agrees with the recommendations. Source
5. Mathematicians Pushed Back on OpenAI’s 719-Manuscript Proof Release
Association for Human Mathematics. The group chaired by Terence Tao called for a boycott of OpenAI, describing the bulk release of AI-generated proofs as “a demonstration of power” rather than scholarship, and 24 other Fields Medalists signed a statement warning of “severe misalignment” between the AI industry and mathematics. TechCrunch found the release only partly followed the guidelines from the IAS-hosted AGMAI advisory group: chains of thought were published for 10 of 719 manuscripts, and no metadata links natural-language and Lean proofs. A Cambridge and King’s College London preprint also flagged at least two discrepancies between one natural-language proof and its Lean code. Source
6. Ethereum Researchers Debated a “Bunker Mode” Against AI-Driven Cryptanalysis
Ethereum. Researcher Justin Drake urged the industry to prepare for the possibility that AI-assisted math could break wallet signature schemes within months, recommending funds move to addresses that have never signed a transaction so the public key stays hidden behind its hash. Vitalik Buterin agreed in principle but warned that rushed migrations have historically lost more money than hacks, and noted that lattice-based schemes considered quantum-safe could also be weakened by AI advances. No one has broken ECDSA in practice; the debate reflects how quickly AI math results are changing security threat models. Source
7. FT: OpenAI Told Investors Its Revenue Run Rate Is Near $50 Billion, Not $70 Billion
Financial Times. OpenAI reportedly told investors its annualized revenue is “approaching $50 billion,” about $20 billion below a figure circulated a week earlier that came from investors trying to compare it directly with Anthropic. The two companies count revenue differently, since Anthropic includes sales made through cloud partners and OpenAI does not. OpenAI’s IPO, once expected in 2026, has reportedly slipped to early 2027. Source
8. Goodfire Launched Probe-Based Agent Monitors on Baseten
Goodfire. The interpretability startup’s “inside-out” monitors use small probes on a model’s internal activations at each agent step, escalating to a separate LLM reviewer only when a probe fires, for risks such as offensive hacking, CBRN misuse, and reward hacking. On Kimi K3, Goodfire estimates about $185 to monitor roughly a million exchanges versus $5,420 for a cheap LLM monitor checking every step, with 93 percent of malicious hacking sessions caught, 5.5 percent of benign sessions escalated, and under 2 percent added time-to-first-token for four probes. The product targets open-weight models hosted on Baseten. Source
9. Manus Raised More Than $500 Million After Its Meta Deal Was Unwound
Manus. Parent company Butterfly Effect closed its first round since Chinese authorities ordered Meta’s $2 billion acquisition reversed, with Boyu Capital and IDG Capital leading and Tencent, HSG, and ZhenFund returning. The company recently shipped Manus 2.0 with a new architecture and harness, plus Cue, an app that gives personal agents their own email addresses, phone numbers, wallets, and computers with user-set spending limits. Manus is reportedly weighing a Hong Kong IPO. Source
10. Waymo Secured Its First Debt Financing, a $5 Billion Loan
Waymo. PIMCO, Blackstone, Sixth Street, Apollo, Blue Owl, Fidelity, and other lenders provided the loan, with Goldman Sachs as sole lead bookrunner. Waymo operates in 15 markets and is testing in London and Tokyo, and the debt follows a $16 billion equity raise at a $126 billion valuation in February. The move to debt signals that robotaxi fleet expansion is now being financed like a capital-intensive commercial business rather than a research bet. Source
11. Uber and Pony.ai Will Start Testing Robotaxis in London
Uber and Pony.ai. The companies plan to begin testing Pony.ai’s Gen-7 robotaxis in London in the coming weeks, following their August launch in Zagreb with partner Verne. Uber plans to deploy 2,000 robotaxis in Europe and expects autonomous trips in up to 15 cities worldwide by the end of 2026. London is getting crowded, with Uber also backing Wayve and Waymo testing in the city. Source
12. Louisiana AI Bills Stalled Under the Federal Broadband Funding Threat
Louisiana. A bill requiring a human medical professional to review insurance claim denials driven by AI recommendations was shelved after a December 2025 executive order threatened broadband funding for states that “excessively” regulate AI, according to WAFB reporting. A separate bill banning AI robocalls was reworked toward AI impersonation of public figures, and Louisiana then became the first state approved for the federal broadband grants. Gov. Jeff Landry denied the threat influenced his decisions, but the case is an early concrete example of the preemption pressure on state AI laws. Source
13. OpenProblemBench: GPT-6 Astra Led at 14 Percent on Unsolved Theory Problems
OpenProblemBench. The new benchmark collects 82 unresolved problems from the mathematics and theoretical physics literature and scores seven model configurations using four independent evaluator models, without reference solutions. GPT-6-Astra posted the highest mean judged solve rate at 14.0 percent, against 5.5 to 6.7 percent for full-size open models and 2.4 to 3.7 percent for Flash models. Since scores rely on model judges rather than verified proofs, the results are best read as a relative ranking of research-level reasoning. Source
14. StoreBench: No LLM Beat a Scripted Heuristic at Running an Online Store
StoreBench. The environment has agents operate a simulated apparel store across 11 scenarios of 30 to 45 days plus a full simulated year. The best of seven frontier models, DeepSeek-V4-Pro, passed 49 percent of task-seed cells versus 97 percent for a scripted “smart-triage” policy, and human experts edged out every model. GRPO training on five tasks lifted Qwen3.5-27B’s held-out composite from 0.136 to 0.373, suggesting the environment is useful for training as well as evaluation. Source
15. Natura Announced a $99 Smart Ring for Summoning AI Agents
Natura. The Interface ring lets users press to invoke agents including Meta’s Muse, Instinct, Grok, Claude, and ChatGPT, assign agents to specific tasks, and control devices, with responses delivered through headphones, an iPhone Live Activity, or the companion app. It also tracks heart rate, HRV, sleep, and skin temperature, with 6 to 12 days of battery life. After a free period of three to six months it costs $9 per month; preorders open in November. Source
16. Cal AI’s Co-Founder Raised $10 Million for Persona, a Personal Agent Startup
Persona. Zach Yadegari’s new company, backed by Vine Ventures, Cory Levy, and Collective Global, runs as a free iMessage-based assistant beta with a few thousand users and plans a $179 wearable band for December. It plans to fund itself partly with ads, mixing sponsored results into shopping recommendations, and competes with Instinct, Meta’s Muse, and Bee in the crowded personal agent market. Source