Daily News
· 1 min read
Apple AI Updates: September 30, 2026
1. Apple Found Most LLM Round-Trip Failures Start When the Model Writes the Problem, Not When It Reads It
Apple. In a NeurIPS paper, Apple researchers including Samy Bengio had one model turn procedurally generated arithmetic expression trees into word problems and another model extract the expression back, checking symbolic equivalence across all pairings of 16 models. Swapping which model generates and which extracts shifted accuracy by up to 60.4 points, the best pair reached 92.9 percent, and at least 73.6 percent of failures originated at generation, with tree depth and branching driving difficulty. Fine-tuning on about 3,600 examples lifted every open-weight model above an untrained Gemini-3.1-Pro, which is relevant for multi-agent pipelines that pass structured data through natural language. Source