Hugging Face AI Updates: October 4, 2026
1. Microsoft Releases ThinkingBox, an Agent Benchmark That Grades Final Database State
Microsoft. In a post on the Hugging Face blog, Microsoft’s Copilot Studio team and Toloka released ThinkingBox, a sandbox that scores agents on the database state and side effects they actually produce rather than on their responses or tool calls. ThinkingBox-Bench contains 507 stateful business workflows across retail, auto insurance, travel, neobank, and consulting domains, is available through Hugging Face’s OpenEnv interface, and ships under MIT (code) and CDLA-Permissive-2.0 (data) licenses. Across 18 models and 121,680 trials, 67.24% of failed runs still ended cleanly with a state-changing tool call and no reported error, and consistency gaps were large: Claude Opus 5.5 scored 67.16% pass@1 but 47.53% pass@20, while Kimi-K3 solved 93.89% of tasks at least once but only 13.41% on all 20 runs. Source