Architecture AI Updates: October 6, 2026
1. Akka Measured Spec-Driven AI Porting Across 65 Open Source Projects
InfoQ. Akka ran a discovery, specification, porting, benchmarking, and improvement loop over 65 open source projects using Claude with Akka Specify, consuming 99.3 hours and 9.41 billion tokens in the first tranche, with 57 of 65 ports showing code size or performance improvements. Sonnet averaged 61 minutes per port while Opus averaged 120 minutes but used roughly 40% fewer tokens, and higher effort settings raised consumption without consistently improving efficiency. Structured specs with claims, evidence, and typed behavior improved first-pass results, though cross-component decisions remained the main gap and each added exit condition (auditors for serialization, PII, idempotency, architectural boundaries) increased porting cost. Source
2. Treating Hallucination Rate as a Platform Metric, Not a Model Property
InfoQ. Aditya Mulik describes a multi-team LLM platform for retail inventory recommendations that cut hallucination rates from 15% to 1.5% over six months without changing foundation models. The design centers on a gateway with per-team rate limits and cost attribution, a versioned prompt registry resolved at runtime for instant rollback, Pydantic schema validation with retries classified by failure type (3-attempt budget), an intent classifier that returns “unclassified” instead of forcing a route, default-deny team-partitioned MCP tool servers, and a golden-set gate requiring a 0.8 tool-trajectory score before prompt promotion. The author argues these primitives should exist before the second application, not the tenth. Source
3. Elastic Consolidated Siloed Agent Evals Into a Shared Framework
InfoQ. In a presentation, Elastic principal data scientist Susan Chang explains how the company replaced per-team agent evaluations with shared LLM-as-judge, deterministic rule, trace-based (tokens, latency, tool calls), and RAG evaluators. LLM judges alone missed issues like hallucinated product IDs, so they are paired with programmatic syntax and factuality checks, and evaluators were moved from Python to TypeScript (via Playwright and an internal framework called Scout) to test production agent code directly. Chang recommends starting with 20 to 50 records, calibrating judges against human ratings, and tracing intermediate steps such as vector searches rather than only final outputs. Source
4. Claude Cowork Moves Both Inference and the Agent VM to the Cloud
Simon Willison. Willison highlighted Anthropic’s Felix Rieseberg explaining that the new Claude Cowork runs model inference and the tool-execution VM in the cloud, with each session in its own sandbox that shares no state with other sessions; the desktop app brokers access when the VM needs local files. The earlier design ran a local VM alongside cloud inference, which drew complaints about disk usage, battery drain, and work stopping when a laptop closed. The shift is a concrete example of the trade-off between local and hosted agent sandboxes, and it enables Cowork on phones. Source