Daily News · 5 min read

AWS AI Updates: October 2, 2026

1. AWS Well-Architected Agent Enters Preview as the Successor to Trusted Advisor

AWS. AWS launched a preview of the Well-Architected Agent, an AI service it describes as the next-generation evolution of AWS Trusted Advisor and the Well-Architected Tool. The agent correlates metrics and application topology against Well-Architected best practices across cost, security, performance, and reliability, ranks recommendations by impact and effort against business goals the team sets, and shows cross-pillar effects (for example, the cost and performance impact of adding multi-AZ failover) before a change is made. Fixes ship as SSM runbooks, CLI scripts, or console walkthroughs, and the agent can review Terraform, CloudFormation, and CDK templates and return the IaC changes needed. It is delivered by AWS Support to customers with a Support plan, accessible in N. Virginia, Ohio, and Oregon, with workloads onboardable from any commercial Region. Source

2. Amazon Quick Apps Can Now Query Governed Datasets Live

AWS. Apps that the Amazon Quick app-builder agent writes from a natural-language description can now query Quick Sight SPICE and Direct Query datasets live, instead of baking in a snapshot at build time. The agent finds the relevant curated datasets, writes the SQL while building the app, and the published app re-runs that SQL each time it opens, executing as the viewing user so existing row-level and column-level security rules apply automatically. AWS says consent is enforced server-side on every query, and the feature is available in all Regions that support Quick apps. Source

3. S3 Vectors Plugged In as a Memory Backend for NVIDIA NeMo Agent Toolkit

AWS. A new walkthrough implements Amazon S3 Vectors as a custom memory provider for NVIDIA’s open-source NeMo Agent Toolkit (tested on version 1.6), alongside its built-in Mem0, MemMachine, Redis, and Zep backends. The plugin implements NAT’s MemoryEditor interface (add_items, search, remove_items), embeds memories with Titan Text Embeddings V2 into a 1024-dimension cosine index, and tags each vector with agent, team, task, and memory-type metadata for scoped retrieval. Agents run as nat serve containers on EKS with IRSA for S3 Vectors access, relying on strong write consistency so every pod sees new memories immediately. AWS notes its claimed gains in groundedness and token usage are directional expectations, not benchmarked results. Source

4. Reference Sample Builds Event-Triggered Ambient Agents on AgentCore

AWS. AWS published a CDK reference implementation for ambient agents, which are triggered by S3 uploads or scheduled events rather than chat prompts and run as jobs on AgentCore Runtime. A single ask_human tool with a canonical response envelope covers approval, clarification, and review steps, letting the agent pause and resume from the same point once a person answers. The stack is serverless, using Lambda for event processing, DynamoDB for state, SQS, API Gateway, Cognito, and a React jobs UI, with Claude Sonnet 4.5 as the default model and each agent turn capped by the 15-minute Lambda timeout. Source

5. Guide Details Three Authentication Paths for Claude Platform on AWS

AWS. A step-by-step guide puts the Claude Platform on AWS subscription in a dedicated AI Services account with separate production and development workspaces, then wires up three access paths. AWS workloads assume a cross-account role and make SigV4-signed calls with no stored keys, developers use a long-lived API key scoped to the development workspace with the standard Anthropic SDK, and workloads outside AWS federate through OIDC to get temporary credentials and short-lived tokens. The post notes that the workspace Region sets the API endpoint (for example aws-external-anthropic.us-east-1.api.aws) but not where inference runs, which is controlled separately (“US” or “Global routing”), and that short-term tokens only work against the Regional endpoint that issued them. Source

6. uniopen Lifts Nova 2 Lite Moderation F1 From 0.59 to 0.86 With SageMaker Fine-Tuning

AWS. Taiwanese retail platform uniopen fine-tuned Amazon Nova 2 Lite with LoRA in SageMaker AI on 3,391 conversation windows to classify nine behavior categories and three subject types. On a 737-window held-out set, per-behavior macro F1 rose from 0.5852 to 0.8364 and subject-type macro F1 from 0.4162 to 0.8302; switching the output from JSON to a line-based format, with no further training, then pushed them to 0.8550 and 0.8491, clearing production targets of 0.85 and 0.82. Nova 2 Pro drafts corrections for reported errors, but each must be human-verified before entering training data, and Argo Workflows on EKS gates promotion with hard regression checks and soft warnings that require admin approval. Source

7. AWS Professional Services Shares a Four-Agent Migration Pattern on AgentCore

AWS. AWS Professional Services described a four-agent pattern, built with Strands Agents on AgentCore and connected to internal systems through MCP tools exposed by AgentCore Gateway, that ran alongside AWS Transform and AWS DMS on an enterprise program covering more than 300 applications. An Intake Agent reads migration inputs from wikis, ticketing, and collaboration tools; an IaC Agent composes approved internal modules; a governance agent reports inside the program’s own tools; and an SRE agent handles operations after cutover. AWS says the IaC agent cut infrastructure-as-code development from 3 to 4 weeks per application to minutes, based on internal project tracking data. Source

8. Strands Labs Open-Sources Strands Decider 2B, a Local Decision Model for Agents

AWS. AWS’s Strands Labs released Strands Decider 2B, a decision model that swaps the language head of Qwen 3.5-2B for one that scores developer-supplied answer options and returns calibrated confidence instead of free text. It runs locally at roughly 115 ms median on an RTX 3090 and about 153 ms on an M3 MacBook, and ranks third of 33 models in the 2B class on JevBench’s public set. Weights are on Hugging Face and the full training data, code, and scripts are on GitHub; in Strands Agents it is used as a pre-tool-call check for whether arguments are grounded in the user’s request, as well as for routing, tool selection, and guardrails. Source