AWS AI Updates: October 3, 2026
1. SageMaker AI Multi-Turn RL Cuts a Search Agent’s BrowseComp-Plus Failure Rate From 22.89% to 0.68%
Amazon SageMaker AI. AWS researchers fine-tuned Qwen3.6-27B as a search agent with BM25 and vector search tools using SageMaker AI multi-turn reinforcement learning (MTRL), a serverless, per-token-priced service that optimizes over full multi-turn trajectories with PPO, CISPO, or importance-sampling losses and GRPO or RLOO advantage estimators. The reward was trajectory-level nDCG@10, with a -1 penalty for hitting the turn or token limit, and only three hyperparameters (max_epochs, global_batch_size, rollout_max_concurrency) were changed from defaults in the MultiTurnRLTrainer SDK. On held-out sets, nDCG@10 rose 23.7% on BrowseComp-Plus (0.5136 to 0.6354), 18.4% on WixQA, and 6% on Wands, with a slight regression on FreshStack (0.4112 to 0.4089). Jobs default to a 24-hour limit and can resume from checkpoints, and the setup ran in US West (Oregon). Source
2. AgentCore Gateway Accepts Private CA Certificates for VPC Targets
Amazon Bedrock AgentCore. AgentCore Gateway now supports TLS certificates signed by private certificate authorities on MCP server, OpenAPI, and HTTP proxy (passthrough) targets, so gateways can connect natively to private endpoints in a VPC without an intermediate Application Load Balancer. Customers register a PEM-encoded CA certificate for targets using private endpoints powered by Amazon VPC Lattice, and the gateway fetches it from Amazon S3 or AWS Secrets Manager to use as the trust anchor for outbound TLS. The feature is available in all Regions where both AgentCore Gateway and VPC Lattice are offered. Source
3. Walkthrough Connects Claude Desktop to AgentCore Web Search Through IAM Identity Center
Amazon Bedrock AgentCore. A new guide wires Claude Desktop, running with Bedrock as its inference provider, to AgentCore’s managed MCP-compatible Web Search target, which AWS says is backed by an Amazon web index spanning tens of billions of documents with no external API keys and query traffic kept inside AWS. Authentication chains IAM Identity Center (SAML) into an Amazon Cognito user pool that issues JWTs via the OAuth 2.0 authorization code flow, which the gateway validates on each request. Claude Desktop discovers the WebSearchTool through MCP tools/list and prompts users for approval before each search. Web Search on AgentCore is currently available in US East (N. Virginia), Europe (Ireland), and Asia Pacific (Tokyo). Source
4. “Adjudicated Query” Pattern Keeps LLMs Out of Compliance Decisions in Amazon Quick
Amazon Quick. AWS published a reference architecture in which an Amazon Quick chat agent only translates questions into calls on six fixed MCP tools (such as sweep_compliance and simulate_rule_change) and narrates results, while a deterministic rules engine with versioned rules makes every pass/fail determination. Each sweep must produce a completeness receipt asserting that compliant, in-breach, ambiguous, and unreadable counts sum to the number scanned before results persist, and no natural language reaches SQL. The stack runs the MCP server on Lambda behind API Gateway with a Cognito JWT authorizer and stores everything in Aurora Serverless v2 with pgvector; Bedrock (Titan Text Embeddings V2 and Claude Sonnet 5) is used only for exploratory clause search. The post also describes guarding against the chat model stripping caveats or extrapolating from sample rows by repeating un-strippable labels at several payload levels. Source
5. Amazon Payments Runs a Multi-Objective Contextual Bandit on SageMaker AI to Pick Generated Content
Amazon SageMaker AI. Amazon Payments described using a multi-objective LinUCB contextual bandit on SageMaker AI to choose which generative AI content variant to show each visitor in a product acquisition funnel. One LinUCB model runs per funnel stage (application start, submission, approval), and their UCB scores are combined with roughly equal weights to avoid optimizing one stage at the expense of another. In a seven-week online A/B test, one customer population showed a high single-digit relative lift in final-funnel conversion while another showed no improvement, which the team attributes to the content rather than the model. A code repository with synthetic data accompanies the post. Source