Daily News · 4 min read

AWS AI Updates: October 1, 2026

1. GPT-6 Astra UltraFast Mode Comes to Bedrock

AWS. Amazon Bedrock now offers UltraFast mode for OpenAI’s GPT-6 Astra, a premium speed tier that OpenAI says delivers up to 6x faster inference in the API at up to 300 tokens per second. AWS targets it at latency-sensitive uses such as real-time coding assistants, interactive agents, and customer-facing experiences. The tier arrives one day after OpenAI announced it at DevDay, and is available through the Bedrock console and supported Bedrock APIs, with Region, inference profile, and pricing details in the Bedrock documentation. Source

AWS. Amazon S3 Vectors indexes can now resolve metadata filters before running similarity search, which AWS says returns up to 5x more of the matching vectors on highly selective filters, such as a single tenant or client in a large multi-tenant RAG store. The release also adds a $startsWith prefix operator for paths, URLs, and hierarchical IDs; each vector carries up to 2 KB of filterable metadata and a query supports up to 100 filter constraints. New indexes use the ENHANCED index mode by default, while existing CLASSIC indexes can switch in place with the UpdateIndexMode API without re-ingestion, and a per-query parameter lets teams compare both modes first. There is no additional cost. Source

3. Aurora Serverless Scales in Bigger Steps for Bursty Agent Workloads

AWS. Aurora serverless can now add up to 16 ACUs to its current capacity within a second and keep scaling up to 256 ACUs, then scale back to zero when the workload finishes. AWS positions the change for agentic AI applications, which tend to have bursts of activity, long idle periods, and unpredictable traffic. It is on by default for clusters on platform version 3 or 4; clusters on versions 1 and 2 need to upgrade to version 4, which can be checked via the ServerlessV2PlatformVersion parameter. Source

4. AWS CLI Can Now Check and Update All Agent Toolkit Skills at Once

AWS. AWS CLI 2.37.0 adds aws agent-toolkit check-skill-updates, which compares installed agent skills against the registry, and aws agent-toolkit update-skill --all, which updates every outdated skill in one command. The Agent Toolkit for AWS bundles the AWS MCP Server (an agent interface to more than 15,000 AWS APIs), agent skills, and plugins, and the CLI already installs and configures them for Kiro, Claude Code, Codex, Cursor, and other coding agents. Previously each skill had to be checked and updated individually. Source

5. AWS Marketplace Ships an Agent Skill That Builds SaaS Metering Integrations

AWS. A generally available AWS Marketplace metering agent skill guides sellers through building a usage-based SaaS metering integration from inside an AI coding assistant. It gathers product type, pricing model, and usage dimensions, generates integration code and a CloudFormation stack, deploys a serverless pipeline using the ResolveCustomer and BatchMeterUsage APIs plus EventBridge subscription events, and runs a live end-to-end test before production. The skill cross-validates dimension names against the seller’s product configuration to catch mistakes that otherwise surface as billing gaps days later, and is served through the AWS MCP Server to any MCP client, including Amazon Q Developer and Kiro. Source

6. Bedrock Knowledge Bases Walkthrough Uses Agentic Retrieval for Cited Insurance Claim Answers

AWS. A new how-to builds a claims assistant on a managed Bedrock knowledge base, where Bedrock selects and operates the embedding model and vector storage, and queries it with the AgenticRetrieveStream API. A foundation model splits multi-part questions into sub-queries and repeats retrieval up to maxAgentIteration rounds until evidence is sufficient, streaming trace events, answer text, and citations back to the client. Metadata sidecar files (up to 10 KB) enable filters on claim ID, type, amount, and dates stored as YYYYMMDD integers, and a Guardrails contextual grounding check blocks unsupported answers. The post uses synthetic data and does not describe a production deployment. Source

7. Three Agents Share One GPU Instance in an AgentCore Runtime Instances Sample

AWS. An AWS sample shows how AgentCore Runtime Instances, the EC2-backed option with sessions up to 14 days, GPUs, and persistent EBS volumes, can colocate several separately deployed agents on one instance by invoking them with the same runtimeSessionId on a shared capacity provider. In the music pipeline, a composition agent uses Claude Sonnet 4.6 for a brief and renders audio with the open-source ACE-Step model on an NVIDIA L4 (20 seconds of 48 kHz stereo in about 9 seconds), while delivery and compliance agents read the file off the shared volume. The post flags two common Strands Agents pitfalls: the handler parameter must be named context to read the session ID, and the Agent must be built inside the handler to avoid re-entrant invocation errors. Source