AWS AI Updates: October 6, 2026
1. Z.ai’s GLM 5.3 Is Generally Available on Amazon Bedrock
AWS. Bedrock added Z.ai’s GLM 5.3, a 753B-parameter mixture-of-experts model with about 40B active parameters per token, a 1M-token context window, up to 128K output tokens, and always-on reasoning with adjustable effort. It is reachable through Bedrock’s OpenAI-compatible Responses and Chat Completions endpoints as well as Invoke and Converse, via the us.zai.glm-5.3 and global.zai.glm-5.3 cross-Region profiles, with explicit prompt caching (minimum 1,024 tokens per breakpoint) and Flex, Standard, and Priority tiers. Access is limited to eligible enterprise customers; Z.ai reports a 50 percent gain over GLM 5.2 on its internal coding benchmark and 84.5 on CyberGym. Source
2. Amazon Nova 2.5 Sonic Improves Reasoning and Tool Calling for Voice Agents
AWS. Amazon released Nova 2.5 Sonic, a speech-to-speech model for real-time voice agents with improved reasoning, instruction following, and tool-calling accuracy compared with Nova 2 Sonic. It supports a 256K context window, seven languages, and lower latency, and is available in Bedrock at the same price as Nova 2 Sonic. Source
3. Claude Opus 5.5 and Sonnet 5.5 Arrive in AWS GovCloud for Claude Code
AWS. Anthropic’s Claude Opus 5.5 and Claude Sonnet 5.5 are now available in the AWS GovCloud (US) Regions with FedRAMP Class D certification, with us-gov-west-1 serving both bedrock-runtime and bedrock-mantle endpoints and US-East serving bedrock-runtime only. The post shows how to point Claude Code at GovCloud using CLAUDE_CODE_USE_BEDROCK=1 and the us-gov.anthropic.claude-sonnet-5-5 and us-gov.anthropic.claude-opus-5-5 model IDs, aimed at regulated and ITAR workloads. Source
4. New aws-ai-ml Agent Skill Benchmarks and Right-Sizes SageMaker Inference Endpoints
AWS. The aws-ai-ml skill in the Agent Toolkit for AWS (installed with npx skills add aws/agent-toolkit-for-aws/skills/aws-ai-ml) lets coding agents such as Kiro, Claude Code, and Codex load-test SageMaker endpoints for throughput, p50/p99 latency, and time-to-first-token. It also ranks candidate instance types for custom, JumpStart, or Hugging Face models and generates SageMaker Python SDK v3 code. The skill asks for confirmation before benchmarking live endpoints and stages Hugging Face models to S3 automatically. Source