AWS AI Updates: September 29, 2026
1. Claude Sonnet 5.5 Arrives on Bedrock, Claude Platform on AWS, and GovCloud
AWS. Anthropic’s Claude Sonnet 5.5 is now available on Amazon Bedrock through the Global cross-Region inference profile (global.anthropic.claude-sonnet-5-5) and on Claude Platform on AWS in North America, with a separate launch in AWS GovCloud (US). AWS positions it as a lower cost-per-task, faster companion to Opus 5.5 for well-scoped work such as alert triage, SQL generation, UI testing, and IDE coding agents, while Opus 5.5 handles release debugging, large PR security review, and long analyses. The model can return a thinking block before the text block, so AWS’s sample code selects the text block by type rather than by a fixed index. Source
2. Grok 4.7 on Bedrock Doubles Output Tokens per Task for Its Benchmark Gains
AWS. xAI’s Grok 4.7 is now on Amazon Bedrock with a 500K token context window, text and image input, four reasoning effort levels (low to xhigh), and support for the Responses, Chat Completions, InvokeModel, and Converse APIs via us.xai.grok-4.7 and global.xai.grok-4.7 inference profiles. AWS cites Artificial Analysis figures showing the Coding Agent Index rising from 47 to 56 and hallucination rate falling from 34 to 29 percent versus Grok 4.6, but output tokens per Intelligence Index task rose from about 38K to 81K at xhigh effort. The model supports implicit prompt caching, Guardrails, structured outputs, and Standard, Priority, and Flex service tiers. Source
3. vLLM-Omni DLC Streams Qwen3-TTS Audio Over SageMaker Bidirectional Streaming
AWS. AWS published a walkthrough that deploys Qwen3-TTS with the AWS vLLM-Omni Deep Learning Container (v1.5) so a client streams text in and receives 24 kHz PCM audio chunks back over one full-duplex HTTP/2 WebSocket on port 8443. The endpoint uses SageMaker instance pools that try ml.g6.xlarge first and fall back to g6e, g5, or g4dn when capacity is short, but SageMaker validates quota for every pool entry at creation and hourly cost changes with whichever instance is selected. It pairs with an earlier post covering the speech-to-text input side with Voxtral. Source
4. The Same vLLM-Omni Container Chains FLUX.2 Image and Wan2.1 Video Generation
AWS. Part 2 of the series deploys one pinned vLLM-Omni DLC image (omni-sagemaker-cuda-v1.6) to two endpoints: a real-time ml.g6.xlarge endpoint running FLUX.2-klein-4B for text-to-image, and an asynchronous ml.g6e.xlarge endpoint running Wan2.1-VACE-1.3B that animates the generated image into an MP4 written to S3. In a single validation run, a warm image call returned in 4.7 seconds and Wan VACE reported 8.9 seconds of model latency, which AWS labels as reproduction checkpoints rather than benchmarks. Source
5. Nova Act Replaces Selector-Based Synthetic Monitoring Scripts With Natural-Language Steps
AWS. A reference architecture runs synthetic user-journey checks with Amazon Nova Act inside AgentCore Runtime and the AgentCore Browser tool, triggered by EventBridge Scheduler calling InvokeAgentRuntime directly and alerting through SNS on failure. Because Nova Act reasons over screenshots instead of DOM selectors, steps such as “Click the checkout button” survive CSS and ID changes, and AWS cites over 90 percent accuracy on browser workflows in early enterprise use. The sample six-step ecommerce journey run every 5 minutes works out to about 8,640 sessions and 207,360 Nova Act calls a month, with browser session duration as the main cost driver. Source
6. Textract Adapter Promotion Across Accounts Still Needs a Support Ticket, So AWS Published a Workaround Pipeline
AWS. A new guide covers moving Amazon Textract Custom Queries adapters from proof of concept to production, noting that cross-account promotion currently requires AWS Support tickets and transfers only trained weights, not query definitions or training data. The pattern stores adapter IDs in Systems Manager Parameter Store so production can swap adapter versions without redeployment, and adds a pre-classification step using DetectDocumentText because Textract allows only one adapter per AnalyzeDocument call per page per feature type. CloudFormation and Terraform templates are included. Source
7. Rekognition Face Liveness Now Explains Why a Check Scored Low
AWS. Amazon Rekognition Face Liveness now returns Feedback Codes in the GetFaceLivenessSessionResults response that identify conditions behind a low score, such as poor lighting, face obstruction, closed eyes, or low video quality. A single session can return multiple codes, so applications can prompt users to fix every issue before one retry, which targets identity verification onboarding and re-authentication flows. Source