Bedrock Brief 23 Sep 2026
Amazon is currently locked in a turf war over who gets to control the shopping cart when AI agents come knocking. Meta's new Muse assistant tried to browse Amazon.com, compare strollers, and place orders without ever asking permission or even identifying itself, prompting Amazon to slam the door shut with a popup telling users that "continued access by an unauthorized AI agent violates Amazon's Conditions of Use." This isn't a one-off spat, either. Amazon has already sued Perplexity and blocked agents from Google and OpenAI. The stakes? A $68 billion ad business that depends on people actually seeing sponsored products instead of delegating their entire shopping experience to a bot that never looks at an ad. The irony is thick: Meta and Amazon are cloud infrastructure partners with a multibillion-dollar deal to run agentic AI on Graviton chips, yet they can't agree on whether an AI agent should be allowed to shop without a hall pass.
While Amazon is busy building moats around its e-commerce site, AWS is sprinting in the opposite direction. Q2 revenue hit $42 billion, up 37% year over year, the fastest growth rate in 18 quarters. Graviton and Trainium are each running at $25 billion-plus annually, with triple-digit growth, and the contracted backlog now sits at $496 billion. That's not speculative demand, either. Anthropic and OpenAI have signed multi-year, multi-gigawatt commitments to Trainium, and 98% of the top 1,000 EC2 customers are using Graviton. The capex burn is real (free cash flow swung to negative $7.6 billion), but most of the AI capacity is already contracted for five-year terms, which means Amazon is building data centers with revenue already locked in before the servers even go online.
And if you're wondering how anyone is supposed to interact with an AI agent that runs in the background for hours or days, AWS thinks the answer is an inbox, not a chat window. The company open-sourced Pizza Bot, a self-hosted app that gives users separate threads for ongoing tasks and a queue for work that's done or needs human input. It's a small signal, but it points to a bigger shift: agents are moving from answering prompts to handling long-running jobs autonomously, and the UI metaphor has to change with them. Between locking down its retail empire, printing record cloud growth, and rethinking how people manage agentic workflows, Amazon is placing bets on both sides of the AI shopping divide.
Fresh Cut
- CloudWatch Omni lets you debug AI agents locally in VS Code without needing an AWS account, and uses natural language chat to help you find bugs by automatically mapping your application's dependencies and telemetry across multiple cloud providers. Read announcement →
- OpenAI's GPT-6 Sol makes half as many mistakes as GPT-5.6 and can debug code end-to-end, while GPT-6 Luna handles high-volume tasks like data extraction—both with 1 million token context windows through AWS Bedrock. Read announcement →
- Claude Opus 5.5 completes tasks using fewer tokens than its predecessor while costing less per token, and it's available on AWS GovCloud with zero data retention for government workloads. Read announcement →
- Security Hub can now scan your Azure VMs using software bill of materials analysis to automatically find AI models and inference endpoints like Ollama or vLLM running there, correlating them with security findings alongside your AWS resources. Read announcement →
- AWS Glue Data Quality uses generative AI to automatically suggest data validation rules for your database tables in seconds, so you don't have to manually write checks for every column when setting up data pipelines. Read announcement →
- AWS's X8i servers with custom Intel Xeon 6 chips are now available in São Paulo, offering up to 6TB of RAM and 3.3x faster memory speeds than the previous generation—useful if you're running massive databases or memory-heavy applications in South America. Read announcement →
- AWS Resilience Hub can now use AI to automatically spot weird patterns in your app's dependencies—like services talking across regions or unexpected connections—so you can find resilience problems before they cause outages. Read announcement →
- Kimi K3, a 2.8 trillion parameter open-weight model with a 1-million-token context window, is available on Amazon Bedrock with explicit prompt caching that cuts costs when you're repeatedly feeding the same large codebase or documents into your AI requests. Read announcement →
- Amazon SNS lets you send messages up to 1 MiB (4x larger than before), so you can pass more data through its notification system without having to split your payloads into chunks or store them separately. Read announcement →
- Amazon's new AgentCore Runtime uses memory snapshots to keep cold start times at 2 seconds for containers up to 2GB (down from 5-30 seconds), and reclaims unused memory during sessions so you only pay for what you're actively using. Read announcement →
The Quarry
Introducing Amazon SageMaker HyperPod Inference Gateway
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes add-on that routes inference requests based on live GPU telemetry instead of treating pods like interchangeable boxes, which can slash first-token latency by up to 82% without you touching a single line of model server code. The clever bit here is GPU-aware routing: instead of round-robin or random distribution, the gateway looks at real-time signals like GPU memory and utilization to pick the pod most ready to handle your request right now. It's a drop-in improvement for EKS clusters running LLM workloads, proving that smarter traffic cops can make a massive difference when every millisecond counts. Read blog →
More posts:
- Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock
- Claude Opus 5.5 is now available on AWS
- Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore
- How Reactiv automates mobile commerce 80% faster with Amazon Bedrock AgentCore
- Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI
Core Sample
In the Field: The exclusive tour inside TBC.co's lab
The Biological Computing Co. is studying nature's neural efficiency to build better AI models, and they're using AWS SageMaker to turn those biological insights into actual working systems. The tour reveals how TBC's lab work translates organic computing principles into practical machine learning architectures that could make AI training and inference far less power-hungry than current brute-force approaches. If you've ever wondered why your GPU bill makes you weep while a fruit fly's brain runs on practically nothing, this is the research trying to close that gap. Watch video →
More videos:
- An Intro to AI Tokenomics
- GameDay Rapid Rewind: Real Estate Adventure
- American Access Institute: Scaling mentorship for 2 million youth with AWS
- Why AI Means Rewiring the Business Model with Jonathan Peachey, CEO, Factory X
- Cloud Report | Let's build a new game with Kiro