Bedrock Brief 07 Oct 2026
AWS just made a very loud statement this week: it's willing to spend a quarter trillion dollars on AI infrastructure, but it really needs you to stop complaining about the data centers that spend depends on. Matt Garman's surprisingly blunt post warned that local moratoriums could cost the U.S. its lead in the AI race, though the $1 billion "Built Together" fund he announced works out to roughly 0.1% of this year's projected capital expenses. It's a fascinating tension: Amazon needs communities to accept data centers, communities want more in return, and meanwhile 75 data center projects worth $130 billion got blocked or delayed in just the first three months of this year. The infrastructure gold rush is real, but it turns out you can't just plop down a power-hungry facility without some serious community relations work.
The other big news is that Amazon just poached Pedram Rezaei, a 20-year Microsoft veteran and former Copilot CTO, to work on Kiro and Quick. That's notable not just because it's a high-profile hire (he's joining as a distinguished engineer), but because it signals how seriously AWS is taking the workplace AI assistant battle. Microsoft has been aggressive with Copilot, and AWS has been playing catch-up with Quick, which only hit general availability last September. Rezaei's background in shipping Copilot at scale could help close that gap, assuming he can translate Microsoft's playbook into Amazon's very different engineering culture.
On the positive side, AWS is betting that falling costs will save the day. Matt Wood told CNBC that AI costs outside frontier models are dropping by 10x to 100x every three to six months, which would make the current infrastructure binge look prescient if it holds true. AWS also launched the Well-Architected Agent, an AI that audits your cloud setup and tells you how to save money and improve performance. It's a clever move: give customers a tool to optimize their spending while simultaneously proving that AWS understands the cost anxiety everyone's feeling right now. Whether that's enough to justify the record capital expenditures remains to be seen, but at least someone's talking about making AI cheaper instead of just bigger.
Fresh Cut
- Amazon Bedrock can run GLM 5.3, a 753-billion parameter AI model with a 1-million-token context window that you can use for coding tasks where the AI needs to remember and work with massive amounts of code at once. Read announcement →
- Amazon's Nova 2.5 Sonic speech-to-speech model can handle multi-step conversations and call external tools during voice interactions, letting you build AI voice agents that can look up data and take actions while talking to users in real-time. Read announcement →
- AI coding assistants can call AWS APIs through a single interface in six new regions including Singapore, Sydney, Tokyo, Ireland, London, and Oregon, letting your tools provision infrastructure and debug issues without leaving your geography. Read announcement →
- Amazon Bedrock's AgentCore Gateway can now verify TLS certificates from your own private certificate authority when connecting to VPC endpoints, eliminating the need for an Application Load Balancer in between. Read announcement →
- Amazon Q can now turn questions like "show me the health of my production environment" into interactive dashboards with charts and status indicators by orchestrating calls across hundreds of AWS APIs, replacing the line-by-line text summaries it used to generate. Read announcement →
- Security Hub groups multiple security issues that share the same root cause into a single fix with step-by-step instructions in formats like Terraform and Python, so you can resolve dozens of problems by fixing one misconfigured resource instead of tackling each issue separately. Read announcement →
- You can update all your AI coding agent skills at once with a single CLI command instead of manually updating each one individually when new versions come out. Read announcement →
- OpenAI's GPT-6 Astra can generate up to 300 tokens per second through Amazon Bedrock's UltraFast mode, making it 6x faster for applications like coding assistants where response time directly impacts user experience. Read announcement →
- Aurora's serverless database can now jump from handling almost no traffic to 16x more capacity in one second (up to 256 ACUs total), which matters if you're building AI agents or anything with unpredictable traffic spikes that would otherwise crash or waste money on idle servers. Read announcement →
- Claude Opus 5.5 and Sonnet 5 can process your AI requests entirely within London's data centers, which matters if you need to keep customer data inside the UK for regulatory compliance. Read announcement →
The Quarry
Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows
Ambient agents flip the traditional chatbot model on its head by reacting to real-world triggers like S3 uploads or CloudWatch alarms instead of sitting around waiting for someone to type a question. This post shows you how to wire up Amazon Bedrock AgentCore with SQS for event queuing, Lambda for orchestration, and DynamoDB for state persistence, all while implementing a single ask_human tool that pauses agent execution until a human reviews and approves the next step via a custom Jobs UI. The pattern is framework-agnostic and proves that agents don't need to live inside a chat window to be useful, they just need the right hooks into your existing infrastructure. Read blog →
More posts:
- Building a context-aware AI assistant on AgentCore and OpenClaw
- Responsible AI governance: How AWS positions customers to align with ISO/IEC 42005:2025
- Best practices for Amazon SageMaker HyperPod administration and governance
- Manage Amazon SageMaker HyperPod Spaces directly from SageMaker Studio
- Build a voice travel concierge with Amazon Bedrock AgentCore, Managed Knowledge Base and Nova Sonic
Core Sample
AWS: AI Answered — What does it take to run AI in production?
Matt Wood argues that getting AI into production isn't just a tech problem but an operational shift, and the old top-down "change management playbook" falls flat because teams need to experiment their way forward rather than follow a grand plan. He digs into why public benchmarks are noisy proxies at best (they rarely match your actual use case), why you should treat models as swappable components to dodge vendor lock-in, and how the cost per unit of intelligence keeps dropping even as your usage climbs. One engineer-friendly nugget: Wood emphasizes building abstraction layers around model APIs so you can A/B test providers or swap foundations without rewriting your entire stack. Watch video →
More videos:
- Amazon Connect Customer - Omnichannel Customer Experience
- Can agents run research and robots?
- GenerAIly Speaking: Let's Talk Tokenomics
- GenerAIly Speaking: So, what exactly is a Token?
- AWS: AI Answered — What does it take to run AI in production?