Bedrock Brief 19 Aug 2026

Share
Bedrock Brief 19 Aug 2026

Well, here's an unexpected plot twist: Amazon, the company that quite literally built its empire on selling books, is now buying rare books in bulk, tearing off their spines, and scanning them for AI training data before destroying them. A 404 Media investigation used an AirTag hidden in a rare book to track it to an Amazon facility in Las Vegas, where workers spend their days scanning volumes that can never be reprinted. The logo for the team? A T-Rex about to devour a book. Subtlety is dead, folks, and Amazon apparently scanned it for training data too. This practice highlights just how desperate the AI arms race has become: with the internet already scraped clean and model collapse lurking around every corner when LLMs train on AI-generated slop, companies are turning to physical texts that predate the generative AI era. Rare books offer unique training data that competitors can't easily replicate, which makes them valuable. What makes them tragic is that Amazon started as an online bookstore promising to democratize access to literature, and now it's literally feeding books to dinosaurs.

Meanwhile, AWS is having a much better week in the finance pages. According to Pershing Square's Q2 investor letter, AI demand has pushed AWS revenue growth from 20% to over 30% this year, a meaningful acceleration that's catching investor attention despite concerns about massive capital expenditures for datacenter buildouts. Amazon's core retail business is also thriving with 15% unit volume growth, the fastest pace since 2021. The irony is thick: Amazon's AI ambitions are simultaneously boosting its cloud business while revealing some deeply questionable practices about how it sources training data. One part of the company is profiting from the AI boom, while another is quietly dismantling cultural artifacts to fuel it.

The book scanning story also reveals a broader industry problem. After a judge ruled that Anthropic's similar book-destroying practices were "transformative" and didn't violate copyright law, the floodgates opened. Booksellers have suspected for months that bulk orders with random selections but consistent ISBNs were destined for AI training, and now we have confirmation that at least one major player is systematically working through the catalog. Companies need unfathomably large amounts of text to train frontier models, and they're willing to destroy irreplaceable physical copies to get it. It's a perfect encapsulation of the AI era: move fast and break things, even if those things are rare books that have survived decades or centuries only to end up as fodder for chatbots.

Fresh Cut

  • OpenAI's GPT-5.6 models (Terra and Luna) can now automatically distribute your API requests across multiple AWS data centers in India while guaranteeing your data never leaves the country, which matters if you're building apps that need to comply with Indian data residency laws. Read announcement →
  • AI agents can autonomously pay for APIs and services using the Machine Payment Protocol (MPP) and x402 protocol, with configurable spending limits to prevent runaway costs. Read announcement →
  • SageMaker Unified Studio can automatically flag when your data drifts from historical patterns without you having to set up manual threshold rules, helping you catch data quality issues before they break your pipelines. Read announcement →
  • AWS servers with Intel Xeon 6 chips and 2.5x faster memory are now in Tel Aviv, running PostgreSQL 30% faster and handling larger databases than the previous generation. Read announcement →
  • You can attach up to 1 GB of JSON, XML, or YAML metadata directly to S3 files that automatically moves and deletes with the file, eliminating the need to build a separate database to track what your stored data actually contains. Read announcement →
  • AWS Bedrock automatically routes your OpenAI GPT requests across multiple regions to handle traffic spikes and cut costs, while letting you control whether data stays within a specific geography or goes anywhere for maximum throughput. Read announcement →
  • Calgary region gets memory-optimized EC2 instances with Intel Xeon 6 processors that run PostgreSQL 30% faster, NGINX 60% faster, and AI recommendation models 40% faster than the previous generation. Read announcement →
  • Amazon QuickSight's AI can now edit your Excel spreadsheets, PowerPoint decks, Word documents, and Outlook emails directly inside Microsoft 365, using your company's data to build financial models, track document changes, and draft emails with full inbox context. Read announcement →
  • Amazon OpenSearch Service can automatically convert your search queries into semantic embeddings that understand meaning instead of just matching keywords, and this feature now works on domains inside private VPCs where it was previously blocked. Read announcement →
  • AWS Console-to-Code can now track your console clicks across different regions and browser tabs to generate infrastructure code, solving the problem where switching regions would previously erase your recorded actions. Read announcement →

The Quarry

Custom reward functions for multi-turn reinforcement learning with Amazon Nova Forge

Amazon Nova Forge lets you write custom multi-turn reward functions that guide reinforcement learning across entire conversations, not just single responses, which means you can penalize a chatbot that gives a correct answer but forgets context from three turns ago. The real trick is building composite rewards that safely execute model-generated code inside the training loop while instrumenting each component so you spot reward collapse before your model learns to game the system instead of solving the actual task. If you've ever wondered why your RL fine-tuning produced a model that's technically correct but useless in practice, it's probably because your reward function measured the wrong thing or silently broke halfway through training. Read blog →

More posts:


Core Sample

8x Faster Genetic Diagnostics While Keeping Patient Data in Germany

Hannover Medical School slashed whole-genome sequencing time from several hours down to one hour by running specialized Illumina DRAGEN instances on AWS EC2, processing over 10,000 patient samples annually while keeping all that sensitive genetic data locked within German borders via AWS European Sovereign Cloud. The team uses AWS Batch to orchestrate the compute-heavy workloads and S3 for storage, achieving an 8x speed boost that directly impacts how quickly doctors can make clinical decisions. Data residency isn't just a nice-to-have here; it's legally mandated, and the sovereign cloud setup lets MHH stay compliant without sacrificing the horsepower needed for high-volume genomic analysis. Watch video →

More videos: