1264410

Your Availability Zone Isn't Highly Available. And Your Region Might Not Be Either.

Format: Talk

There is a line that shows up in nearly every architecture review: “we are running in a single Availability Zone, but AWS is reliable enough, and if something goes wrong we will deal with it.” The math feels fine, because AZ outages are rare and usually measured in hours. You restore from a snapshot, relaunch some instances, write the post-mortem, and move on.

In early 2026 that template broke. Physical attacks took multiple AWS data centers offline across two regions at the same time. Two of three Availability Zones in one region went dark, recovery stretched into months instead of hours, and AWS told customers to migrate out of the region entirely. The obvious disaster-recovery target, the nearest region, was hit too. Every assumption underneath “we will deal with it” failed at once.

This talk is the architecture lesson, not the news story. Why stateful zonal resources like EC2, EBS, and Elastic IPs quietly become anchors the moment a zone goes hard down, including the Elastic IP trap where you cannot move a public address off hardware that is simply off. Why regional services like S3, DynamoDB, and Lambda fared better, and where even they degrade. Why “multi-AZ in one region” and “DR in the nearest region” are weaker guarantees than they look, and what static stability actually costs when you do it properly.

You will leave able to audit your own zonal dependencies and tell which resources are anchors, which can move to regional services, and what a disaster-recovery region that is genuinely independent looks like.

Luka Kladarić
Chaos Guru
Contact Us

Credits

This website uses the open source AWS Community Day Template built by AWSug.nl hosted on Amazon CloudFront and Amazon S3. The website uses bootstrap and hugo.