More Than DNS: The 14 hour AWS us-east-1 outage
On Monday the AWS us-east-1 region had its worst outage in over 10 years. The whole thing lasted over 16 hours and affected 140 AWS services, including, critically, EC2. SLAs were blown, an eight-figure revenue reduction will follow. Before Monday, I’d spent around 7 years in industry and never personally had production nuked by a public cloud outage. I generally regarded AWS’s reliability as excellent, industry-leading. What the hell happened? A number of smart engineers have come to this major bust-up and covered it with the blanket of a simple explanation: brain drain; race condition; it’s always DNS; the cloud is unreliable, go on-prem. You’re not going to understand software reliability if you summarize an outage of this scale in an internet comment. Frankly, I’m not even going to understand it after reading AWS’s 4000 word summary and thinking about it for hours. But I’m going to hold the hot takes and try. I wrote Modal’s internal us-east-1 incident postmortem before AWS publish
I’m on the right, biting a nail nervously. We’re in an Italian hotel because this happened on day 1 of our offsite. On Monday the AWS us-east-1 region had its worst outage in over 10 years. The whole thing lasted over 14 hours and affected 140 AWS services, including, critically, EC2. SLAs were blown, an eight-figure revenue reduction will follow. Before Monday, I’d spent around 7 years in industry and never personally had production nuked by a public cloud outage. I generally regarded AWS’s reliability as excellent, industry-leading. What the hell happened? A number of smart engineers have co
Explore this link on the map →saved by
related reading
- How I Dropped Our Production Database and Now Pay 10% More for AWSalexeyondata.substack.com
- Building and operating a pretty big storage system called S3 | All Things Distributedallthingsdistributed.com
- The real serverless compute to database connection problem, solved - Vercelvercel.com
- Mediumnetflixtechblog.com
- GitHub - chime/terraform-aws-alternat: High availability implementation of AWS NAT instances. · GitHubgithub.com
- Amazon’s Cloud Crisis: How AWS Will Lose The Future Of Computingsemianalysis.com
- The Distributed Computing Manifesto | All Things Distributedallthingsdistributed.com
- 10 Lessons from 10 Years of Amazon Web Services | All Things Distributedallthingsdistributed.com
- A Byzantine failure in the real world | The Cloudflare Blogblog.cloudflare.com
- An unexpected discovery: Automated reasoning often makes systems more efficient and easier to maintain | AWS Security Blogaws.amazon.com
- Cloudflare outage on November 18, 2025 | The Cloudflare Blogblog.cloudflare.com
- AI Agent Bankrupted Their Operator While Trying to Scan DN42 - Lan Tian @ Bloglantian.pub