More Than DNS: The 14 hour AWS us-east-1 outage
On Monday the AWS us-east-1 region had its worst outage in over 10 years. The whole thing lasted over 16 hours and affected 140 AWS services, including, critically, EC2. SLAs were blown, an eight-figure revenue reduction will follow. Before Monday, I’d spent around 7 years in industry and never personally had production nuked by a public cloud outage. I generally regarded AWS’s reliability as excellent, industry-leading. What the hell happened? A number of smart engineers have come to this major bust-up and covered it with the blanket of a simple explanation: brain drain; race condition; it’s always DNS; the cloud is unreliable, go on-prem. You’re not going to understand software reliability if you summarize an outage of this scale in an internet comment. Frankly, I’m not even going to understand it after reading AWS’s 4000 word summary and thinking about it for hours. But I’m going to hold the hot takes and try. I wrote Modal’s internal us-east-1 incident postmortem before AWS publish
I’m on the right, biting a nail nervously. We’re in an Italian hotel because this happened on day 1 of our offsite. On Monday the AWS us-east-1 region had its worst outage in over 10 years. The whole thing lasted over 14 hours and affected 140 AWS services, including, critically, EC2. SLAs were blown, an eight-figure revenue reduction will follow. Before Monday, I’d spent around 7 years in industry and never personally had production nuked by a public cloud outage. I generally regarded AWS’s reliability as excellent, industry-leading. What the hell happened? A number of smart engineers have co
saved by
related reading
- How I Dropped Our Production Database and Now Pay 10% More for AWSalexeyondata.substack.com
- Building and operating a pretty big storage system called S3 | All Things Distributedallthingsdistributed.com
- The real serverless compute to database connection problem, solved - Vercelvercel.com
- Durable Execution Solutionstemporal.io
- On-demand Container Loading in AWS Lambdaarxiv.org
- Mediumnetflixtechblog.com
- 10 Lessons from 10 Years of Amazon Web Services | All Things Distributedallthingsdistributed.com
- Thoughts on the Buildkite Aug 25 incidentsurfingcomplexity.blog
- Amazon’s Cloud Crisis: How AWS Will Lose The Future Of Computingsemianalysis.com
- GitHub - chime/terraform-aws-alternat: High availability implementation of AWS NAT instances.github.com
- A Byzantine failure in the real world | The Cloudflare Blogblog.cloudflare.com
- AI Agent Bankrupted Their Operator While Trying to Scan DN42 - Lan Tian @ Bloglantian.pub