Google SRE - Blameless Postmortem for System Resilience
sre.google · 2,062 words · saved by 2 readers
Blameless postmortems in SRE culture. Incident study that focus on root cause analysis and preventive actions, for culture of continuous improvement.
Written by John Lunney and Sue Lueder Edited by Gary O’ Connor The cost of failure is education. Devin Carraway As SREs, we work with large-scale, complex, distributed systems. We constantly enhance our services with new features and add new systems. Incidents and outages are inevitable given our scale and velocity of change. When an incident occurs, we fix the underlying issue, and services return to their normal operating conditions. Unless we have some formalized process of learning from these incidents in place, they may recur ad infinitum. Left unchecked, incidents can multiply in…
saved by
related reading
- How Complex Systems Failhow.complexsystems.fail
- Prefacesre.google
- Galactic Empire | Purdue Hackersblog.purduehackers.com
- The Kool Aid Factory :: Writing In Public, Inside Your Companykoolaidfactory.com
- Google SRE Principles: SRE Operations and How SRE Teams Worksre.google
- Being oncall taught me everything - Yao Yueyaoyue.org
- Google SRE - IT Service Management: Automate Operationssre.google
- Strategies for Learning from Failurehbr.org
- How SRE Relates to DevOpssre.google
- How Stripe Built a Writing Culture - Knock Down Silos by Slabslab.com
- AI-native incident management platform | Rootlyrootly.com
- Accountability Sinks — LessWronglesswrong.com