Google - Site Reliability Engineering
You might expect Google to try to build 100% reliable services—ones that never fail. It turns out that past a certain point, however, increasing reliability is worse for a service (and its users) rather than better! Extreme reliability comes at a cost: maximizing stability limits how fast new features can be developed and how quickly products can be delivered to users, and dramatically increases their cost, which in turn reduces the numbers of features a team can afford to offer. Further, users typically don’t notice the difference between high reliability and extreme reliability in a service, because the user experience is dominated by less reliable components like the cellular network or the device they are working with. Put simply, a user on a 99% reliable smartphone cannot tell the difference between 99.99% and 99.999% service reliability! With this in mind, rather than simply maximizing uptime, Site Reliability Engineering seeks to balance the risk of unavailability with the goals
Google SRE - Embracing risk and reliability engineering book Chapter 3 - Embracing Risk Table of Contents Foreword Preface Part I - Introduction 1. Introduction 2. The Production Environment at Google, from the Viewpoint of an SRE Part II - Principles 3. Embracing Risk 4. Service Level Objectives 5. Eliminating Toil 6. Monitoring Distributed Systems 7. The Evolution of Automation at Google 8. Release Engineering 9. Simplicity Part III - Practices 10. Practical Alerting 11. Being On-Call 12. Effective Troubleshooting 13. Emergency Response 14. Managing Incidents 15. Postmortem Culture: Learning
Explore this link on the map →related reading
- Google SRE - IT Service Management: Automate Operationssre.google
- Google SRE Principles: SRE Operations and How SRE Teams Worksre.google
- Reliability engineering - Wikipediaen.wikipedia.org
- Google SRE - SRE best practices for production environmentsre.google
- The maze is in the mouse. What ails Google. And how it can turn… | by Praveen Seshadri | Mediummedium.com
- Google SRE monitoring ditributed system - sre golden signalssre.google
- High availability - Wikipediaen.wikipedia.org
- I could do that in a weekend!danluu.com
- Notes on Distributed Systems for Young Bloods – Something Similarsomethingsimilar.com
- Choose Boring Technologyboringtechnology.club
- Mediumnetflixtechblog.com
- Software Engineering at Googleabseil.io