Google - Site Reliability Engineering
This section examines the principles underlying how SRE teams typically work—the patterns, behaviors, and areas of concern that influence the general domain of SRE operations. The first chapter in this section, and the most important piece to read if you want to attain the widest-angle picture of what exactly SRE does, and how we reason about it, is Embracing Risk. It looks at SRE through the lens of risk—its assessment, management, and the use of error budgets to provide usefully neutral approaches to service management. Service level objectives are another foundational conceptual unit for SRE. The industry commonly lumps disparate concepts under the general banner of service level agreements, a tendency that makes it harder to think about these concepts clearly. Service Level Objectives attempts to disentangle indicators from objectives from agreements, examines how SRE uses each of these terms, and provides some recommendations on how to find useful metrics for your own applications
Google SRE Principles: SRE Operations and How SRE Teams Work Part II - Principles Table of Contents Foreword Preface Part I - Introduction 1. Introduction 2. The Production Environment at Google, from the Viewpoint of an SRE Part II - Principles 3. Embracing Risk 4. Service Level Objectives 5. Eliminating Toil 6. Monitoring Distributed Systems 7. The Evolution of Automation at Google 8. Release Engineering 9. Simplicity Part III - Practices 10. Practical Alerting 11. Being On-Call 12. Effective Troubleshooting 13. Emergency Response 14. Managing Incidents 15. Postmortem Culture: Learning from
Explore this link on the map →related reading
- Google SRE - IT Service Management: Automate Operationssre.google
- Google SRE - Embracing risk and reliability engineering booksre.google
- Google SRE monitoring ditributed system - sre golden signalssre.google
- SRE 的工作介绍 | 卡瓦邦噶!kawabangga.com
- The maze is in the mouse. What ails Google. And how it can turn… | by Praveen Seshadri | Mediummedium.com
- Atlassian Engineering's handbook: a guide for autonomous teams - Inside Atlassianatlassian.com
- Operating Well: What I Learned at Stripeevery.to
- Google SRE - SRE best practices for production environmentsre.google
- Google SRE: Time Series Database for Monitoring and Alertingsre.google
- Software Engineering at Googleabseil.io
- AddyOsmani.com - 21 Lessons From 14 Years at Googleaddyosmani.com
- Being oncall taught me everything - Yao Yueyaoyue.org