Google - Site Reliability Engineering
It is a truth universally acknowledged that systems do not run themselves. How, then, should a system—particularly a complex computing system that operates at a large scale—be run? Historically, companies have employed systems administrators to run complex computing systems. This systems administrator, or sysadmin, approach involves assembling existing software components and deploying them to work together to produce a service. Sysadmins are then tasked with running the service and responding to events and updates as they occur. As the system grows in complexity and traffic volume, generating a corresponding increase in events and updates, the sysadmin team grows to absorb the additional work. Because the sysadmin role requires a markedly different skill set than that required of a product’s developers, developers and sysadmins are divided into discrete teams: "development" and "operations" or "ops." The sysadmin model of service management has several advantages. For companies decidi
Google SRE - IT Service Management: Automate Operations Chapter 1 - Introduction Table of Contents Foreword Preface Part I - Introduction 1. Introduction 2. The Production Environment at Google, from the Viewpoint of an SRE Part II - Principles 3. Embracing Risk 4. Service Level Objectives 5. Eliminating Toil 6. Monitoring Distributed Systems 7. The Evolution of Automation at Google 8. Release Engineering 9. Simplicity Part III - Practices 10. Practical Alerting 11. Being On-Call 12. Effective Troubleshooting 13. Emergency Response 14. Managing Incidents 15. Postmortem Culture: Learning from F
Explore this link on the map →related reading
- Google SRE Principles: SRE Operations and How SRE Teams Worksre.google
- Google SRE - Embracing risk and reliability engineering booksre.google
- Google SRE monitoring ditributed system - sre golden signalssre.google
- SRE 的工作介绍 | 卡瓦邦噶!kawabangga.com
- What is an AI SRE? | Clericcleric.io
- The maze is in the mouse. What ails Google. And how it can turn… | by Praveen Seshadri | Mediummedium.com
- Google SRE - SRE best practices for production environmentsre.google
- The Friendship That Made Google Huge | The New Yorkernewyorker.com
- Building and operating a pretty big storage system called S3 | All Things Distributedallthingsdistributed.com
- I could do that in a weekend!danluu.com
- Operating Well: What I Learned at Stripeevery.to
- AddyOsmani.com - 21 Lessons From 14 Years at Googleaddyosmani.com