Google - Site Reliability Engineering
Google datacenters are very different from most conventional datacenters and small-scale server farms. These differences present both extra problems and opportunities. This chapter discusses the challenges and opportunities that characterize Google datacenters and introduces terminology that is used throughout the book. Most of Google’s compute resources are in Google-designed datacenters with proprietary power distribution, cooling, networking, and compute hardware (see [Bar13]). Unlike "standard" colocation datacenters, the compute hardware in a Google-designed datacenter is the same across the board.9 To eliminate the confusion between server hardware and server software, we use the following terminology throughout the book: Machines can run any server, so we don’t dedicate specific machines to specific server programs. There’s no specific machine that runs our mail server, for example. Instead, resource allocation is handled by our cluster operating system, Borg. We realize this us
Google SRE - SRE best practices for production environment Chapter 2 - The Production Environment at Google, from the Viewpoint of an SRE Table of Contents Foreword Preface Part I - Introduction 1. Introduction 2. The Production Environment at Google, from the Viewpoint of an SRE Part II - Principles 3. Embracing Risk 4. Service Level Objectives 5. Eliminating Toil 6. Monitoring Distributed Systems 7. The Evolution of Automation at Google 8. Release Engineering 9. Simplicity Part III - Practices 10. Practical Alerting 11. Being On-Call 12. Effective Troubleshooting 13. Emergency Response 14. M
Explore this link on the map →related reading
- The Friendship That Made Google Huge | The New Yorkernewyorker.com
- Google SRE - IT Service Management: Automate Operationssre.google
- Google SRE - Embracing risk and reliability engineering booksre.google
- The maze is in the mouse. What ails Google. And how it can turn… | by Praveen Seshadri | Mediummedium.com
- How to Rack 30 Petabytes of Storage | blogsi.inc
- NYSRGnotes.ekzhang.com
- Introducing Maelstrom – how Facebook moves traffic among data centers to handle disasterslinkedin.com
- Google SRE: Time Series Database for Monitoring and Alertingsre.google
- Google SRE monitoring ditributed system - sre golden signalssre.google
- Notes on Distributed Systems for Young Bloods – Something Similarsomethingsimilar.com
- What I learned getting acquired by Googleshreyans.org
- Google SRE Principles: SRE Operations and How SRE Teams Worksre.google