Being oncall taught me everything - Yao Yue
Though I have never officially owned the title of DevOps or SRE, the majority of my engineering hours in the first few years of my career were spent on understanding and mitigating incidents. Without exageration, being oncall made me the engineer I am. I was oncall for distributed caching at Twitter for 7.5 years (2010-2017), including the 15 months or so when I managed the team, and the week that officially ended one day past the due date of my first-born. It was not just any service that I was responsible for—Cache had the highest throughput by far, measured by requests per second, of any services at Twitter. And as a load-bearing service, it was far from problem-free—by the time Dan Luu and I co-authored the cache incident survey, we counted no fewer than a dozen high profile (tier 0 or tier 1, which generally meant prolonged site-wide degradation) incidents that were attributed fully or significantly to cache. And I loved it. It’s true that I was pretty young and for most of that p
Though I have never officially owned the title of DevOps or SRE, the majority of my engineering hours in the first few years of my career were spent on understanding and mitigating incidents. Without exageration, being oncall made me the engineer I am. I was oncall for distributed caching at Twitter for 7.5 years (2010-2017), including the 15 months or so when I managed the team, and the week that officially ended one day past the due date of my first-born. It was not just any service that I was responsible for—Cache had the highest throughput by far, measured by requests per second, of any se
Explore this link on the map →saved by
related reading
- AddyOsmani.com - 21 Lessons From 14 Years at Googleaddyosmani.com
- Notes on Distributed Systems for Young Bloods – Something Similarsomethingsimilar.com
- More Than DNS: The 14 hour AWS us-east-1 outage – Jonathon Belotti [thundergolfer]thundergolfer.com
- How I ship projects at big tech companiesseangoedecke.com
- Full Cycle Developers at Netflix — Operate What You Build | by Netflix Technology Blog | Netflix TechBlognetflixtechblog.com
- Engineering spotlight: Jeromy Carriere | Datadogdatadoghq.com
- Building and operating a pretty big storage system called S3 | All Things Distributedallthingsdistributed.com
- Google SRE - IT Service Management: Automate Operationssre.google
- Operating Well: What I Learned at Stripeevery.to
- SRE 的工作介绍 | 卡瓦邦噶!kawabangga.com
- 5 Things I Learned From 5 Years At Vercel | Lee Robinsonleerob.com
- My Failures Onboarding at Splunk | People Workpeople-work.io