Monitoring is a Pain
I have a confession. Despite having been hired multiple times in part due to my experience with monitoring platforms, I have come to hate monitoring. Monitoring and observability tools commit the cardinal sin of tricking people into thinking this is an easy problem. It is very simple to monitor a small application or service. Almost none of those approaches scale. Instead monitoring becomes an endless series of small failures. Metrics disappeared for awhile, logs got dropped for a few hours, the web UI for traces doesn't work anymore. You set up these tools with the mentality of "set and forget" but they actually require ever increasing amounts of maintenance. Some of the tools break and are never fixed. The number of times I join a company to find an unloved broken Jaeger deployed has been far too many. It feels like we have more tools than ever to throw at monitoring but we're not making progress. Instead the focus seems to be on increasing the output of applications to increase the
And we're all doing it wrong (including me) I have a confession. Despite having been hired multiple times in part due to my experience with monitoring platforms, I have come to hate monitoring. Monitoring and observability tools commit the cardinal sin of tricking people into thinking this is an easy problem. It is very simple to monitor a small application or service. Almost none of those approaches scale. Instead monitoring becomes an endless series of small failures. Metrics disappeared for awhile, logs got dropped for a few hours, the web UI for traces doesn't work anymore. You set up thes
Explore this link on the map →saved by
related reading
- Logging, Tracing, Monitoring, et al.alexandruburlacu.github.io
- Scale Prometheus with Chronosphere | Chronospherechronosphere.io
- Why is observability so expensive?mattklein123.dev
- Cloud Native Observability Platform | Chronospherechronosphere.io
- Sensu | An Introduction to Prometheus Monitoring (2021)sensu.io
- What Full-Stack Observability Requires Todaynewrelic.com
- On actionable and actually useful logs | Lanre Adelowolanre.wtf
- The Problem with OpenTelemetrycra.mr
- Mediumgongybable.medium.com
- Prometheus Metric 的实践总结,搞定监控需注意~ - 运维派yunweipai.com
- Google SRE monitoring ditributed system - sre golden signalssre.google
- Everything I know about good system designseangoedecke.com