An introduction to monitoring and alerting with timeseries at scale, with Prometheus

Опубликовано: 07 Октябрь 2024
на канале: Linux.conf.au 2016 -- Geelong, Australia
23,363
201

Jamie Wilkinson
https://linux.conf.au/schedule/30131/...
Monitoring is the foundational bedrock of site reliability, and yet is the bane of most sysadmin's lives. Why? Monitoring sucks when the cost of maintenance scales proportionally with the size of the system being monitored. Recently tools like Riemann and Prometheus have emerged that can address this problem, by scaling out monitoring configurations sublinearly with the size of the system.

In this talk, Jamie will talk about the theory of alert design and timeseries-based alerting methods, and complement that with practical examples in Prometheus that you can deploy in your environment today to reduce the amount of alert spam and help operators keep a healthy level of production hygiene.