Making sure your production system is well behaved is table stakes for any SaaS. The space around monitoring and alerting is complex and moving fast, so it is too easy to end up with results that are worse than useless. Alert fatigue is real and developers must learn to avoid this.
Shahar and Tal, founders of Keep and active SaaS Developer community members joined me to discuss - why is observability so complicated, how alerts fit in, and a lot of best practices for alerting.
If you feel like you have good handle on your alerts today, you should skip to minute 23:40 - where we talk about the future. CI/CD for alerts, semantic layer for alerts and how AI will recommend alerts + ways to resolve them.
Links:
The big observability survey: https://grafana.com/observability-sur...
Problems in Alerting blog: / current-problems-in-the-alerting-space
Our episode about SLO with Ken Finnegan from Workday: • SLO - Best Practices for Reliable SaaS
SaaS Developer Slack: https://launchpass.com/all-about-saas
-----
00:00 intros
02:45 Map of Observability
05:20 OpenTelemetry drove adoption
07:20 OpenTelemetry Ecosystem?
08:55 Big trend in Observability - consolidation
11:03 What do we do with the data?
12:21 types of alerts
15:30 too many alerts
16:52 blame micro services
19:00 Improving alerts
20:20 More challenges with alerts23:40 Alerts are code and need dev tools
29:00 Alerts are like tests for post-production
31:30 Semantic layer for alerts
33:03 Alerting with GPT4
34:45 One bit of alerting advice