At LinkedIn, Kubernetes is more than just an orchestration tool — it's a foundational primitive in our compute platform. In this talk, we dive deep into how we manage thousands of microservices, run large-scale stateful applications with a custom scheduler, and operate a massive fleet of GPUs — all while ensuring zero downtime and no manual intervention during regular maintenance of our bare metal infrastructure.
What You'll Learn:
How LinkedIn leverages Kubernetes at scale
Managing stateful applications with a custom scheduler
Orchestrating large GPU workloads efficiently
Performing live host maintenance without downtime
Insights into LinkedIn’s infrastructure automation strategies
Here's a breakdown of the video's chapters with timestamps:
Introduction to LinkedIn's Compute Platform (0:00-2:36)
Infrastructure as a Service Layer (2:36-6:50)
Coordinated Maintenance Operations (6:50-7:23)
Kubernetes Cluster Organization (7:23-8:46)
Kubernetes Resource Management and Scaling (8:46-11:00)
Workload Platform Layers (11:00-13:40)
Stateful and Stateless Workloads on Kubernetes (13:40-15:40)
User Workflows and API Guardrails (15:40-23:12)
Future Initiatives (23:12-25:50)
Migration Lessons Learned (25:50-26:25)
Q&A Session (26:25-38:00)
🚀 Whether you're building at scale or just starting with Kubernetes, this is a must-watch for SREs, DevOps engineers, and platform teams looking to optimize their infrastructure.
Follow Us on Facebook: / perfology
Connect on Instagram: / perfologys
Network on LinkedIn: / perfology
📌 Subscribe for more deep dives into cloud-native architecture, Kubernetes, and site reliability at scale.
#LinkedIn #Kubernetes #Microservices #CloudNative