Scaling LinkedIn's Compute Platform with Kubernetes: Microservices, Scheduling & ML Jobs

Опубликовано: 07 Май 2026
на канале: Perfology
648
40

At LinkedIn, Kubernetes is more than just an orchestration tool — it's a foundational primitive in our compute platform. In this talk, we dive deep into how we manage thousands of microservices, run large-scale stateful applications with a custom scheduler, and operate a massive fleet of GPUs — all while ensuring zero downtime and no manual intervention during regular maintenance of our bare metal infrastructure.

What You'll Learn:

How LinkedIn leverages Kubernetes at scale

Managing stateful applications with a custom scheduler

Orchestrating large GPU workloads efficiently

Performing live host maintenance without downtime

Insights into LinkedIn’s infrastructure automation strategies


Here's a breakdown of the video's chapters with timestamps:

Introduction to LinkedIn's Compute Platform (0:00-2:36)
Infrastructure as a Service Layer (2:36-6:50)
Coordinated Maintenance Operations (6:50-7:23)
Kubernetes Cluster Organization (7:23-8:46)
Kubernetes Resource Management and Scaling (8:46-11:00)
Workload Platform Layers (11:00-13:40)
Stateful and Stateless Workloads on Kubernetes (13:40-15:40)
User Workflows and API Guardrails (15:40-23:12)
Future Initiatives (23:12-25:50)
Migration Lessons Learned (25:50-26:25)
Q&A Session (26:25-38:00)



🚀 Whether you're building at scale or just starting with Kubernetes, this is a must-watch for SREs, DevOps engineers, and platform teams looking to optimize their infrastructure.

Follow Us on Facebook:   / perfology  
Connect on Instagram:   / perfologys  
Network on LinkedIn:   / perfology  

📌 Subscribe for more deep dives into cloud-native architecture, Kubernetes, and site reliability at scale.

#LinkedIn #Kubernetes #Microservices #CloudNative