Community adoption of Kubernetes (instead of YARN) as a cluster-manager for Apache Spark has been accelerating since initial support was released with Spark 2.3. Companies like to run Spark on Kubernetes because they want to manage their entire infrastructure in a homogeneous, cloud-agnostic way, as well as to take advantage of the improved isolation and resource sharing enabled by k8s.
In this talk, the founders of Data Mechanics, a serverless Spark platform built on Kubernetes, will go over their lessons learned through helping customers adopt Spark-on-k8s. Topics include: Architecture, Initial setup, Docker-based development workflows, I/O and performance optimizations, Dynamic allocation and cost optimizations, and finally, Limitations and future works planned for Spark-on-k8s