Apache Spark Optimization With Holden Karau | DataFlint Webinar

Опубликовано: 21 Сентябрь 2026
на канале: DataFlint
486
16

Watch Holden Karau (Principal Engineer OSS Spark) and Daniel Aronovich (Co-Founder & CTO, DataFlint) as they dive deep in this practical session on Spark optimization that goes beyond generic best practices.
In this webinar, Holden and Daniel covered not just theory, but also real-world strategies to optimize spark, sharing lessons learned, best practices, and concrete suggestions you can apply immediately.

Speakers:
Holden Karau – Principal Engineer OSS Spark
Holden is a leading voice in the Apache Spark community, author of multiple O'Reilly books on Spark, and a frequent conference speaker. Who’s contributed extensively to Spark's core and specializes in making distributed systems accessible and performant.

Daniel Aronovich – Co-Founder & CTO, DataFlint
Brings 15+ years across software engineering, data science, physics, and technical leadership, from hands-on optimizer to team builder. He hosts Data Science Decoded podcast (3K+ subscribers), leads the Israeli & NYC Apache Spark meetups, and writes the Big Data Performance substack, turning real-world production lessons into practical playbooks for data engineering leaders.


00:00 – Welcome + what you’ll get from this session
02:25 – Meet Holden Karau (Spark OSS, Snowflake, Netflix, Databricks)
04:33 – How Holden got into Spark (Scala + distributed systems)
08:28 – Why performance work matters (stories behind High Performance Spark)
14:10 – Defining Spark optimization (mindset + measurement)
20:34 – Demo begins: finding performance issues with the Spark UI
22:29 – The small files problem (why it happens + why it hurts)
29:35 – Practical guidance: “good” Parquet file sizes (rule of thumb)
37:12 – Pro tip: naming jobs + making the Spark UI usable
45:42 – The future: faster PySpark via UDF transpilation
51:23 – Live Q&A: liquid partitioning, debugging pipelines, RDD vs DF, Arrow, Scala