Having small files to process in Spark can be often a pain. The meta data and scheduling overhead slow down Spark significantly. Reasons for that can be recurrent small JSON responses from REST APIs or pushed from streaming services like Kafka.
Find out how the Small File Problem looks like in Spark and how you can deal with it.
~~~~~~~~~~~~~ Subscribe - Like - Comment - Challenge ~~~~~~~~~~~~~
You want to master Data Engineering with PySpark? Subscribe here: https://www.youtube.com/@DataNikktheG...
Feel free to comment or challenge my explanations as always. Happy to learn also myself more by the community.
~~~~~~~~~~~~~~~~~~~~~~~ Resources ~~~~~~~~~~~~~~~~~~~~~~~
Link to Slides: https://github.com/datanikkthegreek/S...
Link to code: https://github.com/datanikkthegreek/S...
Spark Architecture: • Working in Big Data? Understand Spark's ma...
Schema benefit: • The Force of the Schema - Code that matter...
What influences the partition size incl intro to openCostInCosts and maxPartionBytes: • The Secrets of Influencing Spark Partition...
openCostInCosts explained: • The Secrets of Influencing Spark Partition...
~~~~~~~~~~~~~~~~~~~~~~~ Chapters ~~~~~~~~~~~~~~~~~~~~~~~
00:00 - Introduction
01:21 - The Small File Problem
05:17 - Experiment small vs big files, no schema vs schema
11:52 - Discover the Small file problem in the Spark UI
15.44 - Reduce Partition Size to solve Small file Problem
24:49 - Summary
#spark #pyspark #dataengineering #dataengineeringessentials