Predicate Pushdown can improve your loading performance significantly by pushing down your filters to the source and reduce unnecessary data loads and thus improve your query performance. Check out how it works in Spark.
In this session we will:
Learn What Predicate Pushdown is
How Predicate Pushdown behaves in Spark especially for Parquet
You want to master Data Engineering with PySpark? Subscribe here: https://www.youtube.com/@DataNikktheG...
Feel free to comment or challenge my explanations as always. Happy to learn also myself more by the community.
If you missed the session about file formats check it out here: • How Data Formats for Big Data work - Load ...
Link to Slides: https://github.com/datanikkthegreek/S...
Link to code: https://github.com/datanikkthegreek/S...
00:00 - Intro
00:56 - Predicate Pushdown
03:26 - Parquet Recap
05:09 - Predicate Pushdown with Parquet
16:46 - Predicate Pushdown with Avro, JSON, CSV
20:05 - Behaviour of Jobs with Predicate Pushdown
24:05 - Summary
#spark #pyspark #dataengineering #dataengineeringessentials