The Magic behind Parquet File Statistics and Spark - Load Big Data Efficiently (Part 6)

Опубликовано: 24 Июль 2026
на канале: Data with Nikk the Greek
247
3

Predicate Pushdown makes Spark extremely efficient! Have you wondered how Spark decides which data to load when using Predicate Pushdown with Parquet?

This video explores the Parquet statistics and provides you with the connecting dots. Check it out :)

~~~~~~~~~~~~~ Subscribe - Like - Comment - Challenge ~~~~~~~~~~~~~

You want to master Data Engineering with PySpark? Subscribe here: https://www.youtube.com/@DataNikktheG...

Feel free to comment or challenge my explanations as always. Happy to learn also myself more by the community.

~~~~~~~~~~~~~~~~~~~~~~~ Resources ~~~~~~~~~~~~~~~~~~~~~~~

Link to Slides: https://github.com/datanikkthegreek/S...

Link to code: https://github.com/datanikkthegreek/S...

Link to Parquet logs with Parquet-tools: https://github.com/datanikkthegreek/S...

~~~~~~~~~~~~~~~~~~~~~~~ Chapters ~~~~~~~~~~~~~~~~~~~~~~~

00:00 - Introduction
00:34 - Recap Parquet and Predicate Pushdown
05:17 - File Meta Data
09:47 - Row Group Meta Data & Partitions
13:17 - Column Statistics & Experiments

#spark #pyspark #dataengineering #dataengineeringessentials