Have you ever wandered why some data formats should rather been used than others? Already said... The choice of format can have a significant performance on your Big Data processing.
In this episode we will understand how formats like CSV, JSON, PARQUET and AVRO work and perform related to size, write time, load time and schema.
You want to master Data Engineering with PySpark? Subscribe here: https://www.youtube.com/@DataNikktheG...
You missed the previous session? Here the link: • The Data Format Matters - Load Big Data ef...
Feel free to comment or challenge my explanations as always. Happy to learn also myself more by the community.
Link to slides: https://github.com/datanikkthegreek/S...
00:00 - Intro
01:18 - Performance evaluation
04:53 - Structures vs unstructured file formats
05:57 - Culmnar vs. row-wise file formats
09:26 - CSV
12:15 - JSON
14:30 - AVRO
16:24 - PARQUET
21:12 - Summary
#spark #pyspark #dataengineering #dataengineeringessentials