Summarizing a DataFrame in PySpark | min, max, count, percentile, schema

Опубликовано: 16 Июль 2026
на канале: Abhishek mamidi
1,378
28

In this video, I will show you how to summarize a Spark dataframe. You can directly extract basic statistics on large datasets instead of converting the dataframe into pandas and then exploring.

Link to the playlist "Getting started with PySpark" :    • Getting started with PySpark  

Link to "Setting up the PySpark environment on Google Colab" video:    • Setting up the PySpark environment on Goog...  

Link to the GitHub repo: https://github.com/Abhishekmamidi123/...

Check out my "Data Science guide for freshers and enthusiasts" playlist:    • My path to becoming a Data Scientist | Abh...  
I have put my 3 years of learning experience into this playlist.

Contents of this video:
00:00 - Setting up the PySpark environment
00:50 - Initialize Spark Session object
00:58 - Read data from UCI
01:22 - Summarize DataFrame using PySpark
02:07 - Shape of PySpark DataFrame
02:57 - Print schema of PySpark DataFrame
03:37 - Describe a Dataframe in PySpark
05:14 - Percentile in PySpark
07:39 - Summary and Subscribe :)

Please do like, share and subscribe to this channel and share this video with your friends. Keep learning :)

Follow me here:
LinkedIn:   / abhishekmamidi  
Blog: https://www.abhishekmamidi.com/
GitHub: http://github.com/Abhishekmamidi123
Kaggle: http://www.kaggle.com/abhishekmamidi

Tags:
abhishek mamidi, data science, machine learning, deep learning, artificial intelligence, internship, career, college, job, experience, krish naik, ai engineering, fresher, data science enthusiasts, pyspark, apache spark, python, pysparkling