Data Wrangling with PySpark for Data Scientists Who Know Pandas - Andrew Ray

Опубликовано: 06 Август 2026
на канале: Databricks
143,046
2.6k

"Data scientists spend more time wrangling data than making models. Traditional tools like Pandas provide a very powerful data manipulation toolset. Transitioning to big data tools like PySpark allows one to work with much larger datasets, but can come at the cost of productivity.

In this session, learn about data wrangling in PySpark from the perspective of an experienced Pandas user. Topics will include best practices, common pitfalls, performance consideration and debugging.

Session hashtag: #SFds12

Learn more:
Developing Custom Machine Learning Algorithms in PySpark
https://databricks.com/blog/2017/08/3...

Introducing Pandas UDF for PySpark
https://databricks.com/blog/2017/10/3...

Best Practices for Running PySpark
https://databricks.com/session/best-p...

Session Overview:
Why?
What Do i get with pyspark?
Primer
Important Concepts
Architecture
Setup
Run
Load CSV
View Dataframe
Rename Columns
Drop Column
Filtering
Add Column
Fill Nulls
Aggregation
Standard Transformations
Keep it in the JVM
Row Conditional Statements
Python when Required
merge/join dataframes
Pivot table
Summary Statistics
histogram
SQL
Make sure to
Things not to do
If things go wrong
Thank you

About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business.
Read more here: https://databricks.com/product/unifie...

Connect with us:
Website: https://databricks.com
Facebook:   / databricksinc  
Twitter:   / databricks  
LinkedIn:   / databricks  
Instagram:   / databricksinc   Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-nam...