Demystifying Data Pre-processing & Data Wrangling for Data Science | Pariza Kamboj

Опубликовано: 26 Июнь 2026
на канале: WiDS Worldwide
478
10

In the current era, Data Science is rapidly evolving and proving very decisive in ERP (Enterprise Resource Planning). The dataset required for building the analytical model using data science, is collected from various sources such as Government, Academic, Web Scraping, API’s, Databases, Files, Sensors and many more. We cannot use such real-world data for analysis process directly because it is often inconsistent, incomplete, and more likely to contain bulk errors. We often hear the phrase “garbage in, garbage out”. Dirty data or messy data riddled with inaccuracies and errors, result in a bad/improperly trained model which in turn might result in poor business decisions and sometimes even hazardous to the domain. Any powerful algorithm is failed in providing correct analysis when applied to bad data. Therefore, data must be curated, cleaned and refined to be used in data science and products based on data science. To perform these tasks, “Data Preparation” is required which includes two methods that are: Data Pre-processing, and Data Wrangling. Most data scientists spend the majority of their time in data preparation.

This workshop was conducted by Pariza Kamboj, Professor at Sarvajanik College of Engineering & Technology (SCET).

Useful resources for this workshop:
https://bit.ly/jupyter_code
https://bit.ly/cars3_dataset
https://bit.ly/execution_google_colab
https://bit.ly/anaconda_installation_...

Learn more about WiDS Workshops: widsconference.org/workshops

#SCET #CollegeofEngineeringandTechnology #Stanford #StanfordUniversity #WiDS #Womenindatascience #datascience #engineering