Create First RDD(Resilient Distributed Dataset) in PySpark | PySpark 101 | Part 2 | DM | DataMaking

Опубликовано: 28 Сентябрь 2024
на канале: DataMaking
8,003
32

PySpark 101 Tutorial -    • PySpark 101 Tutorial  

====================================================================
====================================================================

Create First PySpark App on Apache Spark 2.4.4 using PyCharm | PySpark 101 |Part 1| DM | DataMaking -    • Create First PySpark App on Apache Sp...  
End to End Project using Spark/Hadoop | Code Walkthrough | Architecture | Part 1 | DM | DataMaking -    • End to End Project using Spark/Hadoop...  
Spark Structured Streaming with Kafka using PySpark | Use Case 2 |Hands-On|Data Making|DM|DataMaking -    • Spark Structured Streaming with Kafka...  
Running First PySpark Application in PyCharm IDE with Apache Spark 2.3.0 | DM | DataMaking -    • Running First PySpark Application in ...  
Access Facebook API using Python in English | Hands-On | Part 3 | DM | DataMaking -    • Access Facebook API using Python in E...  
Real-Time Spark Project |Real-Time Data Analysis|Architecture|Part 1| DM | DataMaking | Data Making -    • Real-Time Spark Project |Real-Time Da...  
Web Scraping using Python and Selenium | Scrape Facebook | Part 5 | Data Making | DM | DataMaking -    • Web Scraping using Python and Seleniu...  
End to End Project using Spark/Hadoop | Code Walkthrough | Kafka Producer | Part 2 | DM | DataMaking -    • End to End Project using Spark/Hadoop...  
Apache Zeppelin | Step-by-Step Installation Guide | Python | Notebook |DM| DataMaking | Data Making -    • Apache Zeppelin | Step-by-Step Instal...  
Create First RDD(Resilient Distributed Dataset) in PySpark | PySpark 101 | Part 2 | DM | DataMaking -    • Create First RDD(Resilient Distribute...  

====================================================================
====================================================================

Course link: https://www.udemy.com/course/spark-pr...

Some of the key points,
1. Compelete step-by-step procedure to build Cloudera Hadoop(CDH 6.3) Cluster like how you work in real-time spark project development
2. Build Cloudera Hadoop(CDH 6.3) on Google Cloud Platform(GCP) for FREE(using Free Trial Account)
3. Setting up IDE like IntelliJ for Spark with Scala and PyCharm for PySpark
4. Building Real-Time Data Processing Pipeline using Spark Structured Streaming using both Spark with Scala and PySpark
5. Apache Kafka integration with Spark Structured Streaming using both Spark with Scala and PySpark
6. Apache Cassandra and MongoDB NoSQL integration with Spark Structured Streaming using both Spark with Scala and PySpark
7. 11 hours of course duration with 46 lectures
8. Apache NiFi for Retail Data Simulator
9. End-to-End project development with source code
10. Solid introduction to technologies which are used in this project

My website: https://www.datamaking.com/

My blog: https://www.datasciencewiki.com/

PySpark 101 Tutorial:    • PySpark 101 Tutorial  

DM, DataMaking, Data Making, Data Science, Data Engineering