This video is a quick and simple introduction to Resilient Distributed Datasets in PySpark. We use the word count example to demonstrate transformations and actions.
The code can be found here: https://github.com/apostolos1927/RDD_...
Follow me on social media:
LinkedIn: www.linkedin.com/in/apostolos-athanasiou-9a0baa119
GitHub: https://github.com/apostolos1927/
00:00 - Intro
01:00 - Spark Architecture & Key Features of RDDs
07:18 - PySpark Count Work Frequencies
22:18 - Conclusion