Brief Summary
Jupyter + Spark is a powerful combination in your data science toolkit to handle big data. In this talk, we explore running Jupyter on a cluster of machines in the cloud. Using PySpark, we construct a pipeline for machine learning (LDA Topic Modeling) on 2.8 million news articles on COVID-19.
Outline
In recent years, cloud providers (AWS, Azure, GCP) have simplified deploying and managing Jupyter Notebooks on Spark clusters. Giving notebooks with computational powers of Spark clusters unlocks data scientists to handle today’s explosion of big data. As an example of how to use Jupyter + Spark to tackle real-world big data problems, I will show you how I analyze and built machine learning models on all news articles available online relating to COVID-19.
In this talk, we will cover: 1) How to spin up a Spark cluster with Jupyter Notebook on Google Cloud; 2) Using 16-node cluster with 1TB+ RAM for LDA topic modeling; 3) Visualizing Spark results in Jupyter Notebook.
This talk is designed for intermediate users of Jupyter with novice users in mind. To get the most out of the talk, attendees should have a general knowledge of distributed computing and cloud computing.
----
JupyterCon brings together data scientists, business analysts, researchers, educators, developers, core Project contributors, and tool creators for in-depth training, insightful keynotes, networking, and practical talks exploring the Project Jupyter ecosystem.
https://jupytercon.com/
JupyterCon is possible thanks to the generous support of our sponsors, and the labor of many volunteer organizers.
https://jupytercon.com/sponsors/
https://jupytercon.com/about/#Organiz...

jupytercon.com
JupyterCon2020
JupyterCon 2020

jupytercon.com
JupyterCon2020
JupyterCon 2020

jupytercon.com
JupyterCon2020
JupyterCon 2020