Brad Miro - Apache Spark 3.0: Big Data Analytics with GPUs and Jupyter Notebooks | JupyterCon 2020

Опубликовано: 16 Март 2026
на канале: JupyterCon
199
2

Brief Summary
Apache Spark is an open source engine for performing data analysis on large amounts of data. Apache Spark 3.0 provides major improvements to its SQL processing capabilities, optimizations for Pandas, and first-class support for GPUs. Apache Spark can be used for exploratory data analysis, data processing or machine learning. In this talk, you’ll learn how to do all of this from a Jupyter Notebook.

Outline
Intro to Apache Spark: History, What it is + How it works, Ecosystem, Data source interoperability, Runtime environments. What’s New in 3.0: GPU compatibility, Efficiency improvements of the Spark engine. Spark + ML: Utilize GPUs, options for available tools. Spark + Jupyter: Discussion of different kernels, options for available tools. Demo: Spark with Jupyter notebooks.
----
JupyterCon brings together data scientists, business analysts, researchers, educators, developers, core Project contributors, and tool creators for in-depth training, insightful keynotes, networking, and practical talks exploring the Project Jupyter ecosystem.

https://jupytercon.com/ 

JupyterCon is possible thanks to the generous support of our sponsors, and the labor of many volunteer organizers. 

https://jupytercon.com/sponsors/ 
https://jupytercon.com/about/#Organiz...


jupytercon.com
JupyterCon2020
JupyterCon 2020


jupytercon.com
JupyterCon2020
JupyterCon 2020


jupytercon.com
JupyterCon2020
JupyterCon 2020