This meetup was recorded in San Francisco on December 10, 2019.
Description:
H2O datatable is a Python package for manipulating 2-dimensional tabular data structures, aka data frames. It is close in spirit to pandas, however, we put specific emphasis on speed and big data support. As the name suggests, the package is closely related to R's data.table and attempts to mimic its core algorithms and API. H2O datatable started in 2017 as a toolkit for performing big data operations on a single-node machine, at the maximum speed possible. Such requirements are dictated by modern machine-learning applications, such as H2O Driverless AI, which need to process large volumes of data and generate many features in order to achieve the best model accuracy. In the talk we will introduce H2O datatable, focusing on its data munging and modeling capabilities, followed by the Q/A session.
Speaker Bio:
Oleksiy Kononenko: Oleksiy is a maker scientist and hacker at H2O.ai, focusing on highly optimized algorithms for machine learning and data analysis. He holds M.S., summa cum laude, and Ph.D. degrees in applied mathematics from National University of Kharkiv, Ukraine. In 2009, Oleksiy was selected as a research fellow by CERN and contributed to R&D for Large Hadron Collider and next generation of high energy particle accelerators. In 2013 he joined SLAC and Stanford University to develop high performance simulation suite for 3D multi-physics modeling. Oleksiy authored more than 60 scientific papers, was an invited speaker at major international conferences, prominent institutions and companies worldwide. In his free time he enjoys snowboarding, playing soccer and basketball, guitar and drums, What? Where? When? and Jeopardy!