SFBigAnalytics_20200116: Elastic Distributed DL Training at large scale on-prem & cloud productions

Опубликовано: 23 Апрель 2026
на канале: SF Big Analytics
152
3

Elastic Distributed Deep Learning Training at large scale on-prem and cloud productions

In this talk, we like to introduce a new high performance distributed deep learning training engine. It transforms static monolithic training into a dynamic resilient process - automatically scales up and down GPU allocation while transparently training the models developed in popular frameworks such as TensorFlow, Pytorch and Caffe.

The architecture and design of dynamic distributed runtime management
Transparent dynamic model scaling (up and down). Transparency means no code change or minimum change (1 line) to models developed in Tensorflow, Pytorch.
Experiences with large scale on-prem and cloud deployments with different (QoS) policies
hyper-parameters

Speaker: Yonggang Hu

Yonggang is IBM Distinguished Engineer, Chief Architect at Spectrum Computing, IBM Systems.He has been working on distributed computing, HPC, grid, cloud and big data analytics for the past 20 years. He is currently focusing on AI runtime at IBM and responsible for roadmap and strategy of IBM Watson Machine Learning Accelerator.