Elastic Distributed Deep Learning Training at large scale on-prem and cloud productions
In this talk, we like to introduce a new high performance distributed deep learning training engine. It transforms static monolithic training into a dynamic resilient process - automatically scales up and down GPU allocation while transparently training the models developed in popular frameworks such as TensorFlow, Pytorch and Caffe.
The architecture and design of dynamic distributed runtime management
Transparent dynamic model scaling (up and down). Transparency means no code change or minimum change (1 line) to models developed in Tensorflow, Pytorch.
Experiences with large scale on-prem and cloud deployments with different (QoS) policies
hyper-parameters
Speaker: Yonggang Hu
Yonggang is IBM Distinguished Engineer, Chief Architect at Spectrum Computing, IBM Systems.He has been working on distributed computing, HPC, grid, cloud and big data analytics for the past 20 years. He is currently focusing on AI runtime at IBM and responsible for roadmap and strategy of IBM Watson Machine Learning Accelerator.