Speaker: Anush Sankaran, Deeplite.ai (https://www.deeplite.ai/)
Host: Nishant Sinha, OffNote Labs (https://offnote.co)
Talk details: https://offnote.substack.com/p/the-sc...
Designing deep learning-based solutions is becoming a race for training deeper models with a greater number of layers. While a large-size deeper model could provide competitive accuracy, it creates a lot of logistical challenges and unreasonable resource requirements during development and deployment. This is one of the key reasons for deep learning models not being very popular in various production environments, especially in edge devices. In this talk, I will talk about the different deep learning model compression techniques and the challenges involved in production ready model optimization. Particularly, I will focus on how at Deeplite, we automate metric-driven automated model compression using a network-in-network approach.
===
Anush Sankaran is currently working as a Senior Research Scientist with Deeplite, Canada. Anush is extremely passionate about democratization of deep learning and works on various usable solutions in bringing deep learning to the hands of all consumers. Anush has co-authored more than 10 journal papers, 20 conference publications, 8 patents, and 2 book chapters.
===
The OffNote Labs AI Talk Series brings you industry experts, researchers and practitioners, passionate to share their learnings and experience -- on innovating and building cutting-edge AI technology / systems which touch and influence the lives of billions of people on this planet.
Our talks are informal, a blend of traditional presentation / podcast, and the audience very technically engaged.
We take delight in unraveling the experience of research and innovation, and celebrate innovators who 'follow the problem', and make complex technology work in the wild.
==
Follow OffNote Labs on LinkedIn. / offnote
Watch previous talks and Subscribe on Youtube: https://bit.ly/31VTMHH
Sign up to our newsletter: http://offnote.substack.com
Read our research articles: / offnote
Web: https://offnote.co
==
00:00 Introduction and Overview
01:46 Deep Learning Drives AI
02:40 Deep Learning Models are Growing Rapidly
04:00 Revolution of Depth
04:53 Deep Learning Models in Production
07:33 What We Need to Achieve in Production
08:26 Time to Deploy AI on Edge Devices
15:20 Challenge #1: Democratization of Deep Learning
18:40 Challenge #2: Multiple Metrics to Optimize
21:33 Challenge #3: Multiple Hardware Support
22:40 Challenge #4: Blackbox Framework
23:33 Challenge #5: Research Paper to Production
24:13 Idea of Optimization
27:33 How do we Optimize Deep Learning Models?
28:40 Parameter Removal- Weight Pruning, Column Pruning/Shape Pruning
34:26 Parameter Removal- Layer Pruning
36:53 Parameter Removal
47:20 Parameter Search
51:33 Parameter Decomposition
54:00 Parameter Quantization
1:02:26 Deeplite- Introduction
1:03:06 Deeplite- Add Optimization to Current Pipelines
1:04:13 Deeplite- Black-Box Optimization
1:06:00 Deeplite- Neutrino Engine
1:10:53 How much to Optimize?
1:13:06 Results on Popular Models
1:16:53 Results on Large Scale Datasets (Resnet18)
1:17:33 Time Take for Optimization
1:19:33 Deep Learning Models in Production: Before Deeplite
1:20:53 Deep Learning Models in Production: Deeplite
1:22:26 Industry Use Case #1
1:23:46 Industry Use Case #2
1:24:13 Industry Use Case #3
1:32:40 Deep Learning Models in Production - With Deeplite AI Optimization
1:39:20 Deeplite Team and Conclusion
===
#modelcompression #edgeml #tinyml #quantization #onnx #ParameterQuantization #ParameterDecomposition #ParameterSearch #ParameterRemoval #LayerPruning #WeightPruning #ColumnPruning #ShapePruning #DLOptimization
#DemocratizationofDeepLearning #DeployAI #EdgeDevices #EdgeML
#ModelsinProduction #NeutrinoEngine #deeplite #compressmodel #ONNX
#MLOPS