Knowledge Distillation, Model Ensemble and Its Application on Visual Recognition

Опубликовано: 26 Июнь 2026
на канале: 2d3d.ai
944
16

This talk will cover the principles and mechanisms of knowledge distillation (KD), as well as several KD framework designs for practical usage, including (1) How KD works and its relationship to Label Smoothing technique; (2) Model ensemble using knowledge distillation for improving the performance; (3) An application of knowledge distillation for self-supervised representation learning.

Lecture slides: https://drive.google.com/file/d/1JbP-...

00:00 Intro - What can knowledge distillation do? Motivation Our main observations.
03:14 MEAL: Multi-Model Ensemble via Adversarial Learning (AAAI 2019)
04:37 Methods: Siamese-like architecture with two-stream networks; Similarity Measurement; Stacked Discriminators
08:04 Experiments
12:37 MEAL V2 An extension of MEAL (NeurIPS 2020 Beyond BackPropagation Workshop)
28:53 Visualization and Analysis
29:38 Is Label Smoothing Truly IncompaDble with Knowledge Distillation: An Empirical Study (ICLR 2021)
35:37 Our observations
37:23 A New Metric to Measure Erased Information Quantitatively
40:19 Results: Neural Machine Translation
42:03 What Circumstances Indeed Will Make LS Less Effective?
46:10 Discussion
51:04 Results: Image Classification
56:12 S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-bit Neural Networks via Guided Distribution Calibration (CVPR 2021)
59:30 Visualizations
01:02:08 Zhiqlang Shen Results on ImageNet
01:04:17 Takeaways
01:12:15 Future works and discussion

[Chapters were auto-generated using our proprietary software - contact us if you are interested in access to the software]


The talk is based on our recent papers:

(1) ICLR 2021 paper 'Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study’,
Project page: http://zhiqiangshen.com/projects/LS_a...

(2) AAAI 2019 paper ‘MEAL: Multi-Model Ensemble via Adversarial Learning’,
Github: https://github.com/AaronHeee/MEAL

(3) NeurIPS 2020 workshop paper: ‘Meal v2: Boosting vanilla resnet-50 to 80%+ top-1 accuracy on imagenet without tricks’,
Github: https://github.com/szq0214/MEAL-V2

(4) CVPR 2021 paper ‘S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-bit Neural Networks via Guided Distribution Calibration’,
Github: https://github.com/szq0214/S2-BNN

Lecture abstract:

Deep neural networks have recently been recognized as one of the most successful techniques for a variety of learning tasks. In addition, it is obvious that the world will be populated with intelligent devices in the not-too-distant future, which will necessitate the execution of deep models on these low-cost, low-power and inexpensive hardware platforms. In this context, efficient deep learning has emerged as a critical concept and research target, with the goal of learning compressed deep models and developing training algorithms to increase the efficiency of model representations, data and label usage, and so on. Architecture search, distillation, pruning, binarization, quantization, and their learning methods for lowering the number of parameters and processing needs of deep models are some of the ways to get compressed networks and improved representation. In this talk, I'll cover one common approach: knowledge distillation for efficient deep learning, based on our previous several studies on understanding soft label in knowledge distillation, efficient knowledge distillation design, and its wide range of applications in visual recognition, self-supervised representation learning, etc.

Presenter Bio:

Zhiqiang Shen is a Postdoctoral Researcher at Carnegie Mellon University and Mohamed bin Zayed University of Artificial Intelligence, working with Prof Marios Savvides and Prof Eric Xing. He was a joint-training PhD student at Fudan University and University of Illinois at Urbana-Champaign. Starting from early 2022, he will join Hong Kong University of Science and Technology as a Research Assistant Professor in the Department of Computer Science and Engineering (CSE). His research interests span broad areas of efficient deep learning, machine learning, computer vision, etc. He has published 30+ top-tier papers on TPAMI, IJCV, ICML, ICLR, CVPR, ICCV, ECCV, AAAI, etc. with 2600+ google scholar citations.
More information about Zhiqiang can be found at http://zhiqiangshen.com/
-------------------------
Find us at:

Newsletter for updates about more events ➜ http://eepurl.com/gJ1t-D
Sub-reddit for discussions ➜   / 2d3dai  
Discord server for, well, discord ➜   / discord  
Blog ➜ https://2d3d.ai