Stanford Seminar - Neural Networks on Chip Design from the User Perspective

Опубликовано: 27 Июль 2026
на канале: Stanford Online
2,088
40

Yu Wang
Tsinghua University

October 9, 2019
To apply neural networks to different applications, various customized hardware architectures are proposed in the past a few years to boost the energy efficiency of deep learning inference processing. Meanwhile, the possibilities of adopting emerging NVM (Non-Volatile Memory) technology for efficient learning systems, i.e., in-memory-computing, are also attractive for both academia and industry. We will briefly review our past effort on Deep learning Processing Unit (DPU) design on FPGA in Tsinghua and Deephi, and then talk about some features, i.e. interrupt and virtualization, we are trying to introduce into the accelerators from the user’s perspective. Furthermore, we will also talk about the challenges for reliability and security issues in NN accelerators on both FPGA and NVM, and some preliminary solutions for now.

View the full playlist:    • Stanford EE380-Colloquium on Computer Syst...  

0:00 Introduction
1:18 Deep Learning for Everything
3:37 The New Era is Waiting for the Next Rising Star
7:46 Why? Power Consumption and Latency Are Crucial
8:25 Development of Energy-Efficient Computing Chips
9:13 Our Previous Work: Software Hardware Co-design for Energy Efficient NN Inference System
12:00 NN Compression: Quantization
12:51 NN Compression: Pruning
13:50 Hardware Architecture - Utilization
19:11 Academic NN Accelerators (Performance vs Power)
22:05 Survey on FPGA based Inference Accelerators
24:14 Application Scenarios: Cloud, Edge, Terminal
26:02 Growing of Computation Power
26:41 Brief Summary
27:28 CNN Greatly Benefits Basic Functions in Robotic Applications
28:12 Accelerator Interrupt for Hardware Conflicts
29:10 Interrupt Respond Latency & Extra Cost
29:56 How to Interrupt?
30:53 Virtual Instruction-Based Interrupt
33:36 DNN Inference Tasks in the Cloud
33:57 How to Support Multiple Tasks in the Cloud?
35:49 How to Support Dynamic Workload in the Cloud?
37:09 Low-overhead Reconfiguration of ISA-based Accelerator
38:02 Design Techniques
38:28 Experiments
40:07 Analysis for NN Fault Problems
42:06 Fault Model in Network Architecture Search (NAS)
43:47 Fault Tolerant Training - NAS Framework
44:24 Discovered Architecture
45:54 Bottleneck of Energy Efficiency Improvement
48:08 Conventional Encryption Incurs Massive Write Operations
48:16 Orders of differences in Write endurance and Write Latency
49:06 SFGE: Sparse Fast Gradient Encryption
49:54 Accuracy Drop vs Encryption Num and Intensity
50:51 Select Encryption Configuration for Different NNS