In this video, we cover the paper "LongNet: Scaling Transformers to 1,000,000,000 Tokens", with a focus on explaining the novel attention mechanism called dilated attention.
We start by discussing the importance of a long sequence length and the limitations of existing attention mechanisms in modeling very long sequences.
We then explain the key factor that helps LongNet to scale up the sequence length so dramatically, which is dilated attention.
We review how a dilated attention block works, explaining the input segmentation and the dilation rate hyperparameter. Then we move to see how mixture of dilated attention blocks are used in order to capture both long-range and short-range information.
We finish by explaining the multi-head dilated attention block which further helps to diversify the information extracted from the input sequence.
👍 Please like & subscribe if you enjoy this content
----------------------------------------------------------------------------------
Support us - https://paypal.me/aipapersacademy
----------------------------------------------------------------------------------
Paper page on arxiv - https://arxiv.org/abs/2307.02486
Blog post - https://aipapersacademy.com/longnet/
Chapters:
0:00 Introduction
0:32 Long Context Importance
1:56 Dilated Attention
3:52 Mixture of Dilated Attention
4:51 Multi-Head Dilated Attention