Code: https://github.com/priyammaz/ManualTr...
MyTorch: https://github.com/priyammaz/MyTorch
PyTorch makes our life easy. You want Layernorm? You got it. You want a Linear layer? No problem! But there is a lot of work that goes into each of these modules that we take for granted. Our goal here today is to build a completely functional Transformer model WITH NO FRAMEWORKS! This means no Autograd, no layers no nothing. We have to implement for every operation the Forward and Backward pass and train a baby GPT2 Model!
The second goal is to actually build a fully functional Deep Learning Framework. This is my MyTorch package I am working on, you can explore it now, but full rewrites of core parts of the framework will come!
(And a little lie, we will use Cupy instead of Numpy so we can train on GPUs!)
Prereqs:
I hope you already know about transformers! There are great resources for this if you want to get started, and here is mine: • Attention is All You Need: Ditching Recurr...
The MinGPT https://github.com/karpathy/minGPT Video/Implementation by Karpathy is also a great place to start, and much of this is modeled after that to keep it as similar as possible! But instead of using PyTorch, we just use nothing to build the same thing!
This is a long one, just break it up by module and really take the time to do the math yourself its important I promise!
Timestamps:
00:00:00 - Introduction
00:00:50 - Structure of Neural Networks
00:01:30 - Backprop is just Chain Rule
00:03:00 - Overview of MyTorch
00:05:10 - Why ManualGrad when AutoGrad is Easier?
00:08:50 - What Backprop Looks Like
00:13:45 - Core Ops for Transformers
00:15:20 - Boilerplate Code
00:18:30 - What is _dict_ dunder method?
00:20:10 - Linear Layer Derivation for single sample
00:25:45 - Linear Layer Derivation w/ Batch Dimension
00:35:00 - Linear Layer Implementation
00:43:40 - Softmax Jacobian Derivation
00:53:50 - Numerical Stability in Softmax (max trick)
00:55:53 - SLOW Softmax Implementation
01:06:45 - Can we avoid the Jacobian? Reinterpret the Softmax
01:14:40 - Efficient Softmax Implementation
01:16:50 - ReLU Activation Implementation
01:20:00 - Embedding Layer Implementation
01:28:30 - Positional Embedding Implementation
01:34:05 - LayerNorm Derivation
02:08:35 - LayerNorm Implementation
02:17:07 - MultiHeadAttention Implementation (Forward Pass)
02:32:05 - MultiHeadAttention Implementation (Backward Pass)
02:53:20 - FeedForward Layer Implementation
02:57:00 - Dropout Implementation
03:02:08 - TransformerBlock Implementation
03:09:01 - SLOW Cross Entropy Derivation (assumes probs)
03:12:00 - SLOW Cross Entropy Implementation
03:16:50 - Cross Entropy on Logits Directly Derivation
03:21:20 - Cross Entropy Implementation
03:24:12 - FlattenForLLM To Make our Shapes Match
03:25:45 - Define the Neural Network Class
03:38:14 - Define the GPT2 Model
03:41:40 - SGD Optimizer Implementation
03:44:45 - Adam Optimizer Implementation
03:52:20 - Character Shakespeare Training Script
04:04:48 - Debugging
04:08:42 - Results
04:12:39 - Why ManualGrad Matters
Socials!
X / data_adventurer
Instagram / nixielights
Linkedin / priyammaz
Discord / discord
🚀 Github: https://github.com/priyammaz
🌐 Website: https://www.priyammazumdar.com/