As language models get larger, it becomes harder and harder to run them on normal consumer hardware. One way to get around this limitation is to split a model over multiple GPUs. One easy way to do this is through the use of the Parallelformers library which I briefly cover in this video.
Parallelformers repo: https://github.com/tunib-ai/parallelf...
Demo code repo: https://github.com/mallorbc/Parallelf...
Discord:
/ discord
Donations(if you want):
Ethereum:
4CE913643909Fa3168297cC2857C0aDdAB389Ad8
Monero: 45mqN96o5JZZhwVTZHHiQJeyhp3WndiPC44hdrAmWeDGeCmaC1c45gTGh5eDUtEhx3JDGbsAnsD3VBXKdiorUhusUydLG22
Timestamps
00:00 - Intro
00:34 - Why Split A Model?
01:20 - Why Not Split A Model This Way?
02:00 - Potential Bug Found
03:17 - Two Identical GPUs? Let Me Know
03:36 - Installing Parallelformers
05:29 - Demo Code Walkthrough
08:20 - Demo Of Parallelformers
09:55 - Demo Of Bug
12:26 - Outro