Enterprise hardware here: https://clck.ru/3NyVEp
Promo code: DONEX
Advertisement from SERVER CENTER LLC
erid: 2VtzqwFWSrW
I ran large neural networks on a home server without a single graphics card: two Xeon E5-2680v4, 128 GB DDR4, and llama.cpp. I'm checking what the processor can really handle—from a Qwen3-Coder 30B to a gpt-oss 120B and a Qwen3-235B. I'm measuring generation speed and understanding what really matters: MoE, quantization, and NUMA. I'm comparing it to an RTX 3090 and honestly showing where the older server excels in capacity and where it loses out in speed. Local ChatGPT without a subscription, without a VPN, and without the internet—we're finding out who really needs it and how much it costs.
🛒 SuperMicro SP3 motherboard:
https://ali.click/afgtf1u?erid=2SDnje...
https://ali.click/52ddh1f?erid=2SDnje...
🛒 Small and budget-friendly SuperMicro for starting a home lab:
https://ali.click/fma8718?erid=2SDnjd...
🛒 Motherboard + CPU + RAM kit (very close to my first build on the channel)
https://ali.click/g6dah1y?erid=2SDnjd...
🛒 Dual-socket with 64GB DDR4:
https://ali.click/1pdah16?erid=2SDnjd...
🛒 Dual-socket similar to the one from this video, but with lower memory speeds 2133. (It's very expensive, only buy it if you really want a ready-made kit with dual-socket 128GB. Although I'd already consider a Supermicro for around 20k with dual-socket and buy your own RAM on Avito. 2680v4 RAM is currently around 1k-1.5k each, and you can also find it on Avito.
https://ali.click/j6eah10?erid=2SDnjd...
You can support the channel with a small donation on Boosty. For subscribers of the Junior Developer level and above, videos are available earlier than on YouTube. Join us:
https://boosty.to/donex
Telegram subscribers get a little more content, and I sometimes post information that doesn't fit the channel's format. So, join us :)
Our Telegram channel: https://t.me/DonExCode
Our Telegram chat: https://t.me/DonExCodeChat
Our Discord: / discord
00:00 - What's the video about
00:59 - Server configuration
02:19 - Enterprise hardware
03:27 - dense and MoE models
05:22 - advantage over a video card
06:26 - quantization
07:25 - server setup and NUMA
09:24 - benchmarks
09:41 - GPT oss 120B benchmark
10:29 - Qwen3.5 122B benchmark
10:51 - GLM 4.5 Air 106B benchmark
11:06 - number of active parameters
11:41 - models for 64GB RAM
14:04 - comparison with the RTX 3090
14:46 - Launching qwen3 235B
15:28 - The quantization problem
16:08 - Tests on specific tasks
16:41 - Qwen3 Coder 30B
17:22 - GLM 4.5 106B
18:22 - GPT OSS 120B
17:14 - Qwen3 235B
20:38 - Conclusions on model performance
21:59 - Nuances of local LLMs
23:08 - System bottleneck
23:37 - At what cost?
26:22 - Why do we need local LLMs?
29:04 - Sponsor badges