━━━━━━━━━━━━━━━━━━
💻 GitHub → https://github.com/tonyd2wild
🐦 Follow on X → https://x.com/Tech2Wild
━━━━━━━━━━━━━━━━━━
GLM 5.2 is officially running locally on 4 DGX Sparks.
REPO - https://github.com/tonyd2wild/GLM-5.2...
In this video, I’m breaking down how I got GLM 5.2 running on my 4x DGX Spark setup, including the larger 655K context version and the faster 200K context recipe I dropped on GitHub.
This is one of the biggest local AI tests I’ve done so far because GLM 5.2 is currently one of the strongest open models out, especially for coding, agent work, planning, and orchestration. Being able to run something this capable locally at home is wild.
We go over:
GLM 5.2 running on 4x DGX Sparks
655K+ context testing
200K context repo version
Around 28–30 tok/s on single prompts
Multi-agent / concurrent testing
4 concurrent instances around 17–20 tok/s
Fast time-to-first-token compared to other local models
Hermes and OpenClaw testing
Why this matters for local AI workflows
Whether it makes more sense to run 4 Sparks on one big model or split them into 2 + 2
How this compares to DeepSeek V4 Flash, MiniMax M3, MiMo V2.5, and Qwen 3.6 models
I’m still running a hybrid system with cloud frontier models and local models, but more and more of my real work is being pushed local.
This is a major step forward for my setup, and honestly, it feels like another sign that the future is local.
Repo:
https://github.com/tonyd2wild/GLM-5.2...
Drop your questions below and let me know what you want me to test next.
#GLM52 #DGXSpark #LocalAI #AIModels #RTX3090 #DeepSeek #Qwen #MiniMax #MiMo #AIAgents #Homelab #NVIDIA #Tech2Wild