If you have an 8GB VRAM graphics card and a computer that's not a total relic, what model should you run? It's Qwen 3.6 35B-A3B.
llama-server \
--host 0.0.0.0 --port 8091 \
-hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-IQ3_XXS \
-c 131072 -ngl 99 --n-cpu-moe 25 \
-b 2048 -ub 2048 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--cache-ram 0 --parallel 1 --jinja \
--reasoning on --reasoning-budget 8192 \
--temp 0.7 --top-k 20 --top-p 0.8 --min-p 0.0 \
--presence-penalty 0.0 --repeat-penalty 1.0 \
--n-predict 32768
Ask a capable agent to set up the 'pi' coding agent to hit the above local server; or use whatever you're comfortable with.
0:00 8GB VRAM + 12GB RAM
0:14 Recommended config
0:59 MoE offload with --n-cpu-moe
2:30 Hardware and prices
4:01 Speed
5:15 Hard task: ComfyUI agent
7:08 Livestream footage
7:55 Bob-bench Wide
9:55 Methodology
12:22 Honorable mentions: REAP, IBM Granite, Bonsai 2, Gemma 4
13:51 Future work
16:20 Wrapup
Recommended configuration, analysis, what to expect, how I reached this conclusion, and next steps.
No BS, no wasting your time, some bad jokes included.
[email protected]
twitter: @rskjr
/ discord