How much RAM do you really need to run a local LLM on a Mac mini, with its full context window? I filled two 4-bit models' context windows to the brim on an M5 Max MacBook Pro, measured memory 10 times a second, and compared what each model needs with what each Mac mini gives its GPU. Every concept is explained from scratch, no background needed.
THE SHORT ANSWER
• Ornith 1.5 (35B mixture of experts): needs about 25.7 GB with its whole 262K-token window full. The 48 GB M5 Pro Mac mini holds that with over 10 GB to spare. The 32 GB Mac mini gives its GPU about 25 GB: enough for 128,000 tokens (about 23 GB), just short of the whole window.
• Qwen 3.8 27B (dense): loads in 16 GB but needs about 35 GB at full context, because its notes (the KV cache) are three times bigger than Ornith's. It just squeezes into the 48 GB Mac mini; the 64 GB gives it real room.
• Speed on a 32 GB M6 Mac mini: people report a little over 9 tokens a second for Qwen (builds that predict a few tokens at once post around 20), and 40 to 65 for Ornith on short prompts. Same memory, same chip.
• Next year (my extrapolation, not a promise): on most tests I found, small open models trail the best by 5 to 11 months, sometimes a year or more. If that holds, by next summer a dense model like Qwen should score about where the best models score today.
WHAT YOU'LL LEARN
Parameters and 4-bit quantization · tokens · the context window · the KV cache · unified memory and the GPU's share · memory bandwidth · dense vs mixture-of-experts (MoE) models · what benchmarks measure · why a smaller model can need more memory · why the first word takes minutes at full context
CHAPTERS
0:00 How much RAM does local AI need?
0:32 What running AI locally means
0:53 Parameters & 4-bit quantization explained
1:22 Tokens & the context window explained
1:53 The KV cache explained
2:10 Unified memory & the GPU’s share
2:31 Dense models & memory bandwidth
3:08 Mixture of experts (MoE) explained
3:37 Qwen 3.8 27B vs Ornith 1.5 benchmarks
3:51 The test: full context on an M5 Max
4:14 Why the smaller model needs more memory
4:24 Which Mac mini fits which model
4:59 The wait before the first word
5:21 What will a Mac mini run next year?
5:49 How far local models trail the best
6:19 The forecast: local AI by next summer
6:37 The answer: which Mac mini to buy
HOW I MEASURED
MacBook Pro M5 Max, 128 GB · 4-bit MLX builds of both models · each window filled to 128K tokens and all the way (261,600 tokens plus a 512-token answer) · memory = the process's physical footprint, sampled every 100 ms, with MLX handing freed memory straight back · memory in GB as macOS shows it · GPU share: 78 % of a 32 GB Mac (one published measurement), 84 % on my Mac; the 48 and 64 GB shares are estimates
SOURCES
Community speeds: omlx.ai benchmarks · Terminal-Bench 2.1: Artificial Analysis · brand-new bugs: SWE-rebench (Nebius) · SWE-bench Verified review: Epoch AI · vendor scores: each model's own card
#LocalAI #MacMini #LLM