Qwen on Apple M4

Can I run Qwen2.5 72B on an Mac Mini M4 Pro 48GB?

Not at Q4_K_M — it needs 44.3 GB against 36.0 GB available. Drop to IQ3_M and it fits, at about 5.29 tokens per second.

Won't fit

44.3 GB of 36.0 GB · 123%
046 GB
Weights 41.2 GB
KV cache 2.5 GB
Runtime overhead 0.6 GB
Over the limit 8.3 GB

Short by 8.3 GB. You can run it with 63 of 80 layers on the Mac Mini M4 Pro 48GB and the rest in system RAM, at roughly 2.43 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation2.43tok/s
Prompt processing49.1tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Qwen2.5 72B on a Mac Mini M4 Pro 48GB

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 135.4 GB 138.5 GB Won't fit 0.37
INT8 / W8A8 8.50 71.4 GB 74.5 GB Won't fit 0.86
Q8_0 (GGUF) 8.50 71.4 GB 74.5 GB Won't fit 0.86
FP8 (E4M3) 8.00 67.7 GB 70.8 GB Won't fit 0.94
Q6_K 6.56 55.5 GB 58.6 GB Won't fit 1.31
Q5_K_M 5.67 48.0 GB 51.1 GB Won't fit 1.73
Q5_K_S 5.52 46.7 GB 49.8 GB Won't fit 1.84
Q4_K_M 4.85 41.2 GB 44.3 GB Won't fit 2.43
AWQ 4-bit 4.25 39.4 GB 42.5 GB Won't fit 2.74
GPTQ 4-bit 4.25 39.4 GB 42.5 GB Won't fit 2.74
MXFP4 4.25 39.4 GB 42.5 GB Won't fit 2.74
Q4_K_S 4.58 39.0 GB 42.1 GB Won't fit 2.84
Q4_0 4.55 38.8 GB 41.9 GB Won't fit 2.86
IQ4_XS 4.25 36.3 GB 39.4 GB Won't fit 3.51
Q3_K_M 3.91 33.6 GB 36.6 GB Won't fit 6K 4.65
IQ3_M 3.70 31.8 GB 34.9 GB Fits, but tight 11K 5.29
IQ3_XXS 3.06 26.6 GB 29.7 GB Runs comfortably 28K 6.29
Q2_K 2.63 23.1 GB 26.2 GB Runs comfortably 39K 7.19
IQ2_XXS 2.06 18.4 GB 21.5 GB Runs comfortably 54K 8.89
IQ1_M 1.75 15.9 GB 19.0 GB Runs comfortably 62K 10.2

Also worth checking