Qwen on NVIDIA Ampere

Can I run Qwen2.5 72B on an RTX A6000?

Not at Q4_K_M — it needs 44.6 GB against 44.3 GB available. Drop to Q4_K_S and it fits, at about 11.9 tokens per second.

Won't fit

44.6 GB of 44.3 GB · 101%
046 GB
Weights 41.2 GB
KV cache 2.5 GB
Runtime overhead 0.9 GB
Over the limit 0.3 GB

Short by 0.3 GB. You can run it with 79 of 80 layers on the RTX A6000 and the rest in system RAM, at roughly 9.90 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation9.90tok/s
Prompt processing217tok/s
Max context7Ktokens
KV per 1K tokens0GB

Every quantisation of Qwen2.5 72B on a RTX A6000

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 135.4 GB 138.8 GB Won't fit 0.40
INT8 / W8A8 8.50 71.4 GB 74.8 GB Won't fit 1.14
Q8_0 (GGUF) 8.50 71.4 GB 74.8 GB Won't fit 1.14
FP8 (E4M3) 8.00 67.7 GB 71.1 GB Won't fit 1.29
Q6_K 6.56 55.5 GB 58.9 GB Won't fit 2.10
Q5_K_M 5.67 48.0 GB 51.4 GB Won't fit 3.67
Q5_K_S 5.52 46.7 GB 50.1 GB Won't fit 4.20
Q4_K_M 4.85 41.2 GB 44.6 GB Won't fit 7K 9.90
AWQ 4-bit 4.25 39.4 GB 42.8 GB Fits, but tight 13K 11.8
GPTQ 4-bit 4.25 39.4 GB 42.8 GB Fits, but tight 13K 11.8
MXFP4 4.25 39.4 GB 42.8 GB Fits, but tight 13K 11.8
Q4_K_S 4.58 39.0 GB 42.4 GB Fits, but tight 14K 11.9
Q4_0 4.55 38.8 GB 42.2 GB Fits, but tight 15K 11.9
IQ4_XS 4.25 36.3 GB 39.7 GB Runs comfortably 23K 12.7
Q3_K_M 3.91 33.6 GB 36.9 GB Runs comfortably 32K 13.7
IQ3_M 3.70 31.8 GB 35.2 GB Runs comfortably 37K 14.4
IQ3_XXS 3.06 26.6 GB 30.0 GB Runs comfortably 54K 17.1
Q2_K 2.63 23.1 GB 26.5 GB Runs comfortably 65K 19.6
IQ2_XXS 2.06 18.4 GB 21.8 GB Runs comfortably 80K 24.1
IQ1_M 1.75 15.9 GB 19.3 GB Runs comfortably 88K 27.7

Also worth checking