Qwen · Qwen3 · 32.8B parameters

Qwen3 32B VRAM requirements

Qwen3 32B has 64 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 18.6 GB, and the best quantisation that fits a 24 GB card is Q4_K_M.

Fits, but tight

21.5 GB of 21.8 GB · 99%
022 GB
Weights 18.6 GB
KV cache 2.0 GB
Runtime overhead 0.8 GB

This fits with almost nothing to spare. A background application claiming VRAM will push it over. Drop to the next quantisation down, or quantise the KV cache to Q8_0 — that halves the cache for no meaningful quality loss.

Generation30.5tok/s
Prompt processing1058tok/s
Max context9Ktokens
KV per 1K tokens0GB

Every quantisation of Qwen3 32B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 61.0 GB 63.9 GB Won't fit 0.85
INT8 / W8A8 8.50 32.1 GB 34.9 GB Won't fit 2.52
Q8_0 (GGUF) 8.50 32.1 GB 34.9 GB Won't fit 2.52
FP8 (E4M3) 8.00 30.5 GB 33.3 GB Won't fit 2.83
Q6_K 6.56 25.0 GB 27.9 GB Won't fit 4.94
Q5_K_M 5.67 21.6 GB 24.5 GB Won't fit 9.34
Q5_K_S 5.52 21.1 GB 23.9 GB Won't fit 10.4
Q4_K_M 4.85 18.6 GB 21.5 GB Fits, but tight 9K 30.5
AWQ 4-bit 4.25 18.3 GB 21.2 GB Fits, but tight 10K 30.9
GPTQ 4-bit 4.25 18.3 GB 21.2 GB Fits, but tight 10K 30.9
MXFP4 4.25 18.3 GB 21.2 GB Fits, but tight 10K 30.9
Q4_K_S 4.58 17.6 GB 20.5 GB Fits, but tight 13K 32.0
Q4_0 4.55 17.5 GB 20.4 GB Fits, but tight 14K 32.2
IQ4_XS 4.25 16.4 GB 19.3 GB Runs comfortably 18K 34.2
Q3_K_M 3.91 15.2 GB 18.0 GB Runs comfortably 23K 36.8
IQ3_M 3.70 14.4 GB 17.3 GB Runs comfortably 26K 38.6
IQ3_XXS 3.06 12.1 GB 15.0 GB Runs comfortably 35K 45.3
Q2_K 2.63 10.6 GB 13.4 GB Runs comfortably 41K 51.3
IQ2_XXS 2.06 8.5 GB 11.3 GB Runs comfortably 50K 62.2
IQ1_M 1.75 7.4 GB 10.2 GB Runs comfortably 54K 70.4

Qwen3 32B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 114
A100 80GB 80 2039 Runs comfortably 128K 63.5
RTX 5090 32 1792 Runs comfortably 39K 59.0
RTX 4090 24 1008 Fits, but tight 9K 30.5
RTX 3090 24 936 Fits, but tight 9K 29.5
Radeon RX 7900 XTX 24 960 Fits, but tight 9K 27.6
L40S 48 864 Runs comfortably 99K 26.2
RTX A6000 48 768 Runs comfortably 99K 24.3
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 23.9
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 17.8
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 9.17
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 67K 8.93
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 6.84
RTX 5080 16 960 Won't fit 4.98
RTX 5070 Ti 16 896 Won't fit 4.95
RTX 5060 Ti 16GB 16 448 Won't fit 4.50
RTX 4080 Super 16 736 Won't fit 4.39
RTX 4070 Ti Super 16 672 Won't fit 4.34
RTX 4060 Ti 16GB 16 288 Won't fit 3.71
RTX 5070 12 672 Won't fit 3.40
RTX 3060 12GB 12 360 Won't fit 3.06
RTX 4070 Super 12 504 Won't fit 3.02
RTX 4070 12 504 Won't fit 3.02
RTX 3080 10GB 10 760 Won't fit 2.79
Arc B580 12 456 Won't fit 2.54

Architecture

Parameters32.8B
Layers64
Hidden size5120
Attention heads / KV heads64 / 8
Head dimension128
Vocabulary151,936
Trained context128K
KV cache per 1K tokens0 GB
Hugging FaceQwen/Qwen3-32B

The Qwen family

Alibaba's series, and the broadest size ladder available — Qwen3 runs from 0.6B to 32B dense, plus 30B-A3B and 235B-A22B as mixture-of-experts. The 152k vocabulary makes the embedding table a large share of a small model's file. Qwen2.5-Coder is the same architecture trained for code.

huggingface.co/Qwen · qwenlm.github.io · all 14 Qwen models

Direct answers