Qwen · Qwen3 · 8.2B parameters

Qwen3 8B VRAM requirements

Qwen3 8B has 36 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 4.7 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.

Runs comfortably

6.7 GB of 21.8 GB · 31%
022 GB
Weights 4.7 GB
KV cache 1.1 GB
Runtime overhead 0.8 GB

Qwen3 8B at Q4_K_M leaves 15.1 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation112tok/s
Prompt processing4231tok/s
Max context115Ktokens
KV per 1K tokens0GB

Every quantisation of Qwen3 8B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 15.3 GB 17.2 GB Runs comfortably 40K 37.9
INT8 / W8A8 8.50 7.8 GB 9.8 GB Runs comfortably 93K 71.0
Q8_0 (GGUF) 8.50 7.8 GB 9.8 GB Runs comfortably 93K 71.0
FP8 (E4M3) 8.00 7.6 GB 9.6 GB Runs comfortably 95K 72.6
Q6_K 6.56 6.3 GB 8.2 GB Runs comfortably 104K 87.0
AWQ 4-bit 4.25 5.8 GB 7.7 GB Runs comfortably 108K 93.8
GPTQ 4-bit 4.25 5.8 GB 7.7 GB Runs comfortably 108K 93.8
MXFP4 4.25 5.8 GB 7.7 GB Runs comfortably 108K 93.8
Q5_K_M 5.67 5.4 GB 7.4 GB Runs comfortably 110K 99.1
Q5_K_S 5.52 5.3 GB 7.2 GB Runs comfortably 111K 101
Q4_K_M 4.85 4.7 GB 6.7 GB Runs comfortably 115K 112
Q4_K_S 4.58 4.5 GB 6.4 GB Runs comfortably 117K 116
Q4_0 4.55 4.5 GB 6.4 GB Runs comfortably 117K 117
IQ4_XS 4.25 4.2 GB 6.2 GB Runs comfortably 119K 123
Q3_K_M 3.91 4.0 GB 5.9 GB Runs comfortably 121K 130
IQ3_M 3.70 3.8 GB 5.7 GB Runs comfortably 122K 135
IQ3_XXS 3.06 3.3 GB 5.2 GB Runs comfortably 126K 152
Q2_K 2.63 2.9 GB 4.9 GB Runs comfortably 128K 167
IQ2_XXS 2.06 2.5 GB 4.4 GB Runs comfortably 128K 192
IQ1_M 1.75 2.2 GB 4.2 GB Runs comfortably 128K 208

Qwen3 8B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 405
A100 80GB 80 2039 Runs comfortably 128K 230
RTX 5090 32 1792 Runs comfortably 128K 214
RTX 5080 16 960 Runs comfortably 62K 117
RTX 4090 24 1008 Runs comfortably 115K 112
RTX 5070 Ti 16 896 Runs comfortably 62K 109
RTX 3090 24 936 Runs comfortably 115K 108
Radeon RX 7900 XTX 24 960 Runs comfortably 115K 101
L40S 48 864 Runs comfortably 128K 96.1
RTX A6000 48 768 Runs comfortably 128K 89.3
RTX 3080 10GB 10 760 Runs comfortably 22K 88.4
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 87.9
RTX 5070 12 672 Runs comfortably 35K 82.6
RTX 4080 Super 16 736 Runs comfortably 62K 82.1
RTX 4070 Ti Super 16 672 Runs comfortably 62K 75.1
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 65.6
RTX 4070 Super 12 504 Runs comfortably 35K 56.5
RTX 4070 12 504 Runs comfortably 35K 56.5
RTX 5060 Ti 16GB 16 448 Runs comfortably 62K 55.4
Arc B580 12 456 Runs comfortably 35K 43.4
RTX 3060 12GB 12 360 Runs comfortably 35K 42.3
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 33.9
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 33.0
RTX 4060 Ti 16GB 16 288 Runs comfortably 62K 32.5
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 25.3

Architecture

Parameters8.2B
Layers36
Hidden size4096
Attention heads / KV heads32 / 8
Head dimension128
Vocabulary151,936
Trained context128K
KV cache per 1K tokens0 GB
Hugging FaceQwen/Qwen3-8B

The Qwen family

Alibaba's series, and the broadest size ladder available — Qwen3 runs from 0.6B to 32B dense, plus 30B-A3B and 235B-A22B as mixture-of-experts. The 152k vocabulary makes the embedding table a large share of a small model's file. Qwen2.5-Coder is the same architecture trained for code.

huggingface.co/Qwen · qwenlm.github.io · all 14 Qwen models

Direct answers