Llama · Llama 3.1 · 8.0B parameters

Llama 3.1 8B VRAM requirements

Llama 3.1 8B has 32 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 4.6 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.

Runs comfortably

6.4 GB of 21.8 GB · 30%
022 GB
Weights 4.6 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB

Llama 3.1 8B at Q4_K_M leaves 15.3 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation116tok/s
Prompt processing4315tok/s
Max context128Ktokens
KV per 1K tokens0GB

Every quantisation of Llama 3.1 8B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 15.0 GB 16.8 GB Runs comfortably 48K 38.8
INT8 / W8A8 8.50 7.7 GB 9.5 GB Runs comfortably 106K 72.6
Q8_0 (GGUF) 8.50 7.7 GB 9.5 GB Runs comfortably 106K 72.6
FP8 (E4M3) 8.00 7.5 GB 9.3 GB Runs comfortably 108K 74.7
Q6_K 6.56 6.1 GB 8.0 GB Runs comfortably 118K 89.6
AWQ 4-bit 4.25 5.4 GB 7.2 GB Runs comfortably 124K 100
GPTQ 4-bit 4.25 5.4 GB 7.2 GB Runs comfortably 124K 100
MXFP4 4.25 5.4 GB 7.2 GB Runs comfortably 124K 100
Q5_K_M 5.67 5.3 GB 7.1 GB Runs comfortably 125K 102
Q5_K_S 5.52 5.2 GB 7.0 GB Runs comfortably 126K 105
Q4_K_M 4.85 4.6 GB 6.4 GB Runs comfortably 128K 116
Q4_K_S 4.58 4.4 GB 6.2 GB Runs comfortably 128K 121
Q4_0 4.55 4.4 GB 6.2 GB Runs comfortably 128K 121
IQ4_XS 4.25 4.1 GB 5.9 GB Runs comfortably 128K 127
Q3_K_M 3.91 3.8 GB 5.7 GB Runs comfortably 128K 135
IQ3_M 3.70 3.7 GB 5.5 GB Runs comfortably 128K 141
IQ3_XXS 3.06 3.2 GB 5.0 GB Runs comfortably 128K 160
Q2_K 2.63 2.8 GB 4.6 GB Runs comfortably 128K 176
IQ2_XXS 2.06 2.3 GB 4.2 GB Runs comfortably 128K 204
IQ1_M 1.75 2.1 GB 3.9 GB Runs comfortably 128K 223

Llama 3.1 8B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 422
A100 80GB 80 2039 Runs comfortably 128K 238
RTX 5090 32 1792 Runs comfortably 128K 222
RTX 5080 16 960 Runs comfortably 70K 121
RTX 4090 24 1008 Runs comfortably 128K 116
RTX 5070 Ti 16 896 Runs comfortably 70K 113
RTX 3090 24 936 Runs comfortably 128K 112
Radeon RX 7900 XTX 24 960 Runs comfortably 128K 105
L40S 48 864 Runs comfortably 128K 99.4
RTX A6000 48 768 Runs comfortably 128K 92.3
RTX 3080 10GB 10 760 Runs comfortably 25K 91.4
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 90.9
RTX 5070 12 672 Runs comfortably 40K 85.4
RTX 4080 Super 16 736 Runs comfortably 70K 84.9
RTX 4070 Ti Super 16 672 Runs comfortably 70K 77.6
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 67.8
RTX 4070 Super 12 504 Runs comfortably 40K 58.4
RTX 4070 12 504 Runs comfortably 40K 58.4
RTX 5060 Ti 16GB 16 448 Runs comfortably 70K 57.3
Arc B580 12 456 Runs comfortably 40K 44.9
RTX 3060 12GB 12 360 Runs comfortably 40K 43.7
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 35.1
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 34.1
RTX 4060 Ti 16GB 16 288 Runs comfortably 70K 33.6
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 26.2

Architecture

Parameters8.0B
Layers32
Hidden size4096
Attention heads / KV heads32 / 8
Head dimension128
Vocabulary128,256
Trained context128K
KV cache per 1K tokens0 GB
Hugging Facemeta-llama/Llama-3.1-8B-Instruct

The Llama family

Meta's open-weight series, and the default target for most local tooling. Every Llama 3.x model uses grouped-query attention with 8 KV heads, so the cache stays modest even at 70B. Llama 4 moved to mixture-of-experts: Scout and Maverick occupy 109B and 400B of memory but read only 17B per token.

huggingface.co/meta-llama · llama.com · all 7 Llama models

Direct answers