Mistral · 7.3B parameters

Mistral 7B v0.3 VRAM requirements

Mistral 7B v0.3 has 32 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 4.1 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.

Runs comfortably

5.9 GB of 21.8 GB · 27%
022 GB
Weights 4.1 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB

Mistral 7B v0.3 at Q4_K_M leaves 15.8 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation128tok/s
Prompt processing4779tok/s
Max context32Ktokens
KV per 1K tokens0GB

Every quantisation of Mistral 7B v0.3 on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 13.5 GB 15.3 GB Runs comfortably 32K 42.8
INT8 / W8A8 8.50 7.1 GB 8.9 GB Runs comfortably 32K 78.2
Q8_0 (GGUF) 8.50 7.1 GB 8.9 GB Runs comfortably 32K 78.2
FP8 (E4M3) 8.00 6.8 GB 8.6 GB Runs comfortably 32K 82.0
Q6_K 6.56 5.5 GB 7.4 GB Runs comfortably 32K 98.2
Q5_K_M 5.67 4.8 GB 6.6 GB Runs comfortably 32K 112
Q5_K_S 5.52 4.7 GB 6.5 GB Runs comfortably 32K 115
Q4_K_M 4.85 4.1 GB 5.9 GB Runs comfortably 32K 128
AWQ 4-bit 4.25 4.0 GB 5.8 GB Runs comfortably 32K 132
GPTQ 4-bit 4.25 4.0 GB 5.8 GB Runs comfortably 32K 132
MXFP4 4.25 4.0 GB 5.8 GB Runs comfortably 32K 132
Q4_K_S 4.58 3.9 GB 5.7 GB Runs comfortably 32K 134
Q4_0 4.55 3.9 GB 5.7 GB Runs comfortably 32K 135
IQ4_XS 4.25 3.6 GB 5.4 GB Runs comfortably 32K 142
Q3_K_M 3.91 3.3 GB 5.2 GB Runs comfortably 32K 152
IQ3_M 3.70 3.2 GB 5.0 GB Runs comfortably 32K 159
IQ3_XXS 3.06 2.7 GB 4.5 GB Runs comfortably 32K 185
Q2_K 2.63 2.3 GB 4.1 GB Runs comfortably 32K 207
IQ2_XXS 2.06 1.8 GB 3.7 GB Runs comfortably 32K 246
IQ1_M 1.75 1.6 GB 3.4 GB Runs comfortably 32K 273

Mistral 7B v0.3 on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 32K 463
A100 80GB 80 2039 Runs comfortably 32K 263
RTX 5090 32 1792 Runs comfortably 32K 245
RTX 5080 16 960 Runs comfortably 32K 134
RTX 4090 24 1008 Runs comfortably 32K 128
RTX 5070 Ti 16 896 Runs comfortably 32K 125
RTX 3090 24 936 Runs comfortably 32K 124
Radeon RX 7900 XTX 24 960 Runs comfortably 32K 116
L40S 48 864 Runs comfortably 32K 110
RTX A6000 48 768 Runs comfortably 32K 102
RTX 3080 10GB 10 760 Runs comfortably 29K 101
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 32K 101
RTX 5070 12 672 Runs comfortably 32K 94.5
RTX 4080 Super 16 736 Runs comfortably 32K 93.9
RTX 4070 Ti Super 16 672 Runs comfortably 32K 85.9
Mac Studio M4 Max 128GB 128 546 Runs comfortably 32K 75.0
RTX 4070 Super 12 504 Runs comfortably 32K 64.7
RTX 4070 12 504 Runs comfortably 32K 64.7
RTX 5060 Ti 16GB 16 448 Runs comfortably 32K 63.4
Arc B580 12 456 Runs comfortably 32K 49.7
RTX 3060 12GB 12 360 Runs comfortably 32K 48.4
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 32K 38.8
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 32K 37.8
RTX 4060 Ti 16GB 16 288 Runs comfortably 32K 37.2
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 32K 29.0

Architecture

Parameters7.3B
Layers32
Hidden size4096
Attention heads / KV heads32 / 8
Head dimension128
Vocabulary32,768
Trained context32K
KV cache per 1K tokens0 GB
Hugging Facemistralai/Mistral-7B-Instruct-v0.3

The Mistral family

Dense models — 7B, NeMo 12B, Small 24B and Large 123B — alongside the two Mixtral mixture-of-experts releases and Codestral for code. Vocabulary size is not consistent across the family: 32k on 7B, Large and Codestral, 131k on NeMo and Small, which changes how much of a small model is embedding table.

huggingface.co/mistralai · mistral.ai · all 7 Mistral models

Direct answers