Mistral · Mixtral · 46.7B parameters · 12.9B active

Mixtral 8x7B VRAM requirements

Mixtral 8x7B has 32 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 26.4 GB, and the best quantisation that fits a 24 GB card is IQ3_XXS.

Won't fit

28.2 GB of 21.8 GB · 130%
029 GB
Weights 26.4 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB
Over the limit 6.4 GB

Short by 6.4 GB. You can run it with 24 of 32 layers on the RTX 4090 and the rest in system RAM, at roughly 15.9 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation15.9tok/s
Prompt processing2686tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Mixtral 8x7B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 87.0 GB 88.8 GB Won't fit 1.95
INT8 / W8A8 8.50 46.2 GB 48.0 GB Won't fit 4.62
Q8_0 (GGUF) 8.50 46.2 GB 48.0 GB Won't fit 4.62
FP8 (E4M3) 8.00 43.5 GB 45.3 GB Won't fit 5.13
Q6_K 6.56 35.7 GB 37.5 GB Won't fit 7.25
Q5_K_M 5.67 30.8 GB 32.6 GB Won't fit 10.0
Q5_K_S 5.52 30.0 GB 31.8 GB Won't fit 11.0
Q4_K_M 4.85 26.4 GB 28.2 GB Won't fit 15.9
Q4_K_S 4.58 24.9 GB 26.7 GB Won't fit 18.5
Q4_0 4.55 24.8 GB 26.6 GB Won't fit 18.6
AWQ 4-bit 4.25 23.5 GB 25.3 GB Won't fit 24.7
GPTQ 4-bit 4.25 23.5 GB 25.3 GB Won't fit 24.7
MXFP4 4.25 23.5 GB 25.3 GB Won't fit 24.7
IQ4_XS 4.25 23.1 GB 25.0 GB Won't fit 25.0
Q3_K_M 3.91 21.3 GB 23.1 GB Won't fit 36.4
IQ3_M 3.70 20.2 GB 22.0 GB Won't fit 6K 59.0
IQ3_XXS 3.06 16.7 GB 18.5 GB Runs comfortably 32K 95.0
Q2_K 2.63 14.4 GB 16.2 GB Runs comfortably 32K 108
IQ2_XXS 2.06 11.3 GB 13.1 GB Runs comfortably 32K 131
IQ1_M 1.75 9.6 GB 11.4 GB Runs comfortably 32K 148

Mixtral 8x7B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 32K 220
A100 80GB 80 2039 Runs comfortably 32K 129
RTX 5090 32 1792 Fits, but tight 17K 120
L40S 48 864 Runs comfortably 32K 55.1
RTX A6000 48 768 Runs comfortably 32K 51.2
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 32K 50.5
Mac Studio M4 Max 128GB 128 546 Runs comfortably 32K 37.8
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 32K 19.7
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 32K 19.1
RTX 3090 24 936 Won't fit 16.4
RTX 4090 24 1008 Won't fit 15.9
Radeon RX 7900 XTX 24 960 Won't fit 15.0
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 32K 14.7
RTX 5080 16 960 Won't fit 9.48
RTX 5070 Ti 16 896 Won't fit 9.44
RTX 5060 Ti 16GB 16 448 Won't fit 8.84
RTX 4080 Super 16 736 Won't fit 8.45
RTX 4070 Ti Super 16 672 Won't fit 8.38
RTX 4060 Ti 16GB 16 288 Won't fit 7.50
RTX 5070 12 672 Won't fit 7.46
RTX 3060 12GB 12 360 Won't fit 6.79
RTX 4070 Super 12 504 Won't fit 6.67
RTX 4070 12 504 Won't fit 6.67
RTX 3080 10GB 10 760 Won't fit 6.59
Arc B580 12 456 Won't fit 5.61

Architecture

Parameters46.7B
Active per token12.9B of 8 experts, top-2
Layers32
Hidden size4096
Attention heads / KV heads32 / 8
Head dimension128
Vocabulary32,000
Trained context32K
KV cache per 1K tokens0 GB
Hugging Facemistralai/Mixtral-8x7B-Instruct-v0.1

The Mistral family

Dense models — 7B, NeMo 12B, Small 24B and Large 123B — alongside the two Mixtral mixture-of-experts releases and Codestral for code. Vocabulary size is not consistent across the family: 32k on 7B, Large and Codestral, 131k on NeMo and Small, which changes how much of a small model is embedding table.

huggingface.co/mistralai · mistral.ai · all 7 Mistral models

Direct answers