Mistral · 23.6B parameters
Mistral Small 24B VRAM requirements
Mistral Small 24B has 40 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 13.4 GB, and the best quantisation that fits a 24 GB card is Q6_K.
Runs comfortably
15.5 GB of 21.8 GB · 71%Mistral Small 24B at Q4_K_M leaves 6.3 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.
Every quantisation of Mistral Small 24B on a RTX 4090
Highlighted row is the highest quality that still fits at 8K context.
| Quantisation | bpw | Weights | Total | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|---|
| FP16 / BF16 | 16.00 | 43.9 GB | 46.0 GB | Won't fit | — | 1.43 |
| INT8 / W8A8 | 8.50 | 23.0 GB | 25.1 GB | Won't fit | — | 7.93 |
| Q8_0 (GGUF) | 8.50 | 23.0 GB | 25.1 GB | Won't fit | — | 7.93 |
| FP8 (E4M3) | 8.00 | 22.0 GB | 24.0 GB | Won't fit | — | 9.38 |
| Q6_K | 6.56 | 18.0 GB | 20.1 GB | Fits, but tight | 19K | 32.2 |
| Q5_K_M | 5.67 | 15.6 GB | 17.6 GB | Runs comfortably | 32K | 37.0 |
| Q5_K_S | 5.52 | 15.1 GB | 17.2 GB | Runs comfortably | 32K | 38.0 |
| AWQ 4-bit | 4.25 | 13.5 GB | 15.6 GB | Runs comfortably | 32K | 42.4 |
| GPTQ 4-bit | 4.25 | 13.5 GB | 15.6 GB | Runs comfortably | 32K | 42.4 |
| MXFP4 | 4.25 | 13.5 GB | 15.6 GB | Runs comfortably | 32K | 42.4 |
| Q4_K_M | 4.85 | 13.4 GB | 15.5 GB | Runs comfortably | 32K | 42.6 |
| Q4_K_S | 4.58 | 12.7 GB | 14.8 GB | Runs comfortably | 32K | 44.8 |
| Q4_0 | 4.55 | 12.6 GB | 14.7 GB | Runs comfortably | 32K | 45.1 |
| IQ4_XS | 4.25 | 11.9 GB | 13.9 GB | Runs comfortably | 32K | 47.9 |
| Q3_K_M | 3.91 | 11.0 GB | 13.1 GB | Runs comfortably | 32K | 51.5 |
| IQ3_M | 3.70 | 10.4 GB | 12.5 GB | Runs comfortably | 32K | 54.0 |
| IQ3_XXS | 3.06 | 8.8 GB | 10.9 GB | Runs comfortably | 32K | 63.3 |
| Q2_K | 2.63 | 7.7 GB | 9.8 GB | Runs comfortably | 32K | 71.7 |
| IQ2_XXS | 2.06 | 6.2 GB | 8.3 GB | Runs comfortably | 32K | 86.8 |
| IQ1_M | 1.75 | 5.4 GB | 7.5 GB | Runs comfortably | 32K | 98.2 |
Mistral Small 24B on each GPU
Q4_K_M weights at 8K context, single card, monitor attached.
| GPU | VRAM | GB/s | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|
| H100 SXM 80GB | 80 | 3350 | Runs comfortably | 32K | 161 |
| A100 80GB | 80 | 2039 | Runs comfortably | 32K | 89.0 |
| RTX 5090 | 32 | 1792 | Runs comfortably | 32K | 82.7 |
| RTX 4090 | 24 | 1008 | Runs comfortably | 32K | 42.6 |
| RTX 3090 | 24 | 936 | Runs comfortably | 32K | 41.3 |
| Radeon RX 7900 XTX | 24 | 960 | Runs comfortably | 32K | 38.6 |
| L40S | 48 | 864 | Runs comfortably | 32K | 36.6 |
| RTX A6000 | 48 | 768 | Runs comfortably | 32K | 34.0 |
| Mac Studio M3 Ultra 256GB | 256 | 819 | Runs comfortably | 32K | 33.4 |
| Mac Studio M4 Max 128GB | 128 | 546 | Runs comfortably | 32K | 24.9 |
| RTX 5080 | 16 | 960 | Won't fit | — | 18.7 |
| RTX 5070 Ti | 16 | 896 | Won't fit | — | 18.2 |
| RTX 4080 Super | 16 | 736 | Won't fit | — | 15.3 |
| RTX 4070 Ti Super | 16 | 672 | Won't fit | — | 14.7 |
| RTX 5060 Ti 16GB | 16 | 448 | Won't fit | — | 13.1 |
| NVIDIA DGX Spark (GB10) | 128 | 273 | Runs comfortably | 32K | 12.8 |
| Mac Mini M4 Pro 48GB | 48 | 273 | Runs comfortably | 32K | 12.5 |
| Ryzen AI Max+ 395 128GB | 128 | 256 | Runs comfortably | 32K | 9.56 |
| RTX 4060 Ti 16GB | 16 | 288 | Won't fit | — | 9.09 |
| RTX 5070 | 12 | 672 | Won't fit | — | 6.91 |
| RTX 4070 Super | 12 | 504 | Won't fit | — | 6.00 |
| RTX 4070 | 12 | 504 | Won't fit | — | 6.00 |
| RTX 3060 12GB | 12 | 360 | Won't fit | — | 5.85 |
| RTX 3080 10GB | 10 | 760 | Won't fit | — | 5.03 |
| Arc B580 | 12 | 456 | Won't fit | — | 4.99 |
Architecture
| Parameters | 23.6B |
| Layers | 40 |
| Hidden size | 5120 |
| Attention heads / KV heads | 32 / 8 |
| Head dimension | 128 |
| Vocabulary | 131,072 |
| Trained context | 32K |
| KV cache per 1K tokens | 0 GB |
| Hugging Face | mistralai/Mistral-Small-24B-Instruct-2501 |
The Mistral family
Dense models — 7B, NeMo 12B, Small 24B and Large 123B — alongside the two Mixtral mixture-of-experts releases and Codestral for code. Vocabulary size is not consistent across the family: 32k on 7B, Large and Codestral, 131k on NeMo and Small, which changes how much of a small model is embedding table.
huggingface.co/mistralai · mistral.ai · all 7 Mistral models
Direct answers
See the verdictMistral Small 24B on RTX 4090
See the verdictMistral Small 24B on RTX 3090
See the verdictMistral Small 24B on RTX 5080
See the verdictMistral Small 24B on RTX 5070 Ti
See the verdictMistral Small 24B on RTX 5070
See the verdictMistral Small 24B on RTX 4070 Ti Super
See the verdictMistral Small 24B on RTX 4070 Super
See the verdictMistral Small 24B on RTX 4070
See the verdict