Mistral on NVIDIA Ada

Can I run Mistral Small 24B on an RTX 4070?

Not at Q4_K_M — it needs 15.5 GB against 10.5 GB available. You would need 2 of these cards.

Won't fit

15.5 GB of 10.5 GB · 148%
016 GB
Weights 13.4 GB
KV cache 1.3 GB
Runtime overhead 0.8 GB
Over the limit 5.0 GB

Short by 5.0 GB. You can run it with 25 of 40 layers on the RTX 4070 and the rest in system RAM, at roughly 6.00 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation6.00tok/s
Prompt processing517tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of Mistral Small 24B on a RTX 4070

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 43.9 GB 46.0 GB Won't fit 1.01
INT8 / W8A8 8.50 23.0 GB 25.1 GB Won't fit 2.33
Q8_0 (GGUF) 8.50 23.0 GB 25.1 GB Won't fit 2.33
FP8 (E4M3) 8.00 22.0 GB 24.0 GB Won't fit 2.52
Q6_K 6.56 18.0 GB 20.1 GB Won't fit 3.38
Q5_K_M 5.67 15.6 GB 17.6 GB Won't fit 4.36
Q5_K_S 5.52 15.1 GB 17.2 GB Won't fit 4.67
AWQ 4-bit 4.25 13.5 GB 15.6 GB Won't fit 5.69
GPTQ 4-bit 4.25 13.5 GB 15.6 GB Won't fit 5.69
MXFP4 4.25 13.5 GB 15.6 GB Won't fit 5.69
Q4_K_M 4.85 13.4 GB 15.5 GB Won't fit 6.00
Q4_K_S 4.58 12.7 GB 14.8 GB Won't fit 6.63
Q4_0 4.55 12.6 GB 14.7 GB Won't fit 6.67
IQ4_XS 4.25 11.9 GB 13.9 GB Won't fit 7.88
Q3_K_M 3.91 11.0 GB 13.1 GB Won't fit 9.54
IQ3_M 3.70 10.4 GB 12.5 GB Won't fit 11.5
IQ3_XXS 3.06 8.8 GB 10.9 GB Won't fit 6K 23.8
Q2_K 2.63 7.7 GB 9.8 GB Fits, but tight 13K 36.1
IQ2_XXS 2.06 6.2 GB 8.3 GB Runs comfortably 22K 43.9
IQ1_M 1.75 5.4 GB 7.5 GB Runs comfortably 27K 49.7

Also worth checking