Microsoft · Phi-4 · 3.8B parameters

Phi-4-mini 3.8B VRAM requirements

Phi-4-mini 3.8B has 32 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 2.2 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.

Runs comfortably

4.0 GB of 21.8 GB · 18%
022 GB
Weights 2.2 GB
KV cache 1.0 GB
Runtime overhead 0.8 GB

Phi-4-mini 3.8B at Q4_K_M leaves 17.7 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation214tok/s
Prompt processing9023tok/s
Max context128Ktokens
KV per 1K tokens0GB

Every quantisation of Phi-4-mini 3.8B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 7.2 GB 9.0 GB Runs comfortably 110K 77.8
INT8 / W8A8 8.50 3.7 GB 5.5 GB Runs comfortably 128K 141
Q8_0 (GGUF) 8.50 3.7 GB 5.5 GB Runs comfortably 128K 141
FP8 (E4M3) 8.00 3.6 GB 5.4 GB Runs comfortably 128K 144
Q6_K 6.56 2.9 GB 4.7 GB Runs comfortably 128K 170
AWQ 4-bit 4.25 2.7 GB 4.5 GB Runs comfortably 128K 180
GPTQ 4-bit 4.25 2.7 GB 4.5 GB Runs comfortably 128K 180
MXFP4 4.25 2.7 GB 4.5 GB Runs comfortably 128K 180
Q5_K_M 5.67 2.5 GB 4.3 GB Runs comfortably 128K 192
Q5_K_S 5.52 2.5 GB 4.3 GB Runs comfortably 128K 196
Q4_K_M 4.85 2.2 GB 4.0 GB Runs comfortably 128K 214
Q4_K_S 4.58 2.1 GB 3.9 GB Runs comfortably 128K 221
Q4_0 4.55 2.1 GB 3.9 GB Runs comfortably 128K 222
IQ4_XS 4.25 2.0 GB 3.8 GB Runs comfortably 128K 232
Q3_K_M 3.91 1.9 GB 3.7 GB Runs comfortably 128K 244
IQ3_M 3.70 1.8 GB 3.6 GB Runs comfortably 128K 252
IQ3_XXS 3.06 1.5 GB 3.3 GB Runs comfortably 128K 280
Q2_K 2.63 1.4 GB 3.2 GB Runs comfortably 128K 303
IQ2_XXS 2.06 1.2 GB 3.0 GB Runs comfortably 128K 339
IQ1_M 1.75 1.1 GB 2.9 GB Runs comfortably 128K 363

Phi-4-mini 3.8B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 128K 741
A100 80GB 80 2039 Runs comfortably 128K 431
RTX 5090 32 1792 Runs comfortably 128K 403
RTX 5080 16 960 Runs comfortably 90K 224
RTX 4090 24 1008 Runs comfortably 128K 214
RTX 5070 Ti 16 896 Runs comfortably 90K 209
RTX 3090 24 936 Runs comfortably 128K 207
Radeon RX 7900 XTX 24 960 Runs comfortably 128K 194
L40S 48 864 Runs comfortably 128K 184
RTX A6000 48 768 Runs comfortably 128K 171
RTX 3080 10GB 10 760 Runs comfortably 45K 170
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 169
RTX 5070 12 672 Runs comfortably 60K 159
RTX 4080 Super 16 736 Runs comfortably 90K 158
RTX 4070 Ti Super 16 672 Runs comfortably 90K 144
Mac Studio M4 Max 128GB 128 546 Runs comfortably 128K 126
RTX 4070 Super 12 504 Runs comfortably 60K 109
RTX 4070 12 504 Runs comfortably 60K 109
RTX 5060 Ti 16GB 16 448 Runs comfortably 90K 107
Arc B580 12 456 Runs comfortably 60K 83.9
RTX 3060 12GB 12 360 Runs comfortably 60K 81.7
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 128K 65.6
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 128K 63.9
RTX 4060 Ti 16GB 16 288 Runs comfortably 90K 62.9
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 128K 49.1

Architecture

Parameters3.8B
Layers32
Hidden size3072
Attention heads / KV heads24 / 8
Head dimension128
Vocabulary200,064
Trained context128K
KV cache per 1K tokens0 GB
Hugging Facemicrosoft/Phi-4-mini-instruct

The Microsoft family

The Phi models are trained on curated and synthetic data to punch above their parameter count. Phi-4 14B keeps a 16K context — short by current standards, but easy on the cache. Phi-4-mini goes to 128K with a 200k vocabulary.

huggingface.co/microsoft · all 2 Microsoft models

Direct answers