DeepSeek · DeepSeek-Coder-V2 · 15.7B parameters · 2.4B active

DeepSeek-Coder-V2-Lite 16B-A2.4B VRAM requirements

DeepSeek-Coder-V2-Lite 16B-A2.4B has 27 layers and uses multi-head latent attention (MLA). At Q4_K_M the weights come to 8.9 GB, and the best quantisation that fits a 24 GB card is INT8 / W8A8.

Runs comfortably

9.9 GB of 21.8 GB · 46%
022 GB
Weights 8.9 GB
KV cache 0.2 GB
Runtime overhead 0.8 GB

DeepSeek-Coder-V2-Lite 16B-A2.4B at Q4_K_M leaves 11.8 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation219tok/s
Prompt processing14438tok/s
Max context160Ktokens
KV per 1K tokens0GB

Every quantisation of DeepSeek-Coder-V2-Lite 16B-A2.4B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 29.2 GB 30.3 GB Won't fit 23.0
INT8 / W8A8 8.50 15.4 GB 16.5 GB Runs comfortably 160K 154
Q8_0 (GGUF) 8.50 15.4 GB 16.5 GB Runs comfortably 160K 154
FP8 (E4M3) 8.00 14.6 GB 15.6 GB Runs comfortably 160K 160
Q6_K 6.56 12.0 GB 13.0 GB Runs comfortably 160K 183
Q5_K_M 5.67 10.4 GB 11.4 GB Runs comfortably 160K 200
Q5_K_S 5.52 10.1 GB 11.1 GB Runs comfortably 160K 203
Q4_K_M 4.85 8.9 GB 9.9 GB Runs comfortably 160K 219
Q4_K_S 4.58 8.4 GB 9.4 GB Runs comfortably 160K 226
Q4_0 4.55 8.4 GB 9.4 GB Runs comfortably 160K 227
AWQ 4-bit 4.25 8.3 GB 9.4 GB Runs comfortably 160K 227
GPTQ 4-bit 4.25 8.3 GB 9.4 GB Runs comfortably 160K 227
MXFP4 4.25 8.3 GB 9.4 GB Runs comfortably 160K 227
IQ4_XS 4.25 7.8 GB 8.9 GB Runs comfortably 160K 235
Q3_K_M 3.91 7.2 GB 8.2 GB Runs comfortably 160K 245
IQ3_M 3.70 6.9 GB 7.9 GB Runs comfortably 160K 252
IQ3_XXS 3.06 5.7 GB 6.7 GB Runs comfortably 160K 276
Q2_K 2.63 4.9 GB 6.0 GB Runs comfortably 160K 294
IQ2_XXS 2.06 3.9 GB 5.0 GB Runs comfortably 160K 322
IQ1_M 1.75 3.4 GB 4.4 GB Runs comfortably 160K 340

DeepSeek-Coder-V2-Lite 16B-A2.4B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 160K 407
A100 80GB 80 2039 Runs comfortably 160K 326
RTX 5090 32 1792 Runs comfortably 160K 315
RTX 5080 16 960 Runs comfortably 154K 226
RTX 4090 24 1008 Runs comfortably 160K 219
RTX 5070 Ti 16 896 Runs comfortably 154K 216
RTX 3090 24 936 Runs comfortably 160K 215
Radeon RX 7900 XTX 24 960 Runs comfortably 160K 205
L40S 48 864 Runs comfortably 160K 198
RTX A6000 48 768 Runs comfortably 160K 189
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 160K 187
RTX 5070 12 672 Fits, but tight 27K 179
RTX 4080 Super 16 736 Runs comfortably 154K 178
RTX 4070 Ti Super 16 672 Runs comfortably 154K 167
Mac Studio M4 Max 128GB 128 546 Runs comfortably 160K 151
RTX 4070 Super 12 504 Fits, but tight 27K 135
RTX 4070 12 504 Fits, but tight 27K 135
RTX 5060 Ti 16GB 16 448 Runs comfortably 154K 133
Arc B580 12 456 Fits, but tight 27K 109
RTX 3060 12GB 12 360 Fits, but tight 27K 107
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 160K 88.9
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 160K 86.9
RTX 3080 10GB 10 760 Won't fit 86.8
RTX 4060 Ti 16GB 16 288 Runs comfortably 154K 85.6
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 160K 68.9

Architecture

Parameters15.7B
Active per token2.4B of 64 experts, top-6
Layers27
Hidden size2048
Attention heads / KV heads16 / 16
Head dimension128
Vocabulary102,400
Trained context160K
KV cache per 1K tokens0 GB
Hugging Facedeepseek-ai/DeepSeek-Coder-V2-Lite-Instruct

The DeepSeek family

V3 and R1 share one 671B mixture-of-experts architecture with 37B active per token, and both use multi-head latent attention — a single compressed KV vector per layer instead of a full set of heads. That gives R1 roughly a fifth of the cache per token of a dense 70B. The R1 distills are ordinary Qwen and Llama models fine-tuned on R1 output, and they fit on one card.

huggingface.co/deepseek-ai · deepseek.com · all 5 DeepSeek models

Direct answers