DeepSeek · DeepSeek-R1 · 671B parameters · 37B active

DeepSeek-R1 671B-A37B VRAM requirements

DeepSeek-R1 671B-A37B has 61 layers and uses multi-head latent attention (MLA). At Q4_K_M the weights come to 379.0 GB.

Won't fit

380.4 GB of 21.8 GB · 1748%
0396 GB
Weights 379.0 GB
KV cache 0.5 GB
Runtime overhead 0.9 GB
Over the limit 358.6 GB

Short by 358.6 GB. You can run it with 3 of 61 layers on the RTX 4090 and the rest in system RAM, at roughly 1.88 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation1.88tok/s
Prompt processing936tok/s
Max context0tokens
KV per 1K tokens0GB

Every quantisation of DeepSeek-R1 671B-A37B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 1249.8 GB 1251.2 GB Won't fit 0.55
INT8 / W8A8 8.50 663.6 GB 665.0 GB Won't fit 1.05
Q8_0 (GGUF) 8.50 663.6 GB 665.0 GB Won't fit 1.05
FP8 (E4M3) 8.00 624.9 GB 626.3 GB Won't fit 1.11
Q6_K 6.56 512.4 GB 513.8 GB Won't fit 1.37
Q5_K_M 5.67 442.9 GB 444.3 GB Won't fit 1.59
Q5_K_S 5.52 431.2 GB 432.6 GB Won't fit 1.63
Q4_K_M 4.85 379.0 GB 380.4 GB Won't fit 1.88
Q4_K_S 4.58 358.0 GB 359.4 GB Won't fit 1.98
Q4_0 4.55 355.6 GB 357.0 GB Won't fit 2.00
AWQ 4-bit 4.25 334.5 GB 335.9 GB Won't fit 2.12
GPTQ 4-bit 4.25 334.5 GB 335.9 GB Won't fit 2.12
MXFP4 4.25 334.5 GB 335.9 GB Won't fit 2.12
IQ4_XS 4.25 332.3 GB 333.7 GB Won't fit 2.13
Q3_K_M 3.91 305.8 GB 307.2 GB Won't fit 2.35
IQ3_M 3.70 289.4 GB 290.8 GB Won't fit 2.48
IQ3_XXS 3.06 239.6 GB 241.0 GB Won't fit 3.03
Q2_K 2.63 206.1 GB 207.5 GB Won't fit 3.55
IQ2_XXS 2.06 161.7 GB 163.1 GB Won't fit 4.55
IQ1_M 1.75 137.5 GB 138.9 GB Won't fit 5.49

DeepSeek-R1 671B-A37B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
Mac Studio M3 Ultra 256GB 256 819 Won't fit 3.10
H100 SXM 80GB 80 3350 Won't fit 2.53
Mac Studio M4 Max 128GB 128 546 Won't fit 2.43
NVIDIA DGX Spark (GB10) 128 273 Won't fit 2.40
A100 80GB 80 2039 Won't fit 2.26
RTX 5090 32 1792 Won't fit 2.10
RTX A6000 48 768 Won't fit 2.05
Mac Mini M4 Pro 48GB 48 273 Won't fit 2.04
RTX 5080 16 960 Won't fit 2.03
RTX 5070 Ti 16 896 Won't fit 2.03
RTX 5060 Ti 16GB 16 448 Won't fit 2.03
RTX 5070 12 672 Won't fit 2.00
L40S 48 864 Won't fit 1.97
RTX 3090 24 936 Won't fit 1.96
Ryzen AI Max+ 395 128GB 128 256 Won't fit 1.90
RTX 3080 10GB 10 760 Won't fit 1.90
RTX 3060 12GB 12 360 Won't fit 1.89
RTX 4090 24 1008 Won't fit 1.88
RTX 4080 Super 16 736 Won't fit 1.85
RTX 4070 Ti Super 16 672 Won't fit 1.84
RTX 4060 Ti 16GB 16 288 Won't fit 1.84
RTX 4070 Super 12 504 Won't fit 1.82
RTX 4070 12 504 Won't fit 1.82
Radeon RX 7900 XTX 24 960 Won't fit 1.78
Arc B580 12 456 Won't fit 1.54

Architecture

Parameters671B
Active per token37B of 256 experts, top-8
Layers61
Hidden size7168
Attention heads / KV heads128 / 128
Head dimension128
Vocabulary129,280
Trained context160K
KV cache per 1K tokens0 GB
Hugging Facedeepseek-ai/DeepSeek-R1

The DeepSeek family

V3 and R1 share one 671B mixture-of-experts architecture with 37B active per token, and both use multi-head latent attention — a single compressed KV vector per layer instead of a full set of heads. That gives R1 roughly a fifth of the cache per token of a dense 70B. The R1 distills are ordinary Qwen and Llama models fine-tuned on R1 output, and they fit on one card.

huggingface.co/deepseek-ai · deepseek.com · all 5 DeepSeek models

Direct answers