Cohere · Command R · 35.0B parameters

Command R 35B VRAM requirements

Command R 35B has 40 layers and uses full multi-head attention — no GQA, so the cache is large. At Q4_K_M the weights come to 20.1 GB.

Won't fit

31.0 GB of 21.8 GB · 142%
032 GB
Weights 20.1 GB
KV cache 10.0 GB
Runtime overhead 0.9 GB
Over the limit 9.2 GB

Short by 9.2 GB. You can run it with 21 of 40 layers on the RTX 4090 and the rest in system RAM, at roughly 3.00 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.

Generation3.00tok/s
Prompt processing991tok/s
Max context640tokens
KV per 1K tokens0GB

Every quantisation of Command R 35B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 65.2 GB 76.0 GB Won't fit 0.63
INT8 / W8A8 8.50 33.7 GB 44.6 GB Won't fit 1.38
Q8_0 (GGUF) 8.50 33.7 GB 44.6 GB Won't fit 1.38
FP8 (E4M3) 8.00 32.6 GB 43.5 GB Won't fit 1.46
Q6_K 6.56 26.7 GB 37.6 GB Won't fit 1.93
Q5_K_M 5.67 23.1 GB 34.0 GB Won't fit 2.35
AWQ 4-bit 4.25 23.0 GB 33.9 GB Won't fit 2.36
GPTQ 4-bit 4.25 23.0 GB 33.9 GB Won't fit 2.36
MXFP4 4.25 23.0 GB 33.9 GB Won't fit 2.36
Q5_K_S 5.52 22.5 GB 33.4 GB Won't fit 2.51
Q4_K_M 4.85 20.1 GB 31.0 GB Won't fit 640 3.00
Q4_K_S 4.58 19.1 GB 30.0 GB Won't fit 1K 3.27
Q4_0 4.55 19.0 GB 29.9 GB Won't fit 2K 3.29
IQ4_XS 4.25 17.9 GB 28.8 GB Won't fit 2K 3.81
Q3_K_M 3.91 16.7 GB 27.6 GB Won't fit 3K 4.50
IQ3_M 3.70 15.9 GB 26.8 GB Won't fit 4K 4.96
IQ3_XXS 3.06 13.7 GB 24.5 GB Won't fit 6K 7.47
Q2_K 2.63 12.1 GB 23.0 GB Won't fit 7K 12.4
IQ2_XXS 2.06 10.1 GB 21.0 GB Fits, but tight 9K 39.7
IQ1_M 1.75 9.0 GB 19.8 GB Fits, but tight 10K 42.9

Command R 35B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 43K 91.5
A100 80GB 80 2039 Runs comfortably 43K 50.3
L40S 48 864 Runs comfortably 19K 20.6
RTX A6000 48 768 Runs comfortably 19K 19.1
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 128K 18.8
Mac Studio M4 Max 128GB 128 546 Runs comfortably 60K 14.0
RTX 5090 32 1792 Won't fit 7K 12.7
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 60K 7.19
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 12K 7.00
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 60K 5.36
RTX 3090 24 936 Won't fit 640 3.12
RTX 4090 24 1008 Won't fit 640 3.00
Radeon RX 7900 XTX 24 960 Won't fit 640 2.84
RTX 5080 16 960 Won't fit 1.96
RTX 5070 Ti 16 896 Won't fit 1.95
RTX 5060 Ti 16GB 16 448 Won't fit 1.93
RTX 4080 Super 16 736 Won't fit 1.77
RTX 4070 Ti Super 16 672 Won't fit 1.77
RTX 4060 Ti 16GB 16 288 Won't fit 1.73
RTX 5070 12 672 Won't fit 1.68
RTX 3080 10GB 10 760 Won't fit 1.59
RTX 3060 12GB 12 360 Won't fit 1.59
RTX 4070 Super 12 504 Won't fit 1.53
RTX 4070 12 504 Won't fit 1.53
Arc B580 12 456 Won't fit 1.29

Architecture

Parameters35.0B
Layers40
Hidden size8192
Attention heads / KV heads64 / 64
Head dimension128
Vocabulary256,000
Trained context128K
KV cache per 1K tokens0 GB
Hugging FaceCohereForAI/c4ai-command-r-v01

MHA, not GQA — the KV cache is enormous at long context.

The Cohere family

Command R 35B uses full multi-head attention rather than grouped-query, so its KV cache costs five times as much per token as Command R+ 104B. On this family it is context length, not weights, that usually runs you out of memory.

huggingface.co/CohereForAI · cohere.com · all 2 Cohere models

Direct answers