Gemma · Gemma 3 · 1B parameters

Gemma 3 1B VRAM requirements

Gemma 3 1B has 26 layers and uses grouped-query attention (1 KV heads). At Q4_K_M the weights come to 0.6 GB, and the best quantisation that fits a 24 GB card is FP16 / BF16.

Runs comfortably

1.4 GB of 21.8 GB · 6%
022 GB
Weights 0.6 GB
KV cache 0.0 GB
Runtime overhead 0.8 GB

Gemma 3 1B at Q4_K_M leaves 20.4 GB spare on a RTX 4090. There is room to raise the context length or move up a quantisation level.

Generation850tok/s
Prompt processing34650tok/s
Max context32Ktokens
KV per 1K tokens0GB

Every quantisation of Gemma 3 1B on a RTX 4090

Highlighted row is the highest quality that still fits at 8K context.

QuantisationbpwWeights TotalVerdictMax ctxtok/s
FP16 / BF16 16.00 1.9 GB 2.7 GB Runs comfortably 32K 304
FP8 (E4M3) 8.00 0.9 GB 1.8 GB Runs comfortably 32K 573
INT8 / W8A8 8.50 0.9 GB 1.7 GB Runs comfortably 32K 578
Q8_0 (GGUF) 8.50 0.9 GB 1.7 GB Runs comfortably 32K 578
AWQ 4-bit 4.25 0.9 GB 1.7 GB Runs comfortably 32K 586
GPTQ 4-bit 4.25 0.9 GB 1.7 GB Runs comfortably 32K 586
MXFP4 4.25 0.9 GB 1.7 GB Runs comfortably 32K 586
Q6_K 6.56 0.8 GB 1.6 GB Runs comfortably 32K 681
Q5_K_M 5.67 0.7 GB 1.5 GB Runs comfortably 32K 771
Q5_K_S 5.52 0.6 GB 1.5 GB Runs comfortably 32K 788
Q4_K_M 4.85 0.6 GB 1.4 GB Runs comfortably 32K 850
Q4_K_S 4.58 0.6 GB 1.4 GB Runs comfortably 32K 877
Q4_0 4.55 0.6 GB 1.4 GB Runs comfortably 32K 880
IQ4_XS 4.25 0.5 GB 1.4 GB Runs comfortably 32K 912
Q3_K_M 3.91 0.5 GB 1.3 GB Runs comfortably 32K 952
IQ3_M 3.70 0.5 GB 1.3 GB Runs comfortably 32K 978
IQ3_XXS 3.06 0.4 GB 1.3 GB Runs comfortably 32K 1068
Q2_K 2.63 0.4 GB 1.2 GB Runs comfortably 32K 1138
IQ2_XXS 2.06 0.4 GB 1.2 GB Runs comfortably 32K 1247
IQ1_M 1.75 0.3 GB 1.2 GB Runs comfortably 32K 1315

Gemma 3 1B on each GPU

Q4_K_M weights at 8K context, single card, monitor attached.

GPUVRAMGB/sVerdictMax ctxtok/s
H100 SXM 80GB 80 3350 Runs comfortably 32K 2386
A100 80GB 80 2039 Runs comfortably 32K 1563
RTX 5090 32 1792 Runs comfortably 32K 1477
RTX 5080 16 960 Runs comfortably 32K 886
RTX 4090 24 1008 Runs comfortably 32K 850
RTX 5070 Ti 16 896 Runs comfortably 32K 835
RTX 3090 24 936 Runs comfortably 32K 827
Radeon RX 7900 XTX 24 960 Runs comfortably 32K 778
L40S 48 864 Runs comfortably 32K 742
RTX A6000 48 768 Runs comfortably 32K 694
RTX 3080 10GB 10 760 Runs comfortably 32K 688
Mac Studio M3 Ultra 256GB 256 819 Runs comfortably 32K 685
RTX 5070 12 672 Runs comfortably 32K 647
RTX 4080 Super 16 736 Runs comfortably 32K 643
RTX 4070 Ti Super 16 672 Runs comfortably 32K 593
Mac Studio M4 Max 128GB 128 546 Runs comfortably 32K 523
RTX 4070 Super 12 504 Runs comfortably 32K 455
RTX 4070 12 504 Runs comfortably 32K 455
RTX 5060 Ti 16GB 16 448 Runs comfortably 32K 446
Arc B580 12 456 Runs comfortably 32K 354
RTX 3060 12GB 12 360 Runs comfortably 32K 345
NVIDIA DGX Spark (GB10) 128 273 Runs comfortably 32K 280
Mac Mini M4 Pro 48GB 48 273 Runs comfortably 32K 272
RTX 4060 Ti 16GB 16 288 Runs comfortably 32K 268
Ryzen AI Max+ 395 128GB 128 256 Runs comfortably 32K 211

Architecture

Parameters1B
Layers26
Hidden size1152
Attention heads / KV heads4 / 1
Head dimension256
Vocabulary262,144
Trained context32K
Sliding window512 (every 6th layer is global)
KV cache per 1K tokens0 GB
Hugging Facegoogle/gemma-3-1b-it

The Gemma family

Google's open models. Two things dominate the memory: a 262k vocabulary — on Gemma 3 1B the embedding table is about a third of the file — and sliding-window attention, where only every sixth layer of Gemma 3 sees the full context, and every second layer on Gemma 2. The cache grows far more slowly than the context length suggests.

huggingface.co/google · ai.google.dev/gemma · all 6 Gemma models

Direct answers