Llama · Llama 4 · 109B parameters · 17B active
Llama 4 Scout 109B-A17B VRAM requirements
Llama 4 Scout 109B-A17B has 48 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 61.7 GB.
Won't fit
64.0 GB of 21.8 GB · 294%Short by 42.3 GB. You can run it with 15 of 48 layers on the RTX 4090 and the rest in system RAM, at roughly 5.18 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.
Every quantisation of Llama 4 Scout 109B-A17B on a RTX 4090
Highlighted row is the highest quality that still fits at 8K context.
| Quantisation | bpw | Weights | Total | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|---|
| FP16 / BF16 | 16.00 | 203.0 GB | 205.4 GB | Won't fit | — | 1.28 |
| INT8 / W8A8 | 8.50 | 107.4 GB | 109.7 GB | Won't fit | — | 2.58 |
| Q8_0 (GGUF) | 8.50 | 107.4 GB | 109.7 GB | Won't fit | — | 2.58 |
| FP8 (E4M3) | 8.00 | 101.5 GB | 103.9 GB | Won't fit | — | 2.79 |
| Q6_K | 6.56 | 83.2 GB | 85.6 GB | Won't fit | — | 3.53 |
| Q5_K_M | 5.67 | 71.9 GB | 74.3 GB | Won't fit | — | 4.15 |
| Q5_K_S | 5.52 | 70.0 GB | 72.4 GB | Won't fit | — | 4.37 |
| Q4_K_M | 4.85 | 61.7 GB | 64.0 GB | Won't fit | — | 5.18 |
| Q4_K_S | 4.58 | 58.3 GB | 60.7 GB | Won't fit | — | 5.45 |
| Q4_0 | 4.55 | 58.0 GB | 60.3 GB | Won't fit | — | 5.64 |
| AWQ 4-bit | 4.25 | 56.8 GB | 59.1 GB | Won't fit | — | 5.75 |
| GPTQ 4-bit | 4.25 | 56.8 GB | 59.1 GB | Won't fit | — | 5.75 |
| MXFP4 | 4.25 | 56.8 GB | 59.1 GB | Won't fit | — | 5.75 |
| IQ4_XS | 4.25 | 54.2 GB | 56.6 GB | Won't fit | — | 6.17 |
| Q3_K_M | 3.91 | 50.0 GB | 52.3 GB | Won't fit | — | 6.84 |
| IQ3_M | 3.70 | 47.4 GB | 49.7 GB | Won't fit | — | 7.40 |
| IQ3_XXS | 3.06 | 39.4 GB | 41.8 GB | Won't fit | — | 9.93 |
| Q2_K | 2.63 | 34.1 GB | 36.4 GB | Won't fit | — | 13.1 |
| IQ2_XXS | 2.06 | 27.0 GB | 29.3 GB | Won't fit | — | 22.2 |
| IQ1_M | 1.75 | 23.1 GB | 25.4 GB | Won't fit | — | 37.8 |
Llama 4 Scout 109B-A17B on each GPU
Q4_K_M weights at 8K context, single card, monitor attached.
| GPU | VRAM | GB/s | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|
| H100 SXM 80GB | 80 | 3350 | Runs comfortably | 63K | 177 |
| A100 80GB | 80 | 2039 | Runs comfortably | 63K | 100 |
| Mac Studio M3 Ultra 256GB | 256 | 819 | Runs comfortably | 692K | 38.5 |
| Mac Studio M4 Max 128GB | 128 | 546 | Runs comfortably | 180K | 28.7 |
| NVIDIA DGX Spark (GB10) | 128 | 273 | Runs comfortably | 180K | 14.8 |
| Ryzen AI Max+ 395 128GB | 128 | 256 | Runs comfortably | 180K | 11.1 |
| RTX A6000 | 48 | 768 | Won't fit | — | 9.64 |
| L40S | 48 | 864 | Won't fit | — | 9.40 |
| RTX 5090 | 32 | 1792 | Won't fit | — | 6.75 |
| Mac Mini M4 Pro 48GB | 48 | 273 | Won't fit | — | 6.52 |
| RTX 3090 | 24 | 936 | Won't fit | — | 5.39 |
| RTX 4090 | 24 | 1008 | Won't fit | — | 5.18 |
| Radeon RX 7900 XTX | 24 | 960 | Won't fit | — | 4.91 |
| RTX 5080 | 16 | 960 | Won't fit | — | 4.91 |
| RTX 5070 Ti | 16 | 896 | Won't fit | — | 4.90 |
| RTX 5060 Ti 16GB | 16 | 448 | Won't fit | — | 4.81 |
| RTX 5070 | 12 | 672 | Won't fit | — | 4.57 |
| RTX 4080 Super | 16 | 736 | Won't fit | — | 4.43 |
| RTX 4070 Ti Super | 16 | 672 | Won't fit | — | 4.42 |
| RTX 4060 Ti 16GB | 16 | 288 | Won't fit | — | 4.28 |
| RTX 3060 12GB | 12 | 360 | Won't fit | — | 4.27 |
| RTX 3080 10GB | 10 | 760 | Won't fit | — | 4.16 |
| RTX 4070 Super | 12 | 504 | Won't fit | — | 4.12 |
| RTX 4070 | 12 | 504 | Won't fit | — | 4.12 |
| Arc B580 | 12 | 456 | Won't fit | — | 3.48 |
Architecture
| Parameters | 109B |
| Active per token | 17B of 16 experts, top-1 |
| Layers | 48 |
| Hidden size | 5120 |
| Attention heads / KV heads | 40 / 8 |
| Head dimension | 128 |
| Vocabulary | 202,048 |
| Trained context | 10240K |
| KV cache per 1K tokens | 0 GB |
| Hugging Face | meta-llama/Llama-4-Scout-17B-16E-Instruct |
The Llama family
Meta's open-weight series, and the default target for most local tooling. Every Llama 3.x model uses grouped-query attention with 8 KV heads, so the cache stays modest even at 70B. Llama 4 moved to mixture-of-experts: Scout and Maverick occupy 109B and 400B of memory but read only 17B per token.
Direct answers
See the verdictLlama 4 Scout 109B-A17B on RTX 4090
See the verdictLlama 4 Scout 109B-A17B on RTX 3090
See the verdictLlama 4 Scout 109B-A17B on RTX 5080
See the verdictLlama 4 Scout 109B-A17B on RTX 5070 Ti
See the verdictLlama 4 Scout 109B-A17B on RTX 5070
See the verdictLlama 4 Scout 109B-A17B on RTX 4070 Ti Super
See the verdictLlama 4 Scout 109B-A17B on RTX 4070 Super
See the verdictLlama 4 Scout 109B-A17B on RTX 4070
See the verdict