Llama · Llama 3.1 · 70.5B parameters
Llama 3.1 70B VRAM requirements
Llama 3.1 70B has 80 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 40.0 GB.
Won't fit
43.4 GB of 21.8 GB · 199%Short by 21.6 GB. You can run it with 36 of 80 layers on the RTX 4090 and the rest in system RAM, at roughly 1.60 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.
Every quantisation of Llama 3.1 70B on a RTX 4090
Highlighted row is the highest quality that still fits at 8K context.
| Quantisation | bpw | Weights | Total | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|---|
| FP16 / BF16 | 16.00 | 131.4 GB | 134.8 GB | Won't fit | — | 0.33 |
| INT8 / W8A8 | 8.50 | 69.3 GB | 72.7 GB | Won't fit | — | 0.72 |
| Q8_0 (GGUF) | 8.50 | 69.3 GB | 72.7 GB | Won't fit | — | 0.72 |
| FP8 (E4M3) | 8.00 | 65.7 GB | 69.1 GB | Won't fit | — | 0.77 |
| Q6_K | 6.56 | 53.9 GB | 57.3 GB | Won't fit | — | 1.01 |
| Q5_K_M | 5.67 | 46.6 GB | 50.0 GB | Won't fit | — | 1.26 |
| Q5_K_S | 5.52 | 45.3 GB | 48.7 GB | Won't fit | — | 1.31 |
| Q4_K_M | 4.85 | 40.0 GB | 43.4 GB | Won't fit | — | 1.60 |
| Q4_K_S | 4.58 | 37.8 GB | 41.2 GB | Won't fit | — | 1.76 |
| AWQ 4-bit | 4.25 | 37.8 GB | 41.2 GB | Won't fit | — | 1.77 |
| GPTQ 4-bit | 4.25 | 37.8 GB | 41.2 GB | Won't fit | — | 1.77 |
| MXFP4 | 4.25 | 37.8 GB | 41.2 GB | Won't fit | — | 1.77 |
| Q4_0 | 4.55 | 37.6 GB | 41.0 GB | Won't fit | — | 1.81 |
| IQ4_XS | 4.25 | 35.2 GB | 38.6 GB | Won't fit | — | 2.02 |
| Q3_K_M | 3.91 | 32.5 GB | 35.9 GB | Won't fit | — | 2.39 |
| IQ3_M | 3.70 | 30.8 GB | 34.2 GB | Won't fit | — | 2.65 |
| IQ3_XXS | 3.06 | 25.7 GB | 29.1 GB | Won't fit | — | 4.26 |
| Q2_K | 2.63 | 22.3 GB | 25.7 GB | Won't fit | — | 6.78 |
| IQ2_XXS | 2.06 | 17.8 GB | 21.2 GB | Fits, but tight | 10K | 31.3 |
| IQ1_M | 1.75 | 15.3 GB | 18.7 GB | Runs comfortably | 18K | 35.9 |
Llama 3.1 70B on each GPU
Q4_K_M weights at 8K context, single card, monitor attached.
| GPU | VRAM | GB/s | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|
| H100 SXM 80GB | 80 | 3350 | Runs comfortably | 107K | 55.4 |
| A100 80GB | 80 | 2039 | Runs comfortably | 107K | 30.5 |
| L40S | 48 | 864 | Fits, but tight | 11K | 12.5 |
| RTX A6000 | 48 | 768 | Fits, but tight | 11K | 11.6 |
| Mac Studio M3 Ultra 256GB | 256 | 819 | Runs comfortably | 128K | 11.4 |
| Mac Studio M4 Max 128GB | 128 | 546 | Runs comfortably | 128K | 8.48 |
| NVIDIA DGX Spark (GB10) | 128 | 273 | Runs comfortably | 128K | 4.37 |
| Ryzen AI Max+ 395 128GB | 128 | 256 | Runs comfortably | 128K | 3.26 |
| RTX 5090 | 32 | 1792 | Won't fit | — | 2.65 |
| Mac Mini M4 Pro 48GB | 48 | 273 | Won't fit | — | 2.63 |
| RTX 3090 | 24 | 936 | Won't fit | — | 1.67 |
| RTX 4090 | 24 | 1008 | Won't fit | — | 1.60 |
| Radeon RX 7900 XTX | 24 | 960 | Won't fit | — | 1.52 |
| RTX 5080 | 16 | 960 | Won't fit | — | 1.35 |
| RTX 5070 Ti | 16 | 896 | Won't fit | — | 1.35 |
| RTX 5060 Ti 16GB | 16 | 448 | Won't fit | — | 1.32 |
| RTX 4080 Super | 16 | 736 | Won't fit | — | 1.22 |
| RTX 4070 Ti Super | 16 | 672 | Won't fit | — | 1.22 |
| RTX 5070 | 12 | 672 | Won't fit | — | 1.21 |
| RTX 4060 Ti 16GB | 16 | 288 | Won't fit | — | 1.17 |
| RTX 3060 12GB | 12 | 360 | Won't fit | — | 1.13 |
| RTX 4070 Super | 12 | 504 | Won't fit | — | 1.09 |
| RTX 4070 | 12 | 504 | Won't fit | — | 1.09 |
| RTX 3080 10GB | 10 | 760 | Won't fit | — | 1.09 |
| Arc B580 | 12 | 456 | Won't fit | — | 0.92 |
Architecture
| Parameters | 70.5B |
| Layers | 80 |
| Hidden size | 8192 |
| Attention heads / KV heads | 64 / 8 |
| Head dimension | 128 |
| Vocabulary | 128,256 |
| Trained context | 128K |
| KV cache per 1K tokens | 0 GB |
| Hugging Face | meta-llama/Llama-3.1-70B-Instruct |
The Llama family
Meta's open-weight series, and the default target for most local tooling. Every Llama 3.x model uses grouped-query attention with 8 KV heads, so the cache stays modest even at 70B. Llama 4 moved to mixture-of-experts: Scout and Maverick occupy 109B and 400B of memory but read only 17B per token.
Direct answers
See the verdictLlama 3.1 70B on RTX 4090
See the verdictLlama 3.1 70B on RTX 3090
See the verdictLlama 3.1 70B on RTX 5080
See the verdictLlama 3.1 70B on RTX 5070 Ti
See the verdictLlama 3.1 70B on RTX 5070
See the verdictLlama 3.1 70B on RTX 4070 Ti Super
See the verdictLlama 3.1 70B on RTX 4070 Super
See the verdictLlama 3.1 70B on RTX 4070
See the verdict