01.AI · Yi-1.5 · 34.4B parameters
Yi-1.5 34B VRAM requirements
Yi-1.5 34B has 60 layers and uses grouped-query attention (8 KV heads). At Q4_K_M the weights come to 19.5 GB, and the best quantisation that fits a 24 GB card is Q4_K_S.
Won't fit
22.2 GB of 21.8 GB · 102%023 GB
Weights
19.5 GB
KV cache
1.9 GB
Runtime overhead
0.9 GB
Over the limit
0.5 GB
Short by 0.5 GB. You can run it with 58 of 60 layers on the RTX 4090 and the rest in system RAM, at roughly 19.7 tok/s — usable for batch work, painful for chat. A smaller quantisation or a shorter context is usually the better trade.
Generation19.7tok/s
Prompt processing1007tok/s
Max context6Ktokens
KV per 1K tokens0GB
Every quantisation of Yi-1.5 34B on a RTX 4090
Highlighted row is the highest quality that still fits at 8K context.
| Quantisation | bpw | Weights | Total | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|---|
| FP16 / BF16 | 16.00 | 64.1 GB | 66.8 GB | Won't fit | — | 0.80 |
| INT8 / W8A8 | 8.50 | 33.8 GB | 36.6 GB | Won't fit | — | 2.27 |
| Q8_0 (GGUF) | 8.50 | 33.8 GB | 36.6 GB | Won't fit | — | 2.27 |
| FP8 (E4M3) | 8.00 | 32.0 GB | 34.8 GB | Won't fit | — | 2.56 |
| Q6_K | 6.56 | 26.3 GB | 29.0 GB | Won't fit | — | 4.27 |
| Q5_K_M | 5.67 | 22.7 GB | 25.5 GB | Won't fit | — | 7.36 |
| Q5_K_S | 5.52 | 22.1 GB | 24.9 GB | Won't fit | — | 8.13 |
| Q4_K_M | 4.85 | 19.5 GB | 22.2 GB | Won't fit | 6K | 19.7 |
| Q4_K_S | 4.58 | 18.4 GB | 21.2 GB | Fits, but tight | 10K | 30.8 |
| Q4_0 | 4.55 | 18.3 GB | 21.1 GB | Fits, but tight | 11K | 31.0 |
| AWQ 4-bit | 4.25 | 18.3 GB | 21.0 GB | Fits, but tight | 11K | 31.1 |
| GPTQ 4-bit | 4.25 | 18.3 GB | 21.0 GB | Fits, but tight | 11K | 31.1 |
| MXFP4 | 4.25 | 18.3 GB | 21.0 GB | Fits, but tight | 11K | 31.1 |
| IQ4_XS | 4.25 | 17.2 GB | 19.9 GB | Fits, but tight | 16K | 33.0 |
| Q3_K_M | 3.91 | 15.8 GB | 18.6 GB | Runs comfortably | 22K | 35.6 |
| IQ3_M | 3.70 | 15.0 GB | 17.8 GB | Runs comfortably | 25K | 37.4 |
| IQ3_XXS | 3.06 | 12.5 GB | 15.3 GB | Runs comfortably | 32K | 44.2 |
| Q2_K | 2.63 | 10.8 GB | 13.6 GB | Runs comfortably | 32K | 50.4 |
| IQ2_XXS | 2.06 | 8.6 GB | 11.4 GB | Runs comfortably | 32K | 61.9 |
| IQ1_M | 1.75 | 7.4 GB | 10.2 GB | Runs comfortably | 32K | 70.6 |
Yi-1.5 34B on each GPU
Q4_K_M weights at 8K context, single card, monitor attached.
| GPU | VRAM | GB/s | Verdict | Max ctx | tok/s |
|---|---|---|---|---|---|
| H100 SXM 80GB | 80 | 3350 | Runs comfortably | 32K | 110 |
| A100 80GB | 80 | 2039 | Runs comfortably | 32K | 61.1 |
| RTX 5090 | 32 | 1792 | Runs comfortably | 32K | 56.8 |
| L40S | 48 | 864 | Runs comfortably | 32K | 25.1 |
| RTX A6000 | 48 | 768 | Runs comfortably | 32K | 23.3 |
| Mac Studio M3 Ultra 256GB | 256 | 819 | Runs comfortably | 32K | 23.0 |
| RTX 4090 | 24 | 1008 | Won't fit | 6K | 19.7 |
| RTX 3090 | 24 | 936 | Won't fit | 6K | 19.6 |
| Radeon RX 7900 XTX | 24 | 960 | Won't fit | 6K | 18.1 |
| Mac Studio M4 Max 128GB | 128 | 546 | Runs comfortably | 32K | 17.1 |
| NVIDIA DGX Spark (GB10) | 128 | 273 | Runs comfortably | 32K | 8.80 |
| Mac Mini M4 Pro 48GB | 48 | 273 | Runs comfortably | 32K | 8.57 |
| Ryzen AI Max+ 395 128GB | 128 | 256 | Runs comfortably | 32K | 6.57 |
| RTX 5080 | 16 | 960 | Won't fit | — | 4.53 |
| RTX 5070 Ti | 16 | 896 | Won't fit | — | 4.50 |
| RTX 5060 Ti 16GB | 16 | 448 | Won't fit | — | 4.13 |
| RTX 4080 Super | 16 | 736 | Won't fit | — | 4.00 |
| RTX 4070 Ti Super | 16 | 672 | Won't fit | — | 3.96 |
| RTX 4060 Ti 16GB | 16 | 288 | Won't fit | — | 3.43 |
| RTX 5070 | 12 | 672 | Won't fit | — | 3.16 |
| RTX 3060 12GB | 12 | 360 | Won't fit | — | 2.85 |
| RTX 4070 Super | 12 | 504 | Won't fit | — | 2.81 |
| RTX 4070 | 12 | 504 | Won't fit | — | 2.81 |
| RTX 3080 10GB | 10 | 760 | Won't fit | — | 2.69 |
| Arc B580 | 12 | 456 | Won't fit | — | 2.36 |
Architecture
| Parameters | 34.4B |
| Layers | 60 |
| Hidden size | 7168 |
| Attention heads / KV heads | 56 / 8 |
| Head dimension | 128 |
| Vocabulary | 64,000 |
| Trained context | 32K |
| KV cache per 1K tokens | 0 GB |
| Hugging Face | 01-ai/Yi-1.5-34B-Chat |
The 01.AI family
Yi-1.5 34B lands in the awkward gap above 30B: 19.5 GB of weights at Q4_K_M, so a 24 GB card takes it only with a modest context.
Direct answers
Yi-1.5 34B on RTX 5090
See the verdictYi-1.5 34B on RTX 4090
See the verdictYi-1.5 34B on RTX 3090
See the verdictYi-1.5 34B on RTX 5080
See the verdictYi-1.5 34B on RTX 5070 Ti
See the verdictYi-1.5 34B on RTX 5070
See the verdictYi-1.5 34B on RTX 4070 Ti Super
See the verdictYi-1.5 34B on RTX 4070 Super
See the verdictYi-1.5 34B on RTX 4070
See the verdict
See the verdictYi-1.5 34B on RTX 4090
See the verdictYi-1.5 34B on RTX 3090
See the verdictYi-1.5 34B on RTX 5080
See the verdictYi-1.5 34B on RTX 5070 Ti
See the verdictYi-1.5 34B on RTX 5070
See the verdictYi-1.5 34B on RTX 4070 Ti Super
See the verdictYi-1.5 34B on RTX 4070 Super
See the verdictYi-1.5 34B on RTX 4070
See the verdict