Head to head
RTX 5090 vs RTX 4090
The RTX 5090 holds more — 32 GB against 24 GB — which decides what you can load at all. The RTX 5090 has more bandwidth at 1792 GB/s, which decides how fast tokens come out once a model fits.
| Model (Q4_K_M, 8K) | RTX 5090 | tok/s | RTX 4090 | tok/s |
|---|---|---|---|---|
| Llama 3.1 8B | Runs comfortably | 222 | Runs comfortably | 116 |
| Llama 3.1 70B | Won't fit | 2.65 | Won't fit | 1.60 |
| Llama 3.2 3B | Runs comfortably | 478 | Runs comfortably | 254 |
| Llama 3.2 1B | Runs comfortably | 1235 | Runs comfortably | 669 |
| Llama 4 Scout 109B-A17B | Won't fit | 6.75 | Won't fit | 5.18 |
| Qwen3 8B | Runs comfortably | 214 | Runs comfortably | 112 |
| Qwen3 14B | Runs comfortably | 126 | Runs comfortably | 65.5 |
| Qwen3 32B | Runs comfortably | 59.0 | Fits, but tight | 30.5 |
| Qwen3 4B | Runs comfortably | 380 | Runs comfortably | 202 |
| Qwen3 30B-A3B | Runs comfortably | 159 | Runs comfortably | 118 |
| Qwen3 235B-A22B | Won't fit | 3.83 | Won't fit | 3.23 |
| Qwen2.5-Coder 32B | Runs comfortably | 59.0 | Fits, but tight | 30.5 |
| Qwen2.5 7B | Runs comfortably | 247 | Runs comfortably | 128 |
| Qwen2.5 72B | Won't fit | 2.50 | Won't fit | 1.53 |
| Gemma 3 4B | Runs comfortably | 405 | Runs comfortably | 215 |
| Gemma 3 12B | Runs comfortably | 151 | Runs comfortably | 78.5 |
Specifications
| RTX 5090 | RTX 4090 | |
|---|---|---|
| Memory | 32 GB | 24 GB |
| Bandwidth | 1792 GB/s | 1008 GB/s |
| FP16 compute | 209 TFLOPS | 165 TFLOPS |
| Launch price | $1,999 | $1,599 |
| Architecture | Blackwell | Ada |