Head to head
RTX 3090 vs RTX 4070 Ti Super
The RTX 3090 holds more — 24 GB against 16 GB — which decides what you can load at all. The RTX 3090 has more bandwidth at 936 GB/s, which decides how fast tokens come out once a model fits.
| Model (Q4_K_M, 8K) | RTX 3090 | tok/s | RTX 4070 Ti Super | tok/s |
|---|---|---|---|---|
| Llama 3.1 8B | Runs comfortably | 112 | Runs comfortably | 77.6 |
| Llama 3.1 70B | Won't fit | 1.67 | Won't fit | 1.22 |
| Llama 3.2 3B | Runs comfortably | 246 | Runs comfortably | 172 |
| Llama 3.2 1B | Runs comfortably | 650 | Runs comfortably | 456 |
| Llama 4 Scout 109B-A17B | Won't fit | 5.39 | Won't fit | 4.42 |
| Qwen3 8B | Runs comfortably | 108 | Runs comfortably | 75.1 |
| Qwen3 14B | Runs comfortably | 63.5 | Runs comfortably | 43.9 |
| Qwen3 32B | Fits, but tight | 29.5 | Won't fit | 4.34 |
| Qwen3 4B | Runs comfortably | 196 | Runs comfortably | 137 |
| Qwen3 30B-A3B | Runs comfortably | 117 | Won't fit | 40.2 |
| Qwen3 235B-A22B | Won't fit | 3.36 | Won't fit | 3.05 |
| Qwen2.5-Coder 32B | Fits, but tight | 29.5 | Won't fit | 4.34 |
| Qwen2.5 7B | Runs comfortably | 125 | Runs comfortably | 86.3 |
| Qwen2.5 72B | Won't fit | 1.59 | Won't fit | 1.18 |
| Gemma 3 4B | Runs comfortably | 209 | Runs comfortably | 146 |
| Gemma 3 12B | Runs comfortably | 76.1 | Runs comfortably | 52.7 |
Specifications
| RTX 3090 | RTX 4070 Ti Super | |
|---|---|---|
| Memory | 24 GB | 16 GB |
| Bandwidth | 936 GB/s | 672 GB/s |
| FP16 compute | 71 TFLOPS | 88 TFLOPS |
| Launch price | $1,499 | $799 |
| Architecture | Ampere | Ada |