Head to head
RTX 5080 vs NVIDIA DGX Spark (GB10)
The NVIDIA DGX Spark (GB10) holds more — 128 GB against 16 GB — which decides what you can load at all. The RTX 5080 has more bandwidth at 960 GB/s, which decides how fast tokens come out once a model fits.
| Model (Q4_K_M, 8K) | RTX 5080 | tok/s | NVIDIA DGX Spark (GB10) | tok/s |
|---|---|---|---|---|
| Llama 3.1 8B | Runs comfortably | 121 | Runs comfortably | 35.1 |
| Llama 3.1 70B | Won't fit | 1.35 | Runs comfortably | 4.37 |
| Llama 3.2 3B | Runs comfortably | 266 | Runs comfortably | 78.1 |
| Llama 3.2 1B | Runs comfortably | 700 | Runs comfortably | 209 |
| Llama 4 Scout 109B-A17B | Won't fit | 4.91 | Runs comfortably | 14.8 |
| Qwen3 8B | Runs comfortably | 117 | Runs comfortably | 33.9 |
| Qwen3 14B | Runs comfortably | 68.7 | Runs comfortably | 19.8 |
| Qwen3 32B | Won't fit | 4.98 | Runs comfortably | 9.17 |
| Qwen3 4B | Runs comfortably | 212 | Runs comfortably | 62.2 |
| Qwen3 30B-A3B | Won't fit | 46.2 | Runs comfortably | 53.5 |
| Qwen3 235B-A22B | Won't fit | 3.36 | Won't fit | 6.20 |
| Qwen2.5-Coder 32B | Won't fit | 4.98 | Runs comfortably | 9.17 |
| Qwen2.5 7B | Runs comfortably | 135 | Runs comfortably | 38.9 |
| Qwen2.5 72B | Won't fit | 1.31 | Runs comfortably | 4.24 |
| Gemma 3 4B | Runs comfortably | 226 | Runs comfortably | 66.4 |
| Gemma 3 12B | Runs comfortably | 82.3 | Runs comfortably | 23.8 |
Specifications
| RTX 5080 | NVIDIA DGX Spark (GB10) | |
|---|---|---|
| Memory | 16 GB | 128 GB |
| Bandwidth | 960 GB/s | 273 GB/s |
| FP16 compute | 112 TFLOPS | 125 TFLOPS |
| Launch price | $999 | $3,999 |
| Architecture | Blackwell | Blackwell |