Head to head

RTX 5090 vs A100 80GB

The A100 80GB holds more — 80 GB against 32 GB — which decides what you can load at all. The A100 80GB has more bandwidth at 2039 GB/s, which decides how fast tokens come out once a model fits.

Model (Q4_K_M, 8K)RTX 5090tok/sA100 80GBtok/s
Llama 3.1 8B Runs comfortably 222 Runs comfortably 238
Llama 3.1 70B Won't fit 2.65 Runs comfortably 30.5
Llama 3.2 3B Runs comfortably 478 Runs comfortably 512
Llama 3.2 1B Runs comfortably 1235 Runs comfortably 1319
Llama 4 Scout 109B-A17B Won't fit 6.75 Runs comfortably 100
Qwen3 8B Runs comfortably 214 Runs comfortably 230
Qwen3 14B Runs comfortably 126 Runs comfortably 136
Qwen3 32B Runs comfortably 59.0 Runs comfortably 63.5
Qwen3 4B Runs comfortably 380 Runs comfortably 407
Qwen3 30B-A3B Runs comfortably 159 Runs comfortably 163
Qwen3 235B-A22B Won't fit 3.83 Won't fit 6.02
Qwen2.5-Coder 32B Runs comfortably 59.0 Runs comfortably 63.5
Qwen2.5 7B Runs comfortably 247 Runs comfortably 265
Qwen2.5 72B Won't fit 2.50 Runs comfortably 29.6
Gemma 3 4B Runs comfortably 405 Runs comfortably 434
Gemma 3 12B Runs comfortably 151 Runs comfortably 162

Specifications

RTX 5090A100 80GB
Memory32 GB80 GB
Bandwidth1792 GB/s2039 GB/s
FP16 compute209 TFLOPS312 TFLOPS
Launch price$1,999$15,000
ArchitectureBlackwellAmpere