RTX 5090
The RTX 5090 is a cloud GPU with 32 GB of memory and 575 W TDP. It is currently available from 8 providers starting at $0.21/hr, with a market median of $0.58/hr across 119 offerings.
Best deal
Hardware specifications
- Architecture
- Blackwell
- VRAM
- 32 GB
- Memory type
- GDDR7
- Memory bandwidth
- 1,792 GB/s
- FP32
- 104 TFLOPS
- TDP
- 575 W
Same across all providers.
Market comparison
All-time price history
Offerings
| Provider | Cloud | Count | vCPU | RAM | Region | Per GPU hr | SPOT | Total/hr | |
|---|---|---|---|---|---|---|---|---|---|
HardwareHQ 0.0% 30dReferral link | — | ×4 | — | 220 GB | SE | $0.47/hr | — | $1.87/hr | Launch ↗ |
HardwareHQ 0.0% 30dReferral link | — | ×4 | — | 126 GB | CN | $0.51/hr | — | $2.03/hr | Launch ↗ |
Vast.ai 0.0% 30dReferral link | — | ×4 | — | — | TW | $0.54/hr | — | $2.14/hr | Launch ↗ |
HardwareHQ 0.0% 30dReferral link | — | ×4 | — | — | — | $0.69/hr | — | $2.76/hr | Launch ↗ |
LLMs that fit
| Model | Params | 16-bit | 8-bit | 4-bit |
|---|---|---|---|---|
DeepSeek V4 Flash DeepSeek | 284B | Not released | Not released | Barely fits on 8× |
Solar Open2 250B Upstage | 250B 15B active | Needs >8 GPUs | Needs >8 GPUs | Barely fits on 8× |
Qwen3 235B-A22B Alibaba | 235B 22B active | Needs >8 GPUs | Needs >8 GPUs | Barely fits on 7× |
Laguna-S 2.1 Poolside | 118B 8B active | Needs >8 GPUs | Barely fits on 6× | 4× · 20 GB free |
gpt-oss-120b OpenAI | 117B 5.1B active | Not released | Not released | Barely fits on 3× |
Llama 3.1 70B Meta | 70B | Barely fits on 6× | Barely fits on 3× | 2× · 10 GB free |
Qwen3.6 35B-A3B Alibaba | 35B 3B active | 4× · 24.8 GB free | Barely fits on 2× | 2× · 30.4 GB free |
Qwen3 32B Alibaba | 32.8B | Barely fits on 3× | 2× · 12.4 GB free | Barely fits on 1× |
Qwen2.5 14B Alibaba | 14B | 2× · 28.6 GB free | 1× · 12.7 GB free | 1× · 20.6 GB free |
Llama 3.1 8B Meta | 8B | 1× · 11.1 GB free | 1× · 20.6 GB free | 1× · 25.3 GB free |
Mistral 7B Mistral AI | 7B | 1× · 13.8 GB free | 1× · 22 GB free | 1× · 26.1 GB free |
Qwen3 4B Alibaba | 4B | 1× · 20.4 GB free | 1× · 25.6 GB free | 1× · 28.3 GB free |
Estimates: model weights plus 20% for KV cache and runtime overhead, at short context lengths. Long contexts and large batches need more. Rows marked “tight” hold the weights but not that full margin. A ×N figure is the total VRAM across N of these GPUs and makes no claim about interconnect throughput.
Similar GPUs
See more GPUs like this in
Frequently asked questions
How much does the RTX 5090 cost per hour?
The cheapest RTX 5090 offering starts at $0.21/hr, and the market median is $0.58/hr.
Which cloud providers offer the RTX 5090?
We track RTX 5090 offerings from HardwareHQ, Vast.ai.
How much VRAM does the RTX 5090 have?
The RTX 5090 has 32 GB of VRAM.
What LLMs can run on the RTX 5090?
With 32 GB of VRAM, the RTX 5090 can run 12 of our tracked models in 4-bit quantization, including DeepSeek V4 Flash, Solar Open2 250B, Qwen3 235B-A22B. 7 models fit in 16-bit precision.
What's the biggest LLM I can run on multiple RTX 5090s?
With 8× RTX 5090 (256 GB total), the largest tracked model that fits in 16-bit is Qwen3.6 35B-A3B (35B params), which needs 4× GPUs.
Where is the RTX 5090 cheapest?
The cheapest RTX 5090 offering in our catalog is HardwareHQ at $0.47/hr per GPU-hour (×4).