← Back to all offerings

RTX 5090

32 GB VRAM575 W TDP8 providers119 offerings91 regions

The RTX 5090 is a cloud GPU with 32 GB of memory and 575 W TDP. It is currently available from 8 providers starting at $0.21/hr, with a market median of $0.58/hr across 119 offerings.

$0.21/hr
Cheapest
on GPU Finder
$0.58/hr
Median
across 119 offerings
$26.67/hr
Most expensive
+17.2%
90-day trend

Best deal

Hardware specifications

Architecture
Blackwell
VRAM
32 GB
Memory type
GDDR7
Memory bandwidth
1,792 GB/s
FP32
104 TFLOPS
TDP
575 W

Same across all providers.

Market comparison

$0.58/hr
Our median / hr
$0.97/hr
gpus.io median / hr
119
Our providers
6
gpus.io providers

View source on gpus.io

13-day price history

Offerings

All ×N×1×2×4×8In stock only · Partners only
ProviderCloudCountvCPURAMRegionPer GPU hrSPOTTotal/hr
HardwareHQ
0.0% 30dReferral link
×4220 GBSE$0.47/hr$1.87/hrLaunch ↗
HardwareHQ
0.0% 30dReferral link
×4126 GBCN$0.51/hr$2.03/hrLaunch ↗
Vast.ai
0.0% 30dReferral link
×4TW$0.54/hr$2.14/hrLaunch ↗
HardwareHQ
0.0% 30dReferral link
×4$0.69/hr$2.76/hrLaunch ↗

LLMs that fit

ModelParams16-bit8-bit4-bit
DeepSeek V4 Flash
DeepSeek
284BNot releasedNot releasedBarely fits on 8×
Solar Open2 250B
Upstage
250B 15B activeNeeds >8 GPUsNeeds >8 GPUsBarely fits on 8×
Qwen3 235B-A22B
Alibaba
235B 22B activeNeeds >8 GPUsNeeds >8 GPUsBarely fits on 7×
Laguna-S 2.1
Poolside
118B 8B activeNeeds >8 GPUsBarely fits on 6×4× · 20 GB free
gpt-oss-120b
OpenAI
117B 5.1B activeNot releasedNot releasedBarely fits on 3×
Llama 3.1 70B
Meta
70BBarely fits on 6×Barely fits on 3×2× · 10 GB free
Qwen3.6 35B-A3B
Alibaba
35B 3B active4× · 24.8 GB freeBarely fits on 2×2× · 30.4 GB free
Qwen3 32B
Alibaba
32.8BBarely fits on 3×2× · 12.4 GB freeBarely fits on 1×
Qwen2.5 14B
Alibaba
14B2× · 28.6 GB free1× · 12.7 GB free1× · 20.6 GB free
Llama 3.1 8B
Meta
8B1× · 11.1 GB free1× · 20.6 GB free1× · 25.3 GB free
Mistral 7B
Mistral AI
7B1× · 13.8 GB free1× · 22 GB free1× · 26.1 GB free
Qwen3 4B
Alibaba
4B1× · 20.4 GB free1× · 25.6 GB free1× · 28.3 GB free

Estimates: model weights plus 20% for KV cache and runtime overhead, at short context lengths. Long contexts and large batches need more. Rows marked “tight” hold the weights but not that full margin. A ×N figure is the total VRAM across N of these GPUs and makes no claim about interconnect throughput.

Similar GPUs

Frequently asked questions

How much does the RTX 5090 cost per hour?

The cheapest RTX 5090 offering starts at $0.21/hr, and the market median is $0.58/hr.

Which cloud providers offer the RTX 5090?

We track RTX 5090 offerings from HardwareHQ, Vast.ai.

How much VRAM does the RTX 5090 have?

The RTX 5090 has 32 GB of VRAM.

What LLMs can run on the RTX 5090?

With 32 GB of VRAM, the RTX 5090 can run 12 of our tracked models in 4-bit quantization, including DeepSeek V4 Flash, Solar Open2 250B, Qwen3 235B-A22B. 7 models fit in 16-bit precision.

What's the biggest LLM I can run on multiple RTX 5090s?

With 8× RTX 5090 (256 GB total), the largest tracked model that fits in 16-bit is Qwen3.6 35B-A3B (35B params), which needs 4× GPUs.

Where is the RTX 5090 cheapest?

The cheapest RTX 5090 offering in our catalog is HardwareHQ at $0.47/hr per GPU-hour (×4).