← Back to all offerings

RTX 5070

12 GB VRAM250 W TDP2 providers2 offerings2 regions

The RTX 5070 is a cloud GPU with 12 GB of memory and 250 W TDP. It is currently available from 2 providers starting at $0.12/hr, with a market median of $0.15/hr across 2 offerings.

$0.12/hr
Cheapest
on HardwareHQ
$0.15/hr
Median
across 2 offerings
$0.18/hr
Most expensive
-16.7%
90-day trend

Hardware specifications

Architecture
Blackwell
VRAM
12 GB
Memory type
GDDR7
Memory bandwidth
672 GB/s
FP32
31 TFLOPS
TDP
250 W

Same across all providers.

Market comparison

$0.15/hr
Our median / hr
$0.29/hr
gpus.io median / hr
2
Our providers
2
gpus.io providers

View source on gpus.io

11-day price history

Offerings

All ×N×1×2×4×8In stock only · Partners only

No offerings found

Try clearing the spot filter or check back later.

LLMs that fit

ModelParams16-bit8-bit4-bit
gpt-oss-120b
OpenAI
117B 5.1B activeNot releasedNot releasedBarely fits on 8×
Llama 3.1 70B
Meta
70BNeeds >8 GPUsBarely fits on 8×Barely fits on 5×
Qwen3.6 35B-A3B
Alibaba
35B 3B activeNeeds >8 GPUsBarely fits on 5×Barely fits on 3×
Qwen3 32B
Alibaba
32.8BBarely fits on 8×5× · 14% spare3× · 6 GB free
Qwen2.5 14B
Alibaba
14BBarely fits on 3×2× · 4.7 GB freeBarely fits on 1×
Llama 3.1 8B
Meta
8BBarely fits on 2×Barely fits on 1×1× · 5.3 GB free
Mistral 7B
Mistral AI
7B2× · 5.8 GB free1× · 2 GB free1× · 6.1 GB free
Qwen3 4B
Alibaba
4BBarely fits on 1×1× · 5.6 GB free1× · 8.3 GB free

Estimates: model weights plus 20% for KV cache and runtime overhead, at short context lengths. Long contexts and large batches need more. Rows marked “tight” hold the weights but not that full margin. A ×N figure is the total VRAM across N of these GPUs and makes no claim about interconnect throughput.

Similar GPUs

Frequently asked questions

How much does the RTX 5070 cost per hour?

The cheapest RTX 5070 offering starts at $0.12/hr, and the market median is $0.15/hr.

Which cloud providers offer the RTX 5070?

No providers currently list the RTX 5070 in our catalog.

How much VRAM does the RTX 5070 have?

The RTX 5070 has 12 GB of VRAM.

What LLMs can run on the RTX 5070?

With 12 GB of VRAM, the RTX 5070 can run 8 of our tracked models in 4-bit quantization, including gpt-oss-120b, Qwen3.6 35B-A3B, Qwen3 32B. 5 models fit in 16-bit precision.

What's the biggest LLM I can run on multiple RTX 5070s?

With 8× RTX 5070 (96 GB total), the largest tracked model that fits in 16-bit is Mistral 7B (7B params), which needs 2× GPUs.