← Back to all offerings

GH200

96 GB VRAM989 FP16 TFLOPS700 W TDP2 providers4 offerings2 regions

The GH200 is a cloud GPU with 96 GB of memory and 700 W TDP. It is currently available from 2 providers starting at $2.29/hr, with a market median of $2.29/hr across 4 offerings.

$2.29/hr
Cheapest
on HardwareHQ
$2.29/hr
Median
across 4 offerings
$2.29/hr
Most expensive
0.0%
90-day trend

Best deal

Hardware specifications

Architecture
Hopper
VRAM
96 GB
Memory type
HBM3
Memory bandwidth
4,000 GB/s
FP16
989 TFLOPS
FP32
67 TFLOPS
TDP
700 W

Same across all providers.

Market comparison

$2.29/hr
Our median / hr
$2.41/hr
gpus.io median / hr
4
Our providers
1
gpus.io providers

View source on gpus.io

10-day price history

Offerings

All ×N×1×2×4×8In stock only · Partners only
ProviderCloudCountvCPURAMRegionPer GPU hrSPOTTotal/hr
HardwareHQ
0.0% 30dReferral link
×1432 GB$2.29/hr$2.29/hrLaunch ↗
GPU Finder
0.0% 30d
×1dulles-usa-3$2.29/hr$2.29/hrLaunch ↗
GPU Finder
0.0% 30d
×1dulles-usa-3$2.29/hr$2.29/hrLaunch ↗
GPU Finder
0.0% 30d
×1us-default, us-east-3$2.29/hr$2.29/hrLaunch ↗

LLMs that fit

ModelParams16-bit8-bit4-bit
GLM-5.2
Z.ai
753BNeeds >8 GPUsNeeds >8 GPUsBarely fits on 8×
DeepSeek R1
DeepSeek
671B 37B activeNot releasedNeeds >8 GPUs7× · 14.1% spare
DeepSeek V4 Flash
DeepSeek
284BNot releasedNot released3× · 57.6 GB free
Solar Open2 250B
Upstage
250B 15B activeBarely fits on 8×5× · 84 GB free3× · 57.6 GB free
Qwen3 235B-A22B
Alibaba
235B 22B activeBarely fits on 8×Barely fits on 4×3× · 70.8 GB free
Laguna-S 2.1
Poolside
118B 8B activeBarely fits on 4×Barely fits on 2×2× · 84 GB free
gpt-oss-120b
OpenAI
117B 5.1B activeNot releasedNot releasedBarely fits on 1×
Llama 3.1 70B
Meta
70BBarely fits on 2×Barely fits on 1×1× · 42 GB free
Qwen3.6 35B-A3B
Alibaba
35B 3B active2× · 88.8 GB free1× · 39.6 GB free1× · 62.4 GB free
Qwen3 32B
Alibaba
32.8BBarely fits on 1×1× · 44.4 GB free1× · 66 GB free
Qwen2.5 14B
Alibaba
14B1× · 60.6 GB free1× · 76.7 GB free1× · 84.6 GB free
Llama 3.1 8B
Meta
8B1× · 75.1 GB free1× · 84.6 GB free1× · 89.3 GB free

Estimates: model weights plus 20% for KV cache and runtime overhead, at short context lengths. Long contexts and large batches need more. Rows marked “tight” hold the weights but not that full margin. A ×N figure is the total VRAM across N of these GPUs and makes no claim about interconnect throughput.

Similar GPUs

Frequently asked questions

How much does the GH200 cost per hour?

The cheapest GH200 offering starts at $2.29/hr, and the market median is $2.29/hr.

Which cloud providers offer the GH200?

We track GH200 offerings from HardwareHQ, GPU Finder.

How much VRAM does the GH200 have?

The GH200 has 96 GB of VRAM.

What LLMs can run on the GH200?

With 96 GB of VRAM, the GH200 can run 14 of our tracked models in 4-bit quantization, including GLM-5.2, DeepSeek R1, DeepSeek V4 Flash. 10 models fit in 16-bit precision.

What's the biggest LLM I can run on multiple GH200s?

With 8× GH200 (768 GB total), the largest tracked model that fits in 16-bit is Qwen3.6 35B-A3B (35B params), which needs 2× GPUs.

Where is the GH200 cheapest?

The cheapest GH200 offering in our catalog is HardwareHQ at $2.29/hr per GPU-hour (×1).