🚀 B200 starting at $3.75/hr. The best price you'll find. DC in US West → (Access it from button on top after login).

Get Your B200 →
Start Building
packet.ai/GPUs/NVIDIA H200
COMING SOON
NVIDIA Hopper141 GB HBM3eSXM5

NVIDIA H200 GPU

Hopper, maxed out.

The NVIDIA H200 is the peak of the Hopper architecture, 141 GB of HBM3e at 4.8 TB/s, built for long-context LLM inference and large-batch training. Where the H100 runs out of memory, the H200 keeps going. Available on packet.ai from $2.49/GPU-hour.

from $2.49/GPU-hr· LAUNCHING SOON
≈ 56% below market median
No contractsHourly billingSSH in <5 min
NVIDIA H200 GPU
141GB
HBM3e memory
4.8TB/s
Memory bandwidth
990TFLOPS
FP16 Tensor compute
900GB/s
4th-gen NVLink
Specifications

NVIDIA H200 specifications.

SpecificationValueGreat for
GPU architecture
NVIDIA HopperGH100 · 80B transistors
Same proven Hopper compute die as H100, battle-tested in production.
GPU memory
141 GBHBM3e
76% more than H100.
Memory bandwidth
4.8 TB/s
1.4× faster than H100.
FP8 Tensor compute
~1.98 PFLOPS
High-precision FP8 training.
FP16 Tensor compute
~990 TFLOPS
Standard mixed-precision training.
Transformer Engine
4th generation
Up to 4× FP8 speedup vs FP16.
NVLink
900 GB/s4th generation
Scales across NVSwitch-connected nodes.
MIG support
Up to 7 instances
Split a single H200 into isolated GPU instances.
Availability
On-demandUS & EU regions
Contact us for H200 capacity.
Architecture

Built on NVIDIA Hopper, with 1.4× more memory.

The H200 uses the same Hopper GH100 die as the H100 but swaps HBM3 for HBM3e, adding 76% more memory capacity and 1.4× more bandwidth.

GH100 Hopper die

Same 80-billion-transistor die as the H100. 4th-gen Tensor Cores, MIG support, NVLink 4.0, with HBM3e delivering 1.4× the bandwidth.

141 GB HBM3e at 4.8 TB/s

76% more memory than H100. Long-context LLMs, multi-modal pipelines, and large-batch training fit without sharding.

4th-gen Transformer Engine

FP8 training and inference with per-tensor scaling for up to 4× speedup over FP16 on Hopper.

NVLink 4.0 at 900 GB/s

Scale across NVSwitch-connected nodes for large training runs where gradient communication is the bottleneck.

Use Cases

What the H200 is built for.

Long-context LLM inference

141 GB fits 70B models with large KV caches, serve 128K+ context windows without degrading throughput.

  • 128K+ context windows
  • 70B on a single GPU
  • No KV cache starvation

Large-batch training

More memory means larger batches and fewer gradient accumulation steps for faster convergence.

  • Larger batch sizes
  • 1.4× more bandwidth than H100
  • Multi-node NVLink

Multi-modal fine-tuning

Vision-language and audio-language models need more VRAM, the H200 provides the headroom.

  • Vision + language models
  • Full 141 GB VRAM
  • Hourly billing

Scientific HPC

Memory-bound genome sequencing, protein folding, and climate modelling that exhaust H100 capacity.

  • 4.8 TB/s bandwidth
  • FP64 capable
  • MIG for multi-tenant
Pricing

H200 pricing.

H200 is launching soon on packet.ai. Join the waitlist to be notified.

Join waitlist →
FAQ

NVIDIA H200, answered.

What is the NVIDIA H200?

The H200 is the memory-upgraded H100, 141 GB HBM3e at 4.8 TB/s, 76% more memory and 1.4× more bandwidth.

How much does H200 cost?

H200 starts at $2.49/GPU-hour dynamic. See pricing →

Difference between H100 and H200?

Same compute die. H200 replaces HBM3 with HBM3e: 141 GB vs 80 GB, 4.8 TB/s vs 3.35 TB/s.

Does H200 support MIG?

Yes, up to 7 isolated MIG instances per GPU.

How do I get an H200?

Available on-demand from $2.49/hr. Deploy now →

Coming soon

H200: Launching soon. Get notified.

Join the notify list and be first to know when H200 capacity opens on packet.ai.

No commitment · we'll notify you by email

NVIDIA H200from $2.49/GPU-hr
Join waitlist →