The NVIDIA H200 is the peak of the Hopper architecture, 141 GB of HBM3e at 4.8 TB/s, built for long-context LLM inference and large-batch training. Where the H100 runs out of memory, the H200 keeps going. Available on packet.ai from $2.49/GPU-hour.

The H200 uses the same Hopper GH100 die as the H100 but swaps HBM3 for HBM3e, adding 76% more memory capacity and 1.4× more bandwidth.
Same 80-billion-transistor die as the H100. 4th-gen Tensor Cores, MIG support, NVLink 4.0, with HBM3e delivering 1.4× the bandwidth.
76% more memory than H100. Long-context LLMs, multi-modal pipelines, and large-batch training fit without sharding.
FP8 training and inference with per-tensor scaling for up to 4× speedup over FP16 on Hopper.
Scale across NVSwitch-connected nodes for large training runs where gradient communication is the bottleneck.
141 GB fits 70B models with large KV caches, serve 128K+ context windows without degrading throughput.
More memory means larger batches and fewer gradient accumulation steps for faster convergence.
Vision-language and audio-language models need more VRAM, the H200 provides the headroom.
Memory-bound genome sequencing, protein folding, and climate modelling that exhaust H100 capacity.
H200 is launching soon on packet.ai. Join the waitlist to be notified.
Join waitlist →The H200 is the memory-upgraded H100, 141 GB HBM3e at 4.8 TB/s, 76% more memory and 1.4× more bandwidth.
H200 starts at $2.49/GPU-hour dynamic. See pricing →
Same compute die. H200 replaces HBM3 with HBM3e: 141 GB vs 80 GB, 4.8 TB/s vs 3.35 TB/s.
Yes, up to 7 isolated MIG instances per GPU.
Available on-demand from $2.49/hr. Deploy now →
Join the notify list and be first to know when H200 capacity opens on packet.ai.
No commitment · we'll notify you by email