packet.ai/Blog/NVIDIA B200 GPU Cloud: Pricing, Specs & Where to Rent in 2026
Guide
NVIDIA B200 GPU Cloud: Pricing, Specs & Where to Rent in 2026
B200 GPU pricing spans $3.75 to $27.04 per hour for identical silicon. Here is what drives the 7x spread across clouds, and how to pick a tier without overpaying.
packet.ai Team
July 7, 2026
NVIDIA B200 cloud pricing on packet.ai starts at $3.75/GPU-hr on Dynamic on-demand, the lowest published rate for B200 as of August 2026. Dedicated single-tenant access starts at $5.90/GPU-hr. Monthly flat-rate plans start at $2,728/month.
Key takeaways
packet.ai B200 SXM on-demand starts at $3.75/hr (Dynamic) and $5.90/hr (Dedicated, 99.99% SLA).
Monthly plans start at $2,728/month for Dynamic and $4,307/month for Dedicated.
The B200 has 192 GB HBM3e, 8 TB/s bandwidth, and up to 20 PFLOPS FP4 compute.
Against six providers surveyed in August 2026, packet.ai has the lowest verified on-demand B200 rate.
A B200 can serve Llama 3.3 70B at FP8 on a single card. H100 requires two cards for the same model.
B200 GPU pricing spans a wide range across cloud providers. The same silicon runs from $3.75/hr to $27.04/hr depending on provider, tenancy model, and contract terms. This guide covers current B200 pricing on packet.ai, how it compares to other providers, and what the B200's specs mean for your workload in practice.
Dynamic instances use scheduler-isolated shared infrastructure. Dedicated instances give you a single-tenant B200 with 99.99% uptime SLA and guaranteed VRAM access. Both tiers provision in under 5 minutes. For multi-GPU cluster deployments, browse available cluster configurations on packet.ai.
B200 Price Comparison Across Cloud Providers
Provider
On-Demand / GPU-hr
Min GPUs
Contract
packet.ai
$3.75
1 GPU
None
Lambda Labs
$6.99
1 GPU
None
RunPod
$5.89
1 GPU
None
CoreWeave
$8.60+
8 GPUs min
Yes
Vast.ai
$5.50
1 GPU
None
AWS (p6)
$14.24
8 GPUs min
No (on-demand)
Prices verified August 2026. packet.ai's $3.75/hr Dynamic rate is 46% below Lambda, 36% below RunPod, and 74% below AWS p6. CoreWeave requires a minimum of 8 GPUs and a sales engagement, effectively gating single-GPU access.
NVIDIA B200 Specifications
192 GB
HBM3e memory
8 TB/s
Memory bandwidth
20 PFLOPS
FP4 compute (peak)
1.8 TB/s
NVLink 5.0
The B200 uses a dual-die Blackwell chiplet design: two TSMC 4NP dies connected by a 10 TB/s internal interconnect. 192 GB of HBM3e at 8 TB/s means a single B200 fits Llama 3.3 70B at FP8 (70 GB) with 122 GB of headroom for KV cache and batching. H100 SXM (80 GB HBM3) requires two cards to run the same model without quantization. For a full architecture breakdown, see the Blackwell architecture guide.
Frequently asked questions
On packet.ai, B200 starts at $3.75/hr Dynamic (scheduler-isolated shared) or $5.90/hr Dedicated (single-tenant, 99.99% SLA). Monthly plans start at $2,728/month Dynamic and $4,307/month Dedicated.
192 GB of HBM3e fits Llama 3.3 70B at FP8 (~70 GB), Llama 3.1 405B at FP4 (~101 GB), or any 13B-34B model at full FP16 with generous KV cache headroom. Only models over 192 GB at inference precision require multi-GPU sharding on a B200.
Dynamic uses packet.ai's scheduler-isolated shared infrastructure, meaning your workload is isolated at the software level but the physical GPU may be shared. Dedicated gives you a physical B200 exclusively, with a 99.99% uptime SLA and no noisy-neighbour risk. Dynamic is better for burst workloads; Dedicated is better for production inference with latency SLAs.
For 70B+ model inference and large-scale training, yes. The B200 delivers up to 2.5x the LLM throughput of H100 at FP8 and adds FP4 precision that H100 does not support. For sub-34B model inference where H100 SXM has sufficient memory, H100 is cheaper per GPU-hour. The break-even depends on model size and utilisation; see the B200 vs H200 ROI guide for the full framework.