AWS GPU pricing for the NVIDIA B200 hit $14.24/hr per GPU on the p6-b200.48xlarge in 2026. packet.ai Blackwell bare metal starts at $3.75/hr, a 74% reduction for the same silicon.
Key takeaways
GPU cloud bills are rising at hyperscalers while compute prices at specialist providers are falling. AWS raised EC2 Capacity Block prices for H200 instances by 15% in January 2026, citing supply and demand. It raised prices again in July 2026, up to 20% across P5, P5e, and P6 families. The announcement went out on a Saturday.
The math is not subtle. An 8xB200 cluster on AWS costs $113.93/hr on-demand. The same 8-GPU B200 Dynamic cluster on packet.ai runs $34,456/mo at 730 hours. That difference compounds fast on any sustained training or inference workload.
This post breaks down the verified per-GPU rates across every major NVIDIA GPU model: H100, H200, and B200. It shows exactly where the gap comes from and what it means for your GPU bill.
AWS GPU instances are organised into families. The P5 family covers H100 GPUs, P5e and P5en cover H200s, and the newer P6 family covers Blackwell B200 and B300 GPUs. On-demand pricing below is for us-east-1.
Sources: Vantage EC2 instances (p6-b200.48xlarge, updated July 21 2026); Economize Cloud (p5.48xlarge); p5en.48xlarge on-demand rate as provided. On-demand pricing; spot and reserved rates differ.
AWS also layers additional costs on top of the headline instance rate: EBS storage at $0.08/GB/month, data egress at $0.09/GB after the first 100GB, and SageMaker surcharges if you run managed training or inference. A team pulling 100GB model checkpoints regularly adds real money to every line item.
packet.ai operates bare metal GPU infrastructure with no hypervisor overhead, no noisy neighbour tax, and no inflated margin for brand premiums. The table below compares verified on-demand rates per GPU per hour.
packet.ai B200 Dynamic on-demand starts at $3.75/hr per GPU, 74% below the AWS p6-b200.48xlarge on-demand rate of $14.24/hr for the same NVIDIA Blackwell silicon.
Per-GPU hourly rates are one number. Monthly cluster costs are where the gap becomes tangible. The calculations below use 730 hours (one month) on an 8-GPU B200 node.
Monthly cost: 8x B200, 730 hrs (on-demand)
8 GPUs x 730 hrs. AWS: $14.24/GPU/hr. packet.ai: 8x B200 Dynamic at $5.90/GPU/hr. Excludes egress, storage, SageMaker.
Running an 8-GPU B200 cluster for one month costs $83,162 on AWS on-demand and $34,456 on packet.ai Dynamic, a $48,706 monthly saving on a single node without giving up a single GPU-hour or switching silicon.
For larger clusters, the savings compound further. A 32-GPU Blackwell build on AWS p6 runs approximately $332,646/mo. The equivalent via GPUaaS.com GPU clusters comes in roughly 30% below that, with pricing available on request for multi-node configurations.
Three structural factors push AWS GPU pricing above specialist cloud providers.
Hyperscaler margin. AWS, Azure, and GCP operate GPU infrastructure as a premium product. Their overhead covers enterprise support, multi-region redundancy, compliance certifications, and per-service billing complexity, all priced into every GPU-hour. Specialist providers like packet.ai run leaner infrastructure with direct data centre partnerships and pass the difference to customers.
Quota friction. New AWS accounts default to 0 P-series vCPUs. A p5.48xlarge consumes 192 vCPUs. Quota increases require written business justification and take 3 to 7 business days to approve. P5e and P5en capacity is harder still. packet.ai provisions B200 clusters in under two minutes without a support ticket.
Price direction. Specialist cloud GPU rates have trended downward through 2025 and into 2026 as supply improved and competition increased. AWS moved in the opposite direction: a 15% H200 Capacity Block increase in January 2026, followed by a 20% increase across P5, P5e, and P6 families in July 2026. The July increase was communicated via documentation update rather than a customer announcement.
Advisory
AWS Capacity Block for ML pricing is a separate product from standard on-demand EC2 pricing. On-demand rates for P5 and P5en instances were reduced by up to 45% in June 2025. The January and July 2026 increases applied specifically to Capacity Blocks, not on-demand pricing. If you compare these, ensure you are comparing the same billing model.
The per-instance hourly rate is the floor, not the ceiling. Several line items routinely inflate the real AWS GPU bill.
Data Egress
$0.09/GB out of AWS after the first 100GB/month. Teams regularly moving model checkpoints or inference outputs pay hundreds to thousands of dollars per month in transfer fees alone.
EBS Storage
P5 local NVMe storage is ephemeral. It disappears when the instance stops. Persistent storage requires EBS at $0.08/GB/month plus snapshot costs. For large dataset or checkpoint workflows, this adds significant monthly spend.
SageMaker Overhead
Teams using SageMaker for managed training or inference pay an additional surcharge on top of EC2 instance costs, typically 10-40% depending on the managed feature used.
Regional Premium
Rates above are for us-east-1. Other AWS regions carry 5-15% surcharges. US West (N. California) P5e Capacity Block rates are $49.75/hr, 25% above the already-elevated standard rate.
packet.ai charges no egress fees. Persistent storage is included with cluster allocations. There are no managed-service surcharges. What you see on the packet.ai GPU cluster page is what you pay.
Lower pricing does not mean lower-spec hardware. packet.ai runs the same NVIDIA B200 SXM silicon as AWS: the same 192GB HBM3e VRAM, the same 8TB/s memory bandwidth, the same NVLink 5.0 interconnect. The hardware spec is identical because it is the same NVIDIA GPU.
The B200 runs 70B-parameter models with substantial KV cache headroom. A single 8-GPU node holds Llama 4 Maverick 400B in 4-bit quantisation without tensor parallelism across nodes. For inference workloads using vLLM or SGLang on B200, FP4 quantisation drops cost-per-token to a fraction of H100 rates despite the higher hourly GPU price. For pre-training runs that fit within a single node, bare metal B200 on packet.ai eliminates both the hyperscaler premium and the managed-service overhead.
B200 clusters on packet.ai are available without quota requests. Deploy a B200 cluster on packet.ai in under two minutes from sign-up.
Not every workload needs a B200. The right GPU depends on model size, VRAM headroom, and how cost-per-token compares across GPU generations. Here is the full picture across the three current NVIDIA data centre GPUs.
For teams running models between 70B and 140B parameters, the H200 is the cost-optimal choice on packet.ai. The 141GB HBM3e loads Llama 4 70B in full BF16 without tensor parallelism across nodes, at $2.49/hr per GPU. For inference workloads on models that fit comfortably in 80GB VRAM, the packet.ai H100 at $2.50/hr remains the most cost-efficient option in the catalogue. Explore H200 SXM on packet.ai for long-context inference workloads.
Last reviewed: July 2026. AWS pricing sourced from Vantage (p6-b200.48xlarge, updated July 21 2026), Economize Cloud (p5.48xlarge), and The Register / DevZero (Capacity Block increases). packet.ai pricing from packet.ai B200 page. GPU prices change frequently. Verify current rates before committing to a cluster.
Same models. Same API. Fraction of the cost. Start free — no credit card required.
Start Building →