Start Building →
Engineering

GB300 NVL72 vs GB200 NVL72: Is Blackwell Ultra Worth the Premium in 2026?

GB300 vs GB200 NVL72 looks like a 1.5x upgrade on paper. Benchmarks say 5% to 45%, and GB300 quietly cuts FP64 by 97%. If you are cost-sensitive, here is when the premium pays and when a B200 at $3.75/hr beats both.

Author photo
packet.ai Team
September 28, 2026

GB300 NVL72 is worth its premium over GB200 NVL72 mainly for long-context, FP4 reasoning inference: Blackwell Ultra adds 50% more dense FP4 compute, about 2x attention throughput and up to 288 GB of HBM3e per GPU, which produced 14% to 45% more DeepSeek-R1 throughput per GPU in MLPerf, not the 1.5x the spec sheet implies.

This post is written for engineering and infrastructure teams already evaluating 72-GPU rack-scale systems. If you are a solo builder, indie team or startup running models on a single node, jump straight to the B200 section below. A single B200 at $3.75/GPU-hr handles Llama 4 405B, Qwen3 235B FP4 and most fine-tune workloads without rack-scale overhead or long-term contracts.

GB300 NVL72 swaps the 72 Blackwell GPUs in GB200 NVL72 for 72 Blackwell Ultra GPUs, raising dense FP4 compute from 720 to 1,080 PFLOPS per rack and HBM3e from 13.4 TB to 20 TB. The NVLink fabric, memory bandwidth and Grace CPU topology stay unchanged. Measured DeepSeek-R1 throughput gains run from 14% to 45% per GPU depending on the serving scenario, not the 1.5x the spec sheet implies.

What this post covers

  • GB300 NVL72 keeps GB200 NVL72's 72-GPU, 130 TB/s NVLink domain and 576 TB/s memory bandwidth, while raising dense FP4 from 720 to 1,080 PFLOPS and HBM3e from 13.4 TB to 20 TB per rack.
  • Peak FP8 and BF16 throughput are identical on both racks. GB300 cuts FP64 from 2,880 to 100 TFLOPS and INT8 from 720 to 24 POPS.
  • Measured DeepSeek-R1 gains per GPU range from 14% (CoreWeave, MLPerf v6.0 server) to 45% (NVIDIA, MLPerf v5.1 offline). InferenceX shows as little as 5% at 231 tok/s/user.
  • GB300 lowers cost per token only when its hourly premium over GB200 is smaller than its throughput gain at your latency target.
  • Vera Rubin NVL72 racks began shipping on September 23, 2026, and NVIDIA reports up to 3.7x GB300 NVL72 throughput in MLPerf Inference v6.1.
  • For solo builders and small teams: a B200 on packet.ai at $3.75/GPU-hr handles most LLM workloads without rack-scale overhead. No minimum commitment, SSH-ready in under 5 minutes.

The GB300 vs GB200 NVL72 choice looks like a clean generational upgrade. It is narrower than that. Both racks use the same NVLink Switch fabric, the same 36 Grace CPUs and the same per-GPU memory bandwidth. NVIDIA changed the GPU's math units and memory capacity and left the interconnect alone.

That matters because the interconnect is what made NVL72 worth considering in the first place. If you are still deciding between rack-scale and node-scale Blackwell, read our GB200 NVL72 vs 8-GPU B200 guide first. This post assumes you already need a 72-GPU NVLink domain and want to know whether Blackwell Ultra earns its higher quote.

GB300 vs GB200 NVL72 specs: what Blackwell Ultra changes

GB300 NVL72 changes the GPU, not the rack. Each Blackwell Ultra GPU gets 15 PFLOPS of dense NVFP4 (up from 10), up to 288 GB of HBM3e (versus 186 GB per GPU in GB200 NVL72), and roughly double the special-function throughput that attention softmax depends on. NVLink bandwidth, memory bandwidth and rack topology do not change, according to NVIDIA's GB300 NVL72 and GB200 NVL72 spec pages.

SpecificationGB200 NVL72GB300 NVL72Change
Configuration72 Blackwell GPUs, 36 Grace CPUs72 Blackwell Ultra GPUs, 36 Grace CPUsSame topology
HBM3e per rack13.4 TB20 TB+49%
HBM3e per GPU186 GBUp to 288 GB+55%
Memory bandwidth per rack576 TB/sUp to 576 TB/sSame
NVLink bandwidth per rack130 TB/s130 TB/sSame
Dense NVFP4720 PFLOPS1,080 PFLOPS+50%
Sparse NVFP41,440 PFLOPS1,440 PFLOPSSame
FP8/FP6 (sparse)720 PFLOPS720 PFLOPSSame
FP16/BF16 (sparse)360 PFLOPS360 PFLOPSSame
INT8 (sparse)720 POPS24 POPS-97%
FP642,880 TFLOPS100 TFLOPS-97%
Attention SFU per GPU5 TeraExponentials/s10.7 TeraExponentials/s+114%
Max GPU power (TGP)Up to 1,200 WUp to 1,400 W+17%

Sources: NVIDIA GB200 NVL72 and GB300 NVL72 spec pages; per-GPU attention and TGP figures from NVIDIA's Blackwell Ultra technical blog (Blackwell vs Blackwell Ultra). Per-GPU GB200 memory derived from NVIDIA's 372 GB per two-GPU superchip.

Two details in that table change how you read vendor charts. First, the 1.5x FP4 gain is dense only; sparse FP4 is 1,440 PFLOPS on both racks. If your serving stack does not use 2:4 structured sparsity, the dense row is the one to compare.

Second, memory bandwidth per GPU stays at 8 TB/s. Token generation at small batch sizes is usually limited by how fast weights and KV cache stream from HBM, so a bigger compute budget does less for low-batch decode than the headline suggests.

Blackwell Ultra raises dense NVFP4 compute by 50% to 15 PFLOPS per GPU, but its FP8 throughput stays at 5 PFLOPS dense, so FP8 serving and BF16 training see no peak-compute uplift from the upgrade.

Scale-out networking does improve. GB300 NVL72 uses ConnectX-8 SuperNICs at 800 Gb/s per GPU, and on the 8-GPU side NVIDIA's HGX spec table lists 1.6 TB/s of networking for HGX B300 versus 0.8 TB/s for HGX B200. That helps multi-rack training and disaggregated serving that crosses rack boundaries. For background on the format and memory changes, see our explainers on NVFP4 precision and HBM3e memory bandwidth.

What GB300 NVL72 gives up: FP64 and INT8 throughput

GB300 NVL72 trades double-precision and INT8 math for FP4 density. NVIDIA lists 100 TFLOPS of FP64 per GB300 rack against 2,880 TFLOPS for GB200, and 24 POPS of INT8 against 720 POPS. The same pattern holds at node scale: HGX B300 lists 10 TFLOPS of FP64 versus 296 TFLOPS for HGX B200.

For most LLM inference teams this costs nothing, because modern serving runs in FP8 or NVFP4. It is a real downgrade for two groups: HPC users running FP64 solvers, and anyone serving checkpoints quantized to INT8 rather than FP8 or FP4.

⚠ Watch out

If your pipeline serves INT8 (W8A8) checkpoints or runs double-precision simulation, GB300 can be slower than the GB200 rack it replaces. Requantize to FP8 or NVFP4 before you move, or keep those jobs on GB200, B200 or Hopper. Our FP8 vs FP16 vs BF16 guide covers the precision trade-offs.

GB300 vs GB200 performance in MLPerf and InferenceX benchmarks

Measured per-GPU gains for GB300 over GB200 on DeepSeek-R1 run from 5% to 45%. The spread depends on who ran the test, the scenario, and the interactivity target. Offline batch throughput shows the largest gains; latency-bound serving shows the smallest.

Source and scenarioGB200 NVL72GB300 NVL72Gain
MLPerf Inference v5.1, NVIDIA, DeepSeek-R1 offline4,024 tok/s/GPU5,842 tok/s/GPU+45%
MLPerf Inference v5.1, NVIDIA, DeepSeek-R1 server2,327 tok/s/GPU2,907 tok/s/GPU+25%
MLPerf Inference v6.0, CoreWeave, DeepSeek-R1 offline7,323 tok/s/GPU9,821 tok/s/GPU+34%
MLPerf Inference v6.0, CoreWeave, DeepSeek-R1 server4,868 tok/s/GPU5,574 tok/s/GPU+14%
InferenceX, DeepSeek-R1 FP4 8K/1K at 93 tok/s/user11,532 tok/s/chip12,534 tok/s/chip+9%
InferenceX, DeepSeek-R1 FP4 8K/1K at 162 tok/s/user4,192 tok/s/chip5,394 tok/s/chip+29%
InferenceX, DeepSeek-R1 FP4 8K/1K at 231 tok/s/user1,012 tok/s/chip1,065 tok/s/chip+5%

Sources: NVIDIA MLPerf v5.1 results, CoreWeave MLPerf v6.0 results, SemiAnalysis InferenceX (interpolated, disaggregated configs, retrieved September 2026). Per-GPU throughput is not a primary MLPerf metric, and figures from different sources use different software and should not be compared across rows.

DeepSeek-R1: measured GB300 gain vs the 1.5x spec

Per-GPU throughput gain, GB300 NVL72 over GB200 NVL72

Dense FP4 spec delta
+50%
MLPerf v5.1 offline (NVIDIA)
+45%
MLPerf v6.0 offline (CoreWeave)
+34%
InferenceX @ 162 tok/s/user
+29%
MLPerf v5.1 server (NVIDIA)
+25%
MLPerf v6.0 server (CoreWeave)
+14%
InferenceX @ 93 tok/s/user
+9%
InferenceX @ 231 tok/s/user
+5%

Sources: NVIDIA, CoreWeave, SemiAnalysis InferenceX. Blue bar is the spec-sheet ceiling; teal bars are measured results.

In MLPerf Inference v5.1, GB300 NVL72 served 5,842 DeepSeek-R1 tokens per second per GPU offline versus 4,024 on GB200 NVL72, a 45% gain, while the latency-bound server scenario improved 25%.

Why the gain shrinks under latency targets

Offline runs can pack large batches into the extra memory and keep FP4 Tensor Cores busy, which is where Blackwell Ultra's added compute and capacity show up. NVIDIA itself describes the larger HBM as a way to raise batch size. Tight per-user latency targets cap batch size, so decode leans on memory bandwidth, and that number did not change.

Power efficiency follows the same curve. InferenceX measured GB300 at about 13% more tokens per megawatt than GB200 at 162 tok/s/user, but GB200 came out slightly ahead at 93 and 231 tok/s/user.

Software moved the numbers as much as silicon

CoreWeave reports it doubled its GB300 NVL72 DeepSeek-R1 server throughput between MLPerf v5.1 and v6.0 on the same hardware, in about six months. That is a bigger jump than the GB200 to GB300 hardware step in any row above. When you compare quotes, compare them on the same serving stack (TensorRT-LLM, SGLang or Dynamo version), or you are measuring software drift. Our prefill vs decode disaggregation explainer covers the serving technique behind most of these results.

Training shows a similar pattern. NVIDIA's MLPerf Training v6.0 summary reports GB300 NVL72 trained up to 1.6x faster than GB200 NVL72 at the same scale. Peak FP8 and BF16 are unchanged, so the gain likely comes from NVFP4 training, larger memory and software; NVIDIA's summary does not break it down.

Is Blackwell Ultra worth the premium? A break-even test

GB300 is worth it when its price premium is smaller than its throughput gain at your latency target. Cost per token scales with hourly price divided by tokens per second. A GB300 quote at 1.3x the GB200 quote needs at least 1.3x the throughput on your workload to break even.

1

Fix your operating point

Pick the tokens-per-second-per-user target your product needs. The gain at 93 tok/s/user and at 162 tok/s/user can differ by more than 3x.

2

Measure both racks on the same stack

Run your model on the same serving framework and version on both systems, or use a public benchmark that matches your model, precision and sequence lengths.

3

Get quotes on matching terms

Ask both providers for the same term length, region and rack count. Rack-scale capacity is quote-based, and term length moves the rate.

4

Compare two ratios

Divide the GB300 quote by the GB200 quote, then divide GB300 throughput by GB200 throughput. GB300 wins on cost per token only if the throughput ratio is larger.

As a hypothetical, assume a GB300 quote lands at 1.3x its GB200 equivalent. NVIDIA's 45% offline gain clears that bar. The 25% MLPerf server gain and CoreWeave's 14% do not, and InferenceX's 29% at 162 tok/s/user is close to break-even. The same rack can be a good buy for a batch pipeline and a poor one for a chat product.

Note

We don't cite third-party GB200 or GB300 rental rates here. Rack pricing is quote-based and changes with term, region and volume, so use the numbers from your own quotes. Memory can also change the math: 288 GB per GPU may let a model fit on fewer GPUs per replica, which raises throughput per rack beyond what per-GPU benchmarks show.

For solo builders and indie teams

The GB300 vs GB200 question is beside the point for your workload. A single B200 on packet.ai covers Llama 4 405B at FP8 (needs around 202 GB), Qwen3 235B at FP4 (around 118 GB), and 70B fine-tuning runs, all within the 192 GB HBM3e of one B200.

packet.ai B200 Dynamic starts at $3.75/GPU-hr, with no platform fee, no egress charge and no long-term contract. That is the lowest published on-demand Blackwell rate currently tracked by the SemiAnalysis GPU pricing comparison.

Which workloads justify GB300 NVL72 over GB200

GB300 pays off where attention, FP4 compute and memory capacity are the constraints at the same time. It adds little where memory bandwidth, FP8 math or small models set the ceiling.

✓ GB300 NVL72 is right for

  • Long-context reasoning inference, where softmax in attention becomes a latency bottleneck
  • Large MoE models served in NVFP4 with wide expert parallelism across the rack
  • High-concurrency serving that needs more KV cache per GPU
  • Offline batch generation and evaluation runs that can use large batches
  • Training runs with a tested NVFP4 recipe

✗ Stay on GB200 or B200 for

  • FP64 simulation and INT8 checkpoints
  • Low-batch FP8 or BF16 decode, limited by the unchanged 8 TB/s per GPU
  • Models whose weights and KV cache fit inside one 8-GPU node
  • Multi-year commitments signed without pricing in Vera Rubin
  • Teams whose serving stack has not been tuned for Blackwell Ultra yet

GB300 vs GB200 vs Vera Rubin: timing the upgrade in 2026

Vera Rubin changes the math for long commitments, not for capacity you need this quarter. Supermicro began shipping Vera Rubin NVL72 racks on September 23, 2026, each with 72 Rubin GPUs, 36 Vera CPUs and 20.7 TB of HBM4. NVIDIA reports up to 3.7x the throughput of GB300 NVL72 in Rubin's first MLPerf Inference v6.1 submission.

First rack shipments are not broad cloud availability. For the next few quarters, GB300 and GB200 are the rack-scale options most teams can actually get.

A 12 to 36 month GB300 commitment signed today will overlap Rubin's cloud ramp. Prefer shorter terms, or ask providers about refresh options. GB200 is not obsolete either. CoreWeave's MLPerf v6.0 GB200 NVL72 result, 7,323 DeepSeek-R1 tokens per second per GPU offline, is higher than the 5,842 that GB300 posted in NVIDIA's v5.1 debut six months earlier. Different submitters and software versions, which is exactly the point: tuning keeps finding headroom on older silicon.

When 8-GPU B200 nodes beat both rack-scale options

For most solo builders and indie teams, the GB300 vs GB200 debate is irrelevant. The real question is: does your workload fit in one B200? One B200 has 192 GB of HBM3e at 8 TB/s, enough for Llama 4 405B at FP8, Qwen3 235B at FP4 and most fine-tunes, at $3.75/GPU-hr with no minimum term on packet.ai. If your tensor or expert parallelism fits inside an 8-GPU NVLink island, an 8-GPU node skips the rack-scale question entirely. Our NVSwitch explainer covers where that 8-GPU NVLink domain ends and why that boundary matters.

NVIDIA

B200 SXM 192GB

$3.75/GPU/hr

on packet.ai · Dynamic · $6.99/hr Dedicated

VRAM

192 GB HBM3e

Memory BW

8 TB/s

NVLink

1.8 TB/s

On packet.ai, a single B200 starts at $3.75/GPU-hr on Dynamic and $6.99/GPU-hr on Dedicated, and a full 8x B200 Dedicated node costs $55.92/hr with hourly billing.

packet.ai B200 Dynamic at $3.75/GPU-hr is the lowest published on-demand Blackwell rate currently tracked by the SemiAnalysis GPU pricing comparison, 36% below the next-lowest published rate. No platform fee, no egress charge, no long-term contract. Dynamic nodes are SSH-ready in under 5 minutes; monthly commits cut up to 20% off the hourly rate, per the packet.ai pricing page. If you outgrow a single node, packet.ai cluster options cover multi-node B200 over InfiniBand on quoted terms.

Frequently asked questions

GB300 NVL72 swaps the 72 Blackwell GPUs in GB200 NVL72 for 72 Blackwell Ultra GPUs. The rack topology, 130 TB/s NVLink domain, 36 Grace CPUs and 576 TB/s memory bandwidth stay the same. What changes is dense FP4 compute (720 to 1,080 PFLOPS per rack), HBM3e capacity (13.4 TB to 20 TB), attention throughput (about 2x), and FP64 and INT8, which drop sharply.
Yes, at the same scale. In MLPerf Training v6.0, NVIDIA reports GB300 NVL72 trained up to 1.6x faster than GB200 NVL72. Peak FP8 and BF16 throughput are identical on both racks, so the gain likely comes from NVFP4 training, the larger 288 GB memory per GPU and software tuning. Teams training in BF16 without FP4 recipes should expect a smaller gap.
GB300 NVL72 rental is quote-based and varies by provider, region, term length and volume, so we don't cite third-party rates. Use a break-even test instead: GB300 is cheaper per token only when its price premium over a GB200 quote is smaller than its throughput gain at your target interactivity. For single-node Blackwell, packet.ai B200 starts at $3.75/GPU-hr on Dynamic.
No. B300 is the Blackwell Ultra GPU itself, usually sold as an 8-GPU HGX B300 node with x86 host CPUs and an 8-GPU NVLink domain. GB300 pairs two Blackwell Ultra GPUs with one Grace CPU per superchip, and GB300 NVL72 links 36 of those superchips into a single 72-GPU NVLink domain. Same GPU silicon, different system and interconnect scope.
Per GPU, yes. NVIDIA lists maximum TGP of up to 1,400 W for Blackwell Ultra versus up to 1,200 W for Blackwell. Tokens per megawatt depends on the operating point: InferenceX DeepSeek-R1 data shows GB300 about 13% more efficient at 162 tok/s/user, but GB200 slightly more efficient at 93 and 231 tok/s/user. Check efficiency at your own latency target.
For capacity you need this quarter, no. Supermicro began shipping Vera Rubin NVL72 racks on September 23, 2026, but first rack shipments are not broad cloud availability. For multi-year commitments, factor Rubin in: NVIDIA reports up to 3.7x GB300 NVL72 throughput in MLPerf Inference v6.1. Prefer shorter terms or refresh clauses on anything signed now.
It can, but slowly. NVIDIA lists 100 TFLOPS of FP64 per GB300 NVL72 rack versus 2,880 TFLOPS for GB200 NVL72, roughly a 97% cut. At node scale, HGX B300 lists 10 TFLOPS FP64 against 296 TFLOPS for HGX B200. Scientific simulation and double-precision solvers should stay on Blackwell GB200, B200 or Hopper-class hardware.

Last reviewed: September 28, 2026. Specs from NVIDIA product pages and technical blogs; benchmarks from MLPerf submissions and SemiAnalysis InferenceX as cited. Before you price a rack, test your workload on a B200 on packet.ai from $3.75/GPU-hr.

Waste less compute.

Same models. Same API. Fraction of the cost. Start free — no credit card required.

Start Building →

More from the blog