
GB300 vs GB200 NVL72 looks like a 1.5x upgrade on paper. Benchmarks say 5% to 45%, and GB300 quietly cuts FP64 by 97%. If you are cost-sensitive, here is when the premium pays and when a B200 at $3.75/hr beats both.
GB300 NVL72 is worth its premium over GB200 NVL72 mainly for long-context, FP4 reasoning inference: Blackwell Ultra adds 50% more dense FP4 compute, about 2x attention throughput and up to 288 GB of HBM3e per GPU, which produced 14% to 45% more DeepSeek-R1 throughput per GPU in MLPerf, not the 1.5x the spec sheet implies.
This post is written for engineering and infrastructure teams already evaluating 72-GPU rack-scale systems. If you are a solo builder, indie team or startup running models on a single node, jump straight to the B200 section below. A single B200 at $3.75/GPU-hr handles Llama 4 405B, Qwen3 235B FP4 and most fine-tune workloads without rack-scale overhead or long-term contracts.
GB300 NVL72 swaps the 72 Blackwell GPUs in GB200 NVL72 for 72 Blackwell Ultra GPUs, raising dense FP4 compute from 720 to 1,080 PFLOPS per rack and HBM3e from 13.4 TB to 20 TB. The NVLink fabric, memory bandwidth and Grace CPU topology stay unchanged. Measured DeepSeek-R1 throughput gains run from 14% to 45% per GPU depending on the serving scenario, not the 1.5x the spec sheet implies.
What this post covers
The GB300 vs GB200 NVL72 choice looks like a clean generational upgrade. It is narrower than that. Both racks use the same NVLink Switch fabric, the same 36 Grace CPUs and the same per-GPU memory bandwidth. NVIDIA changed the GPU's math units and memory capacity and left the interconnect alone.
That matters because the interconnect is what made NVL72 worth considering in the first place. If you are still deciding between rack-scale and node-scale Blackwell, read our GB200 NVL72 vs 8-GPU B200 guide first. This post assumes you already need a 72-GPU NVLink domain and want to know whether Blackwell Ultra earns its higher quote.
GB300 NVL72 changes the GPU, not the rack. Each Blackwell Ultra GPU gets 15 PFLOPS of dense NVFP4 (up from 10), up to 288 GB of HBM3e (versus 186 GB per GPU in GB200 NVL72), and roughly double the special-function throughput that attention softmax depends on. NVLink bandwidth, memory bandwidth and rack topology do not change, according to NVIDIA's GB300 NVL72 and GB200 NVL72 spec pages.
Sources: NVIDIA GB200 NVL72 and GB300 NVL72 spec pages; per-GPU attention and TGP figures from NVIDIA's Blackwell Ultra technical blog (Blackwell vs Blackwell Ultra). Per-GPU GB200 memory derived from NVIDIA's 372 GB per two-GPU superchip.
Two details in that table change how you read vendor charts. First, the 1.5x FP4 gain is dense only; sparse FP4 is 1,440 PFLOPS on both racks. If your serving stack does not use 2:4 structured sparsity, the dense row is the one to compare.
Second, memory bandwidth per GPU stays at 8 TB/s. Token generation at small batch sizes is usually limited by how fast weights and KV cache stream from HBM, so a bigger compute budget does less for low-batch decode than the headline suggests.
Blackwell Ultra raises dense NVFP4 compute by 50% to 15 PFLOPS per GPU, but its FP8 throughput stays at 5 PFLOPS dense, so FP8 serving and BF16 training see no peak-compute uplift from the upgrade.
Scale-out networking does improve. GB300 NVL72 uses ConnectX-8 SuperNICs at 800 Gb/s per GPU, and on the 8-GPU side NVIDIA's HGX spec table lists 1.6 TB/s of networking for HGX B300 versus 0.8 TB/s for HGX B200. That helps multi-rack training and disaggregated serving that crosses rack boundaries. For background on the format and memory changes, see our explainers on NVFP4 precision and HBM3e memory bandwidth.
GB300 NVL72 trades double-precision and INT8 math for FP4 density. NVIDIA lists 100 TFLOPS of FP64 per GB300 rack against 2,880 TFLOPS for GB200, and 24 POPS of INT8 against 720 POPS. The same pattern holds at node scale: HGX B300 lists 10 TFLOPS of FP64 versus 296 TFLOPS for HGX B200.
For most LLM inference teams this costs nothing, because modern serving runs in FP8 or NVFP4. It is a real downgrade for two groups: HPC users running FP64 solvers, and anyone serving checkpoints quantized to INT8 rather than FP8 or FP4.
⚠ Watch out
If your pipeline serves INT8 (W8A8) checkpoints or runs double-precision simulation, GB300 can be slower than the GB200 rack it replaces. Requantize to FP8 or NVFP4 before you move, or keep those jobs on GB200, B200 or Hopper. Our FP8 vs FP16 vs BF16 guide covers the precision trade-offs.
Measured per-GPU gains for GB300 over GB200 on DeepSeek-R1 run from 5% to 45%. The spread depends on who ran the test, the scenario, and the interactivity target. Offline batch throughput shows the largest gains; latency-bound serving shows the smallest.
Sources: NVIDIA MLPerf v5.1 results, CoreWeave MLPerf v6.0 results, SemiAnalysis InferenceX (interpolated, disaggregated configs, retrieved September 2026). Per-GPU throughput is not a primary MLPerf metric, and figures from different sources use different software and should not be compared across rows.
DeepSeek-R1: measured GB300 gain vs the 1.5x spec
Per-GPU throughput gain, GB300 NVL72 over GB200 NVL72
Sources: NVIDIA, CoreWeave, SemiAnalysis InferenceX. Blue bar is the spec-sheet ceiling; teal bars are measured results.
In MLPerf Inference v5.1, GB300 NVL72 served 5,842 DeepSeek-R1 tokens per second per GPU offline versus 4,024 on GB200 NVL72, a 45% gain, while the latency-bound server scenario improved 25%.
Offline runs can pack large batches into the extra memory and keep FP4 Tensor Cores busy, which is where Blackwell Ultra's added compute and capacity show up. NVIDIA itself describes the larger HBM as a way to raise batch size. Tight per-user latency targets cap batch size, so decode leans on memory bandwidth, and that number did not change.
Power efficiency follows the same curve. InferenceX measured GB300 at about 13% more tokens per megawatt than GB200 at 162 tok/s/user, but GB200 came out slightly ahead at 93 and 231 tok/s/user.
CoreWeave reports it doubled its GB300 NVL72 DeepSeek-R1 server throughput between MLPerf v5.1 and v6.0 on the same hardware, in about six months. That is a bigger jump than the GB200 to GB300 hardware step in any row above. When you compare quotes, compare them on the same serving stack (TensorRT-LLM, SGLang or Dynamo version), or you are measuring software drift. Our prefill vs decode disaggregation explainer covers the serving technique behind most of these results.
Training shows a similar pattern. NVIDIA's MLPerf Training v6.0 summary reports GB300 NVL72 trained up to 1.6x faster than GB200 NVL72 at the same scale. Peak FP8 and BF16 are unchanged, so the gain likely comes from NVFP4 training, larger memory and software; NVIDIA's summary does not break it down.
GB300 is worth it when its price premium is smaller than its throughput gain at your latency target. Cost per token scales with hourly price divided by tokens per second. A GB300 quote at 1.3x the GB200 quote needs at least 1.3x the throughput on your workload to break even.
Fix your operating point
Pick the tokens-per-second-per-user target your product needs. The gain at 93 tok/s/user and at 162 tok/s/user can differ by more than 3x.
Measure both racks on the same stack
Run your model on the same serving framework and version on both systems, or use a public benchmark that matches your model, precision and sequence lengths.
Get quotes on matching terms
Ask both providers for the same term length, region and rack count. Rack-scale capacity is quote-based, and term length moves the rate.
Compare two ratios
Divide the GB300 quote by the GB200 quote, then divide GB300 throughput by GB200 throughput. GB300 wins on cost per token only if the throughput ratio is larger.
As a hypothetical, assume a GB300 quote lands at 1.3x its GB200 equivalent. NVIDIA's 45% offline gain clears that bar. The 25% MLPerf server gain and CoreWeave's 14% do not, and InferenceX's 29% at 162 tok/s/user is close to break-even. The same rack can be a good buy for a batch pipeline and a poor one for a chat product.
Note
We don't cite third-party GB200 or GB300 rental rates here. Rack pricing is quote-based and changes with term, region and volume, so use the numbers from your own quotes. Memory can also change the math: 288 GB per GPU may let a model fit on fewer GPUs per replica, which raises throughput per rack beyond what per-GPU benchmarks show.
For solo builders and indie teams
The GB300 vs GB200 question is beside the point for your workload. A single B200 on packet.ai covers Llama 4 405B at FP8 (needs around 202 GB), Qwen3 235B at FP4 (around 118 GB), and 70B fine-tuning runs, all within the 192 GB HBM3e of one B200.
packet.ai B200 Dynamic starts at $3.75/GPU-hr, with no platform fee, no egress charge and no long-term contract. That is the lowest published on-demand Blackwell rate currently tracked by the SemiAnalysis GPU pricing comparison.
GB300 pays off where attention, FP4 compute and memory capacity are the constraints at the same time. It adds little where memory bandwidth, FP8 math or small models set the ceiling.
✓ GB300 NVL72 is right for
✗ Stay on GB200 or B200 for
Vera Rubin changes the math for long commitments, not for capacity you need this quarter. Supermicro began shipping Vera Rubin NVL72 racks on September 23, 2026, each with 72 Rubin GPUs, 36 Vera CPUs and 20.7 TB of HBM4. NVIDIA reports up to 3.7x the throughput of GB300 NVL72 in Rubin's first MLPerf Inference v6.1 submission.
First rack shipments are not broad cloud availability. For the next few quarters, GB300 and GB200 are the rack-scale options most teams can actually get.
A 12 to 36 month GB300 commitment signed today will overlap Rubin's cloud ramp. Prefer shorter terms, or ask providers about refresh options. GB200 is not obsolete either. CoreWeave's MLPerf v6.0 GB200 NVL72 result, 7,323 DeepSeek-R1 tokens per second per GPU offline, is higher than the 5,842 that GB300 posted in NVIDIA's v5.1 debut six months earlier. Different submitters and software versions, which is exactly the point: tuning keeps finding headroom on older silicon.
For most solo builders and indie teams, the GB300 vs GB200 debate is irrelevant. The real question is: does your workload fit in one B200? One B200 has 192 GB of HBM3e at 8 TB/s, enough for Llama 4 405B at FP8, Qwen3 235B at FP4 and most fine-tunes, at $3.75/GPU-hr with no minimum term on packet.ai. If your tensor or expert parallelism fits inside an 8-GPU NVLink island, an 8-GPU node skips the rack-scale question entirely. Our NVSwitch explainer covers where that 8-GPU NVLink domain ends and why that boundary matters.
On packet.ai, a single B200 starts at $3.75/GPU-hr on Dynamic and $6.99/GPU-hr on Dedicated, and a full 8x B200 Dedicated node costs $55.92/hr with hourly billing.
packet.ai B200 Dynamic at $3.75/GPU-hr is the lowest published on-demand Blackwell rate currently tracked by the SemiAnalysis GPU pricing comparison, 36% below the next-lowest published rate. No platform fee, no egress charge, no long-term contract. Dynamic nodes are SSH-ready in under 5 minutes; monthly commits cut up to 20% off the hourly rate, per the packet.ai pricing page. If you outgrow a single node, packet.ai cluster options cover multi-node B200 over InfiniBand on quoted terms.
Last reviewed: September 28, 2026. Specs from NVIDIA product pages and technical blogs; benchmarks from MLPerf submissions and SemiAnalysis InferenceX as cited. Before you price a rack, test your workload on a B200 on packet.ai from $3.75/GPU-hr.
Same models. Same API. Fraction of the cost. Start free — no credit card required.
Start Building →