AI Infra Interviews logo
NVIDIA · GPU · reviewed 2026-09-12

NVIDIA GB300 NVL72

A rack-scale system, not a single PCIe card. This comparison uses 279 GB per GPU from GCP’s four-GPU VM configuration; the full rack includes 72 GPUs and 36 Grace CPUs.

Memory per accelerator
279 GB

HBM3e

Memory bandwidth
8 TB/s

Published peak, not measured application throughput

Remember this

CPU LPDDR memory is not GPU HBM. Adding the two into a “fast memory” marketing total does not make them equally fast or automatically available to every tensor. The per-GPU numbers here must not be read as full-rack totals.

When this is a sensible choice

Start here if…

Use a rack-scale NVLink allocation for large models whose expert or tensor-parallel traffic benefits from staying in one fast domain. Confirm how many GPUs your reservation actually exposes and whether the ranks share that domain.

Choose another configuration if…

For a small replica, a simpler x86 GPU node may be easier to operate. Arm host binaries, containers, storage throughput and placement all need validation.

Specifications with their boundaries attached

Architecture
Grace Blackwell Ultra
Memory
279 GB HBM3e
Memory bandwidth
8 TB/s per accelerator
Peak compute
GCP dense per GPU: 2,500 TFLOPS BF16; 5,000 FP8
Scale-up interconnect
NVLink 5 rack domain: 72 GPUs; about 130 TB/s aggregate
Host attachment
Grace Arm CPU; NVLink-C2C CPU/GPU link
Power
Liquid-cooled rack; use full-system power plan
Partitioning
Allocation and partitioning are provider-specific
Catalogue status
Documented product

Compute figures are theoretical peaks at the stated precision. Dense and structured-sparse rates must not be mixed. Bandwidth labelled bidirectional combines both directions. See the source documents.

Follow the bytes · conceptual topology

Three bandwidths, three different jobs

Local memory279 GB HBM3e
Compute enginesExecute kernels on these bytes
① Memory bandwidth: 8 TB/s
Accelerator AOwn local memory
Accelerator BOwn local memory
Scale-up: NVLink, Infinity Fabric or PCIe
NVLink 5 rack domain: 72 GPUs; about 130 TB/s aggregate
Server AAccelerators + host
Server BAnother fabric domain
③ Scale-out: NICs + switches + placement
InfiniBand, RoCE, EFA or provider-specific transport
A 400 Gb/s NIC has a 50 GB/s raw line-rate equivalent before overhead. A 900 GB/s bidirectional NVLink figure counts traffic in both directions. Neither is the bandwidth at which a GPU reads its own HBM. This diagram explains the boundaries; it is not a wiring diagram for a particular cloud machine.

Where it appears in provider documentation

Documented configurations, checked September 12, 2026. Listing does not guarantee regional stock, quota, allocation size or an on-demand purchase.
Provider / machineNetwork scopeWhat changes the decision
Google Cloud
A4X Max · a4x-maxgpu-4g-metal
3,600 Gb/s maximum machine egressBare metal: four GPUs, 1,116 GB total GPU memory, two Grace CPUs. Capacity reservation required.
Azure
ND GB200 v6 / ND GB300 v6
Check the reserved system and fabric domainListed in Azure’s CUDA platform catalogue. Listing does not establish quota or capacity in your subscription.

Model fit and software support

These publisher or serving-engine documents mention this hardware family. They have not been reproduced on our machines.

2026-07-27

Kimi K3

Multimodal MoE · KDA / MLA · Kimi K3 License

2,800B · total parameter proxy · 104B active in the main model

2.8T parameters imply a 1.4 TB raw 4-bit floor, before scales and non-4-bit tensors. Eight 180 GB B200s leave only 40 GB above that idealized floor. Native MXFP4 and hybrid attention require a matching serving stack; active parameters do not make this a single-card model. The vLLM recipe specifies at least eight GB300 GPUs, or eight MI355X/MI350X on ROCm; these are published configurations, not measurements reproduced here.

Published serving recipe ↗
Compatibility evidence; no local throughput or latency measurement.

vLLM’s dated day-zero post confirms public weights on July 27. Date evidence ↗

What is the memory floor?

Start with total parameters, then add the memory the workload needs. This arithmetic does not certify a serving configuration. All output sizes below are decimal GB.

Override device capacity with the memory exposed by your allocation, particularly for cloud B300 and partitioned devices. The starting 32 GB budget and 15% reserve are editable teaching assumptions. They are not measurements for the selected model. Mixed-precision tensors, quantization scales, vision encoders and draft models can increase the weight payload.

320 GB

Raw weights only
320B × 8 bits ÷ 8

237.15 GB

Budget per device
279 GB × (1 − 15%)

2 devices

Arithmetic lower bound
round up ((weights + 32) ÷ budget)

WeightsCache + runtimeRemaining budget

This assumes perfectly balanced sharding. A result of three does not prove that a three-device parallel layout is supported. Check layer/expert divisibility, actual allocatable memory, precision kernels and the fabric before renting.

320B is the model card’s total; 18B active is not its storage size. The vLLM recipe reports about 306 GiB for the native FP8 checkpoint. Hopper requires BF16 KV for this model; the documented ROCm path targets gfx950, not every Instinct GPU. Read the model source ↗

For full training, also budget gradients, optimizer states, activations and communication buffers. The inference weight estimate above is insufficient.

Tokens per second, TTFT and TPOT are not certified for this hardware in our reference. Use the benchmark checklist to compare an exact model, software revision and workload.

Nearby memory capacities, different tradeoffs

These are comparison candidates selected by memory capacity, not performance rankings or drop-in replacements.

NVIDIA B300 HGX

288 GB · 8 TB/s

Blackwell Ultra’s larger HBM budget helps large models and long context. A cloud allocation may expose less memory than the 288 GB physical-product figure.

AMD Instinct MI355X

288 GB · 8 TB/s

288 GB of HBM on an AMD CDNA 4 accelerator. The memory is useful only after the model’s kernels and serving engine work on the ROCm target.

AMD Instinct MI350X

288 GB · 8 TB/s

MI350X and MI355X share 288 GB of HBM3e and 8 TB/s peak bandwidth. MI350X has a lower board-power specification and lower compute peaks.

Sources and review scope

Manufacturer and cloud documentation checked 2026-09-12. We reviewed specifications and the stated product boundaries; we did not run training or serving benchmarks on this device. Cloud memory and availability can differ by configuration.

  1. NVIDIA GB300 NVL72 specifications · checked 2026-09-12
  2. Google Compute Engine GPU machine types · checked 2026-09-12

Cite this reference

AI Infra Interviews. NVIDIA GB300 NVL72: specifications and workload fit. Reviewed . https://aiinfrainterviews.com/hardware/gb300

Include your access date when citing a changing specification. Link to the specification section for hardware figures or the explanation for a sizing or workload decision.

Research method and limits

We compile manufacturer specifications, cloud documentation, model cards and serving recipes. The reference preserves source units and distinguishes individual devices from nodes and racks. Conflicting figures and unknown fields remain labelled.

Our contribution is the comparison, unit reconciliation, worked arithmetic and workload explanation. Published peaks are vendor specifications. Calculator results are estimates under the displayed assumptions. Neither is a measurement from our own accelerator lab.

For a manufacturer’s specification, consult the original source documents. Cite our page when using its analysis, and retain primary-source attribution for underlying figures. This curated reference does not establish market share, live availability or a universal performance ranking.

Found a discrepancy? Send a correction with the page URL, exact variant, disputed figure and supporting primary document.