NVIDIA H200 NVL
H200 memory capacity in a PCIe card for air-cooled servers. It retains 141 GB and 4.8 TB/s, with a lower compute peak than H200 SXM.
HBM3e
Published peak, not measured application throughput
The 141 GB figure is per GPU, not per bridge group. Seven 16.5 GB partitions do not expose the entire nominal card capacity. NVIDIA marks the table specifications as preliminary.
When this is a sensible choice
Start here if…
Evaluate NVL when the server must use PCIe cards and the workload needs more memory than H100 NVL. Check the actual two- or four-card bridge layout for sharded models.
Choose another configuration if…
An eight-GPU server does not imply an eight-GPU NVSwitch domain. Prefer a documented HGX topology if the workload needs all ranks tightly connected.
Specifications with their boundaries attached
- Architecture
- Hopper
- Memory
- 141 GB HBM3e
- Memory bandwidth
- 4.8 TB/s per accelerator
- Peak compute
- 835.5 TFLOPS BF16/FP16; 1,670.5 TFLOPS FP8, dense
- Scale-up interconnect
- Two- or four-way NVLink bridge: 900 GB/s per GPU
- Host attachment
- PCIe 5.0 x16; dual-slot air-cooled card
- Power
- Up to 600 W, configurable
- Partitioning
- Up to seven MIG instances; product table lists 16.5 GB each
- Catalogue status
- Documented product
Compute figures are theoretical peaks at the stated precision. Dense and structured-sparse rates must not be mixed. Bandwidth labelled bidirectional combines both directions. See the source documents.
Three bandwidths, three different jobs
Two- or four-way NVLink bridge: 900 GB/s per GPU
InfiniBand, RoCE, EFA or provider-specific transport
Where it appears in provider documentation
No provider configuration has been verified for this exact variant in this reference. Consult the linked manufacturer documentation; catalogue absence here is not evidence that the hardware is unavailable.
Model fit and software support
These publisher or serving-engine documents mention this hardware family. They have not been reproduced on our machines.
No model-specific recipe in our reviewed set certifies this exact hardware. The memory calculator can narrow candidates, but it cannot establish software support. Read the model register.
What is the memory floor?
Start with total parameters, then add the memory the workload needs. This arithmetic does not certify a serving configuration. All output sizes below are decimal GB.
Override device capacity with the memory exposed by your allocation, particularly for cloud B300 and partitioned devices. The starting 32 GB budget and 15% reserve are editable teaching assumptions. They are not measurements for the selected model. Mixed-precision tensors, quantization scales, vision encoders and draft models can increase the weight payload.
Raw weights only
320B × 8 bits ÷ 8
Budget per device
141 GB × (1 − 15%)
Arithmetic lower bound
round up ((weights + 32) ÷ budget)
This assumes perfectly balanced sharding. A result of three does not prove that a three-device parallel layout is supported. Check layer/expert divisibility, actual allocatable memory, precision kernels and the fabric before renting.
320B is the model card’s total; 18B active is not its storage size. The vLLM recipe reports about 306 GiB for the native FP8 checkpoint. Hopper requires BF16 KV for this model; the documented ROCm path targets gfx950, not every Instinct GPU. Read the model source ↗
For full training, also budget gradients, optimizer states, activations and communication buffers. The inference weight estimate above is insufficient.
Tokens per second, TTFT and TPOT are not certified for this hardware in our reference. Use the benchmark checklist to compare an exact model, software revision and workload.
Nearby memory capacities, different tradeoffs
These are comparison candidates selected by memory capacity, not performance rankings or drop-in replacements.
NVIDIA H200 SXM
Hopper with a larger, faster memory system: 141 GB at 4.8 TB/s. Its value over H100 is easiest to see in memory capacity and memory traffic.
NVIDIA B200 HGX
An HBM-rich Blackwell GPU for large training and inference jobs. Use the 180 GB HGX usable-memory specification, not an early 192 GB announcement.
NVIDIA GB200 NVL72
A rack-scale system, not a single PCIe card. This comparison uses 186 GB per GPU from GCP’s four-GPU VM configuration; the full rack includes 72 GPUs and 36 Grace CPUs.
Sources and review scope
Manufacturer and cloud documentation checked 2026-09-12. We reviewed specifications and the stated product boundaries; we did not run training or serving benchmarks on this device. Cloud memory and availability can differ by configuration.
- NVIDIA H200 specifications ↗ · checked 2026-09-12
Cite this reference
AI Infra Interviews. NVIDIA H200 NVL: specifications and workload fit. Reviewed . https://aiinfrainterviews.com/hardware/h200-nvl
Include your access date when citing a changing specification. Link to the specification section for hardware figures or the explanation for a sizing or workload decision.
Research method and limits
We compile manufacturer specifications, cloud documentation, model cards and serving recipes. The reference preserves source units and distinguishes individual devices from nodes and racks. Conflicting figures and unknown fields remain labelled.
Our contribution is the comparison, unit reconciliation, worked arithmetic and workload explanation. Published peaks are vendor specifications. Calculator results are estimates under the displayed assumptions. Neither is a measurement from our own accelerator lab.
For a manufacturer’s specification, consult the original source documents. Cite our page when using its analysis, and retain primary-source attribution for underlying figures. This curated reference does not establish market share, live availability or a universal performance ranking.
Found a discrepancy? Send a correction with the page URL, exact variant, disputed figure and supporting primary document.
