AI Infra Interviews logo
NVIDIA · DPU · reviewed 2026-09-12

NVIDIA BlueField-3 DPU

A DPU moves infrastructure work such as networking, security and storage off the host. It is not another name for Google TPU.

Memory per accelerator
DPU

Infrastructure offload

Memory bandwidth
N/A

Does not provide model-weight HBM

Remember this

A SuperNIC and a DPU can share technology while serving different roles. Read the actual server topology: GPU-to-GPU training traffic, storage traffic and management traffic need not traverse the same adapters.

When this is a sensible choice

Start here if…

Use a DPU when the platform needs isolated infrastructure services, storage offload or network policy enforcement. In the referenced HGX design, BlueField handles the north/south infrastructure network while separate SuperNICs serve GPU traffic.

Choose another configuration if…

Do not include DPU memory in an LLM weight budget. Buying a faster DPU does not automatically accelerate a model’s matrix multiplication.

Specifications with their boundaries attached

Architecture
Infrastructure offload
Memory
Not a model-memory device
Memory bandwidth
Not a model HBM bandwidth specification
Peak compute
Infrastructure processing; not an LLM Tensor Core accelerator
Scale-up interconnect
400 Gb/s-class networking; exact ports depend on board
Host attachment
PCIe 5.0 x16; board variants differ
Power
Board-specific; some require auxiliary power
Partitioning
DPU mode and NIC ownership are deployment-specific
Catalogue status
Documented product

Compute figures are theoretical peaks at the stated precision. Dense and structured-sparse rates must not be mixed. Bandwidth labelled bidirectional combines both directions. See the source documents.

Follow the bytes · conceptual topology

Where a DPU sits

Model acceleratorGPU or TPU computes the model
DPUNetwork / storage / isolation
The DPU handles infrastructure traffic; it does not add GPU HBM.
Accelerator AOwn local memory
Accelerator BOwn local memory
Scale-up: NVLink, Infinity Fabric or PCIe
Moves sharded tensors and collective results between devices
Server AAccelerators + host
Server BAnother fabric domain
③ Scale-out: NICs + switches + placement
InfiniBand, RoCE, EFA or provider-specific transport
A 400 Gb/s NIC has a 50 GB/s raw line-rate equivalent before overhead. A 900 GB/s bidirectional NVLink figure counts traffic in both directions. Neither is the bandwidth at which a GPU reads its own HBM. This diagram explains the boundaries; it is not a wiring diagram for a particular cloud machine.

Where it appears in provider documentation

No provider configuration has been verified for this exact variant in this reference. Consult the linked manufacturer documentation; catalogue absence here is not evidence that the hardware is unavailable.

Model fit and software support

These publisher or serving-engine documents mention this hardware family. They have not been reproduced on our machines.

No model-specific recipe in our reviewed set certifies this exact hardware. The memory calculator can narrow candidates, but it cannot establish software support. Read the model register.

Tokens per second, TTFT and TPOT are not certified for this hardware in our reference. Use the benchmark checklist to compare an exact model, software revision and workload.

Sources and review scope

Manufacturer and cloud documentation checked 2026-09-12. We reviewed specifications and the stated product boundaries; we did not run training or serving benchmarks on this device. Cloud memory and availability can differ by configuration.

  1. NVIDIA HGX B300 reference architecture · checked 2026-09-12

Cite this reference

AI Infra Interviews. NVIDIA BlueField-3 DPU: specifications and workload fit. Reviewed . https://aiinfrainterviews.com/hardware/bluefield-3

Include your access date when citing a changing specification. Link to the specification section for hardware figures or the explanation for a sizing or workload decision.

Research method and limits

We compile manufacturer specifications, cloud documentation, model cards and serving recipes. The reference preserves source units and distinguishes individual devices from nodes and racks. Conflicting figures and unknown fields remain labelled.

Our contribution is the comparison, unit reconciliation, worked arithmetic and workload explanation. Published peaks are vendor specifications. Calculator results are estimates under the displayed assumptions. Neither is a measurement from our own accelerator lab.

For a manufacturer’s specification, consult the original source documents. Cite our page when using its analysis, and retain primary-source attribution for underlying figures. This curated reference does not establish market share, live availability or a universal performance ranking.

Found a discrepancy? Send a correction with the page URL, exact variant, disputed figure and supporting primary document.