HBM stacks DRAM dies on top of each other and wires them to the GPU through a silicon interposer with a 1,024-bit bus per stack. That design sets how much bandwidth and capacity a card can have, why the two scale together, and why they have grown more slowly than FLOPS across three generations.
Explain what HBM is and why memory bandwidth, not compute, is the wall for LLM inference.
HBM stacks DRAM dies on top of each other and wires them to the GPU through a silicon interposer with a 1,024-bit bus per stack. That design sets how much bandwidth and capacity a card can have, why the two scale together, and why they have grown more slowly than FLOPS across three generations.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on a correct physical picture of HBM (stacks, interposer, wide bus), the per-stack arithmetic, and the generation-over-generation ratio that shows bandwidth falling behind compute.
No comments yet — be the first to share your approach.
