08How much faster is an H200 than an H100, really? Which workloads see the gain and which do not?▼mediumNewNVIDIACoreWeaveLambda4 repliesunlockedThe H200 has the same compute die as the H100 and costs more per hour. The datasheet gives two ratios, 1.43x bandwidth and 1.76x memory, and those two numbers decide exactly which workloads pay back the premium and which ones lose money on it.Open full answer →
12Explain what HBM is and why memory bandwidth, not compute, is the wall for LLM inference.▼mediumNewNVIDIAAMDMicron4 replies○ sign inHBM stacks DRAM dies on top of each other and wires them to the GPU through a silicon interposer with a 1,024-bit bus per stack. That design sets how much bandwidth and capacity a card can have, why the two scale together, and why they have grown more slowly than FLOPS across three generations.Open full answer →