03InfiniBand or RoCE version 2 for a new GPU training cluster. Make the call and say what would reverse it.▼medium★ EssentialNewNVIDIACrusoexAI4 repliesunlockedBoth carry RDMA at the same line rate. One arrives lossless because the transport was designed that way, the other becomes lossless only if a set of switch settings is correct on every port. What that difference costs in operations, what it saves in money and hiring, and the fleet size where the answer flips.Open full answer →
17Why do large training clusters provision roughly 400 gigabits per second per GPU rather than more or less?▼mediumNewNVIDIAMetaxAI4 replies○ sign inThe number comes from one requirement: the gradient reduction has to finish inside the backward pass that produces it. Working that requirement backward gives a bandwidth per GPU, and the answer lands near the port speed the industry ships, which is not a coincidence.Open full answer →