512 GPUs at 400 Gb/s each is 205 Tb/s of injection bandwidth a two-tier fabric must carry with no oversubscribed link. The port arithmetic that gives 16 leaves and 8 spines, the rail wiring that keeps NCCL traffic one hop away, and the line where 900 GB/s of NVLink becomes 50 GB/s of InfiniBand.
Design a non-blocking network fabric for 512 H100s. How many switches, how are they wired, and where does the NVLink domain end?
512 GPUs at 400 Gb/s each is 205 Tb/s of injection bandwidth a two-tier fabric must carry with no oversubscribed link. The port arithmetic that gives 16 leaves and 8 spines, the rail wiring that keeps NCCL traffic one hop away, and the line where 900 GB/s of NVLink becomes 50 GB/s of InfiniBand.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the port-counting derivation for a 1:1 Clos, understanding of rail-optimized wiring and why it matches NCCL's traffic, and a clear statement of the NVLink-to-network bandwidth cliff.
No comments yet — be the first to share your approach.
