Two all-to-alls per layer across sixty layers is over a hundred collectives on the critical path of every decode step, each one small and latency-bound. The per-GPU volume worked out, the time on each link type, and why a single large NVLink domain changes the design rather than just improving it.
A mixture-of-experts model does an all-to-all twice per layer. What does that demand of the fabric, and what changes on a rack-scale system?
Two all-to-alls per layer across sixty layers is over a hundred collectives on the critical path of every decode step, each one small and latency-bound. The per-GPU volume worked out, the time on each link type, and why a single large NVLink domain changes the design rather than just improving it.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on computing the per-GPU all-to-all volume rather than the aggregate, on recognising the collectives as latency-bound rather than bandwidth-bound, and on why a wide NVLink domain changes what is buildable.
No comments yet — be the first to share your approach.
