The first question is whether 1T is total or active, because storage follows one and compute follows the other. The state budget, the layout that falls out of it, why expert parallelism belongs inside NVLink, and the two failure modes a dense-model plan does not have: router imbalance and all-to-all congestion.
Design a training cluster for a one-trillion-parameter MoE. Size it, choose the parallel layout, and map it onto the fabric.
The first question is whether 1T is total or active, because storage follows one and compute follows the other. The state budget, the layout that falls out of it, why expert parallelism belongs inside NVLink, and the two failure modes a dense-model plan does not have: router imbalance and all-to-all congestion.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on separating total from active parameters in the first minute, on deriving the layout from the memory and communication budgets rather than naming a fashionable one, and on treating expert imbalance as a straggler problem with a measurable signal.
No comments yet — be the first to share your approach.
