Every collective ends when the last rank arrives, so a single throttled GPU taxes 16,383 others. The arithmetic of that tax, the per-rank timing that finds the rank in one step, and the ranked list of causes from a hot GPU to a NIC with symbol errors, each with the command that confirms it.
One GPU out of 16,000 is 15% slow and the whole run is 15% slow. How do you find it, and what is usually wrong with it?
Every collective ends when the last rank arrives, so a single throttled GPU taxes 16,383 others. The arithmetic of that tax, the per-rank timing that finds the rank in one step, and the ranked list of causes from a hot GPU to a NIC with symbol errors, each with the command that confirms it.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on explaining why synchronous training turns one slow rank into a fleet-wide loss, on instrumenting per-rank compute time rather than reading NCCL time, and on a cause list with a check for each.
No comments yet — be the first to share your approach.
