A collective stops when one participant stops, and a job of a thousand ranks has no way to distinguish a link that will come back in two seconds from a rank that has died. What the job actually does during those seconds, the cost per flap, and why the fault passes every test you would run.
One fabric link goes down and comes back once an hour. What does that do to a 1,024-GPU training job, and how would you find it?
A collective stops when one participant stops, and a job of a thousand ranks has no way to distinguish a link that will come back in two seconds from a rank that has died. What the job actually does during those seconds, the cost per flap, and why the fault passes every test you would run.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on knowing a collective blocks all ranks rather than degrading, on computing the cost per flap against the timeout policy, and on naming continuous counters as the only instrument that finds an intermittent fault.
No comments yet — be the first to share your approach.
