Two restarts a day at fifteen minutes each is the 2%. The loss decomposed into detection, rescheduling, reload and rewound work, the lever on each, the in-memory checkpoint that makes the interval a minute, the square-root pareto of interval against write cost, and the residual that only fewer failures can remove.
Your 4,096-GPU run loses about 2% of every day to restarts. Fix it, and tell me where the floor is.
Two restarts a day at fifteen minutes each is the 2%. The loss decomposed into detection, rescheduling, reload and rewound work, the lever on each, the in-memory checkpoint that makes the interval a minute, the square-root pareto of interval against write cost, and the residual that only fewer failures can remove.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on decomposing the loss per failure into named stages with minutes on each, on attacking the largest term first, on knowing that the interval and the write cost trade against each other, and on the residual set by the failure rate itself.
No comments yet — be the first to share your approach.
