Three one-line policy differences produce different fragmentation, and the metric that separates them is not utilization. The measured outcome on the same job sequence, why the largest free block is what matters on a training cluster, and the case where the intuitive policy is exactly wrong.
Place GPU jobs onto nodes. Compare first fit, best fit and worst fit, and say which one a training cluster wants.
Three one-line policy differences produce different fragmentation, and the metric that separates them is not utilization. The measured outcome on the same job sequence, why the largest free block is what matters on a training cluster, and the case where the intuitive policy is exactly wrong.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on measuring the largest contiguous free block rather than total free capacity, on best or first fit for gang scheduling, and on knowing worst fit spreads and destroys whole-node availability.
No comments yet — be the first to share your approach.
