← 🗂️ Scheduling & Orchestration
Advanced
Gang Scheduling with Kueue and Volcano
A distributed training job is 64 pods that start together or not at all: if 40 are running and 24 are Pending, the 40 hold their GPUs idle at a collective barrier waiting for ranks that may never come, and two such jobs can deadlock a whole cluster. Gang scheduling makes the job the unit of admission. Kueue and Volcano add queues, quotas, priorities and preemption on top, which is what turns a pile of GPUs into a platform several teams can share without starving each other.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS
Kubernetes, Slurm & GPU SchedulingWhat is gang scheduling, and what goes wrong on a Kubernetes cluster that does not have it?→Kubernetes, Slurm & GPU SchedulingEight research teams share 1,024 GPUs. Design the quota and fairness policy, and tell me how they will game it.→Kubernetes, Slurm & GPU SchedulingA 64-node gang-scheduled job has been pending for six hours while the cluster shows free GPUs. Debug it live.→Kubernetes, Slurm & GPU SchedulingDesign a job queue for 100k GPU jobs with preemption: what state, what ordering, and what happens when a quota owner returns?→Kubernetes, Slurm & GPU SchedulingSlurm or Kubernetes for a 2,000-GPU training cluster? Make the case, and tell me what you lose either way.→AI Infrastructure System DesignDesign a job scheduler for 100,000 jobs on a shared GPU cluster, with preemption and checkpointing. Show me the state machine.→
