Platform and fleet
Run the machine that other people's jobs run on, and keep it honest under many tenants.
You want the cluster itself: scheduling, the fabric, node health, quota and the on-call that comes with all of it. The largest share of open roles, and the path that skips kernels on purpose.
A broken cluster to fix, a scheduling design under multi-tenancy, and fleet reliability questions with real failure signatures.
The course sequence
In this order. Each assumes the one before it.
The question tracks to drill
90 questions across 3 tracks, in the order this loop weights them. Device plugins, MIG, gang scheduling, containers, Linux, a broken cluster to fix, bare-metal fleets.
Device plugins and dynamic resource allocation, MIG, MPS and time-slicing, gang scheduling with Kueue and Volcano, topology-aware placement, multi-tenancy and quotas, Slurm versus Kubernetes, containers and cold starts: the platform round at CoreWeave, Modal, Nebius and every GPU cloud.
30 questionsDCGM, the XID taxonomy, ECC and row remapping, NVLink faults, stragglers and hangs, thermal and power events, node health checks, SLOs for training and serving, incident response and postmortems at fleet scale. The on-call reality most prep sites skip.
30 questionsNCCL and the collective algorithms, RDMA, InfiniBand versus RoCE, rail-optimized and fat-tree fabrics, congestion control, GPUDirect, parallel filesystems versus object storage, data loading and checkpoint I/O: the fabric and the disks that decide whether ten thousand GPUs act like one.
30 questionsWho hires for this
Grouped by the kind of employer, because archetype predicts the loop better than the brand does.
