Research infrastructure
Plan, run and rescue training jobs large enough that failure is a schedule item.
You want to sit next to researchers and keep very large runs moving: parallelism layout, collectives, checkpoints and the failure math.
Parallelism strategy derived rather than recited, collective cost arithmetic, and a run that degraded for you to diagnose.
The course sequence
In this order. Each assumes the one before it.
The question tracks to drill
92 questions across 3 tracks, in the order this loop weights them. Parallelism strategy, collectives, RDMA fabrics, a slow all-reduce to debug, fault-tolerant training.
DDP, ZeRO and FSDP, tensor, pipeline, context and expert parallelism, collectives and their cost, MFU, activation checkpointing, elastic and fault-tolerant training, checkpoint economics and the RL post-training stack. Owning the training run at cluster scale.
32 questionsNCCL and the collective algorithms, RDMA, InfiniBand versus RoCE, rail-optimized and fat-tree fabrics, congestion control, GPUDirect, parallel filesystems versus object storage, data loading and checkpoint I/O: the fabric and the disks that decide whether ten thousand GPUs act like one.
30 questionsDCGM, the XID taxonomy, ECC and row remapping, NVLink faults, stragglers and hangs, thermal and power events, node health checks, SLOs for training and serving, incident response and postmortems at fleet scale. The on-call reality most prep sites skip.
30 questionsWho hires for this
Grouped by the kind of employer, because archetype predicts the loop better than the brand does.
