← 🕸️ Distributed Training
Advanced
Forward Pass, Backpropagation and Optimizer Steps
Follow one parameter through prediction, loss, gradient and update. Understand why backward does not itself train the weights, why training needs extra memory, and how microbatches differ from optimizer steps.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS
Distributed Training & ParallelismIn data-parallel training, what actually gets communicated between GPUs, and how much is it per step?→CUDA, Triton & Kernel EngineeringA training step runs at 20 percent model FLOPs utilization. Profile it and find where the missing time goes.→Napkin Math, Cost & CapacityEstimate how long it takes to write a checkpoint for a 405B training run→Napkin Math, Cost & CapacityHow long does an all-reduce of a 70B model's gradients take on eight GPUs?→Napkin Math, Cost & CapacityHow many H100s do you need to train a 70B model on 15 trillion tokens in 30 days?→Napkin Math, Cost & CapacityHow long does that 70B run take on 16,384 H100s at 40% MFU?→
