← 🧩 GPU & Accelerator Architecture
Advanced
Trainium and Inferentia
AWS's accelerators trade the GPU's general-purpose flexibility for a compiler-driven design with separate tensor, vector, scalar and GPSIMD engines, software-managed on-chip SRAM, and a proprietary NeuronLink fabric. Trainium2 delivers 667 dense bf16 TFLOPS with 96 GB at 2.9 TB/s, Trainium3 about the same bf16 with 2.5 PFLOPS of fp8 and 4.9 TB/s. The pitch is cost per FLOP; the price is a kernel ecosystem you may have to build yourself, which is exactly what the AWS loop probes.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
GPU & Accelerator ArchitectureWould you move a cost-sensitive training and serving fleet from H100 to Trainium2? What do you gain, and what do you have to plan for?→Napkin Math, Cost & CapacityRank the H100, MI300X and Trainium2 by cost per token for decode→Hardware, Cabling & Cluster Build-OutA vendor claims their accelerator beats an H100 at half the price. How do you evaluate that?→Open-Weights Models & Serving EnginesYou must serve a frontier open-weights model on non-NVIDIA accelerators. Plan it.→GPU & Accelerator ArchitectureWalk me through the CUDA execution model: what are grids, blocks and warps, and what does the hardware actually schedule?→GPU & Accelerator ArchitectureDescribe the GPU memory hierarchy. Where can a byte live on an H100, and what does each level cost?→
