AI Infra Interviews logo

A tiled matrix multiply is the shape every fast kernel borrows

A tiled matmul has six parts: an owned output tile, a slice loop, coalesced staging into shared memory, compute into register accumulators, two barriers per slice, and an epilogue. Attention, convolution and grouped expert matmuls are this shape with different answers to what is owned, streamed and finished.

18 MIN · PREMIUM

Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew