Launch overhead is real, and small kernels are where it eats you
Every launch pays a fixed host cost to submit it and a fixed device cost to fill and drain the grid, and neither shrinks with the work. A small-batch decode step of hundreds of short kernels can idle as long between kernels as inside them. Fusion removes launches, graphs remove submission and gaps, batching amortises.
15 MIN · PLUS
a free account unlocks the core curriculum tier · no card
