Normalization reads a row and writes a row, so a copy sets the ceiling and everything else is overhead you can remove. The six changes in the order a reviewer wants them, what each is worth, and the one that is a correctness fix rather than a speed fix, with the measured error that proves it.
A take-home gives you a working layernorm kernel at a tenth of memory bandwidth. Make it fast and justify every change.
Normalization reads a row and writes a row, so a copy sets the ceiling and everything else is overhead you can remove. The six changes in the order a reviewer wants them, what each is worth, and the one that is a correctness fix rather than a speed fix, with the measured error that proves it.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on computing the bandwidth ceiling before optimizing, on the ordering of the fixes, on knowing that fp32 accumulation is correctness rather than speed, and on reporting against the ceiling rather than against the starting point.
No comments yet — be the first to share your approach.
