Registers are allocated in fixed steps out of a fixed budget per multiprocessor, so a small increase in live values can cost a whole resident block. The compiler output that shows it in two lines, the occupancy cliff arithmetic, and the two different failure modes that produce the same symptom.
You added three local variables to a working kernel and it got 30 percent slower. Explain what happened and how you would confirm it.
Registers are allocated in fixed steps out of a fixed budget per multiprocessor, so a small increase in live values can cost a whole resident block. The compiler output that shows it in two lines, the occupancy cliff arithmetic, and the two different failure modes that produce the same symptom.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on reading the ptxas register and spill lines, on the occupancy-step arithmetic rather than a vague appeal to pressure, and on separating a lost resident block from an actual spill to local memory.
No comments yet — be the first to share your approach.
