The source translation is a script and takes an afternoon. The assumptions underneath it are the problem: a wavefront is 64 lanes rather than 32, so a warp-level reduction written for 32 ports cleanly and reduces half the data. What breaks silently, what breaks loudly, and the numbers that change every tuning decision.
Port a hand-written CUDA kernel to MI300X. What translates mechanically, and what silently computes the wrong answer?
The source translation is a script and takes an afternoon. The assumptions underneath it are the problem: a wavefront is 64 lanes rather than 32, so a warp-level reduction written for 32 ports cleanly and reduces half the data. What breaks silently, what breaks loudly, and the numbers that change every tuning decision.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on naming wavefront 64 as the silent correctness hazard, on the local-memory and matrix-instruction differences that force retuning, and on knowing that the machine balance differs so the tuned configuration does not transfer.
No comments yet — be the first to share your approach.
