Decode is bandwidth-bound, so cost per token tracks dollars per terabyte per second, and on spec the MI300X wins by 2x. The table, the memory-capacity term that changes the replica shape, and the software-efficiency discount that decides whether the spec ranking survives a benchmark.
Rank the H100, MI300X and Trainium2 by cost per token for decode
Decode is bandwidth-bound, so cost per token tracks dollars per terabyte per second, and on spec the MI300X wins by 2x. The table, the memory-capacity term that changes the replica shape, and the software-efficiency discount that decides whether the spec ranking survives a benchmark.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
The interviewer wants the candidate to pick bandwidth per dollar as the metric, build the table, and then discount it for measured efficiency and capacity. A ranking by TFLOPS is the wrong metric for decode and is the common miss.
No comments yet — be the first to share your approach.
