Each card reads an eighth of the weights in 5.3 ms, then the step pays 160 latency-bound all-reduces and hundreds of kernel launches that do not shrink with sharding. The chain to a 9 to 12 ms step, the communication floor, and why TP8 gives 4x rather than 8x at batch one.
Estimate the latency of one decode step for a 70B model under tensor parallelism across eight H100s
Each card reads an eighth of the weights in 5.3 ms, then the step pays 160 latency-bound all-reduces and hundreds of kernel launches that do not shrink with sharding. The chain to a 9 to 12 ms step, the communication floor, and why TP8 gives 4x rather than 8x at batch one.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
The interviewer wants three terms with numbers: the sharded weight read, the per-layer all-reduce latency times the count, and launch overhead. Then the sentence that the communication term is a floor independent of batch.
No comments yet — be the first to share your approach.
