Two bounds, bandwidth and compute, and five questions about what the number counts. For a 70B on an H100 the claim exceeds the bf16 compute peak and is impossible; for an 8B at high batch it is routine; for input tokens it is easy. The chain that tells the cases apart.
A vendor claims 10,000 tokens per second per GPU. Sanity-check it.
Two bounds, bandwidth and compute, and five questions about what the number counts. For a 70B on an H100 the claim exceeds the bf16 compute peak and is impossible; for an 8B at high batch it is routine; for input tokens it is easy. The chain that tells the cases apart.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
This is the capstone: every earlier chain in one answer. The score is on the candidate computing both bounds for a stated model and precision, then asking the five clarifying questions in the right order and naming the configuration under which the claim is true.
No comments yet — be the first to share your approach.
