6ND to a compute budget, the fleet equation to a GPU count, the rental rate to a bill: 64 H100s, about 19 days, roughly $75k plus a buffer. The chain, the reasons the answer is 64 and not 59, and the memory check that says the fleet is compute-sized, not memory-sized.
A startup wants to train a 7B model on 1 trillion tokens in three weeks. What cluster do they rent?
6ND to a compute budget, the fleet equation to a GPU count, the rental rate to a bill: 64 H100s, about 19 days, roughly $75k plus a buffer. The chain, the reasons the answer is 64 and not 59, and the memory check that says the fleet is compute-sized, not memory-sized.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
The interviewer wants the smallest fleet that meets the deadline with stated MFU, rounded to whole nodes, plus a budget with a buffer and the memory check. A candidate who quotes GPUs without a node count or a calendar has not planned a rental.
No comments yet — be the first to share your approach.
