Capacity is the easy half and almost never the constraint. The number that sizes the system is the checkpoint write burst, which is thousands of times the steady dataset read, and the tier that serves each is different. The full derivation, and the tiering that avoids paying burst prices for archive bytes.
Size the storage for a 2,048-GPU training cluster. What numbers actually drive it?
Capacity is the easy half and almost never the constraint. The number that sizes the system is the checkpoint write burst, which is thousands of times the steady dataset read, and the tier that serves each is different. The full derivation, and the tiering that avoids paying burst prices for archive bytes.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on sizing from the checkpoint burst rather than dataset capacity, on separating burst bandwidth from steady read, and on tiering so archive capacity is not bought at burst prices.
No comments yet — be the first to share your approach.
