Sixteen bytes per parameter is 6.5 TB per checkpoint; over a 100 GB/s filesystem that is about a minute, which at a 30-minute cadence is 3% of the run. The chain, the storage tiers that meet it, and why the write is made asynchronous.
Estimate how long it takes to write a checkpoint for a 405B training run
Sixteen bytes per parameter is 6.5 TB per checkpoint; over a 100 GB/s filesystem that is about a minute, which at a 30-minute cadence is 3% of the run. The chain, the storage tiers that meet it, and why the write is made asynchronous.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
The interviewer wants the candidate to know what is in a checkpoint (optimizer state dominates), to size it, and to convert the write time into lost MFU at a given cadence. The async answer is the follow-up.
No comments yet — be the first to share your approach.
