One is a write burst of terabytes in a minute followed by half an hour of silence; the other is a steady read that never stops and never spikes. Provisioning for the peak of the first wastes most of its capacity, and provisioning for the average of the second fails the moment a checkpoint lands.
Checkpoints and datasets both live on storage. Why does one system sized for both usually serve neither well?
One is a write burst of terabytes in a minute followed by half an hour of silence; the other is a steady read that never stops and never spikes. Provisioning for the peak of the first wastes most of its capacity, and provisioning for the average of the second fails the moment a checkpoint lands.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on characterizing the two access patterns by burstiness rather than volume, on the provisioning arithmetic that shows why one system serves neither, and on the tier assignment that follows.
No comments yet — be the first to share your approach.
