A three-month run on 8,000 GPUs is a large bet against everything that can happen to one building. The recovery point and the recovery time as numbers, the checkpoint replication that sets the first, the second cluster and its warm state that set the second, and the drills that make the numbers true.
Design disaster recovery for a three-month training run: checkpoint replication, cluster failover, and the RTO you can promise.
A three-month run on 8,000 GPUs is a large bet against everything that can happen to one building. The recovery point and the recovery time as numbers, the checkpoint replication that sets the first, the second cluster and its warm state that set the second, and the drills that make the numbers true.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on RPO and RTO as derived numbers, on replication that keeps up with the checkpoint cadence, on a failover cluster whose readiness is measured in hours not weeks, and on drills as part of the plan rather than a wish.
No comments yet — be the first to share your approach.
