18How many spare GPUs, nodes, cables and transceivers do you hold for a 2,048-GPU fleet?▼mediumNewCoreWeaveMetaLambda Labs5 replies○ sign inThe spares number falls out of the failure rate times the repair turnaround, and the cheap items are the ones people forget. The arithmetic for GPUs, the very different arithmetic for transceivers, and the rule that decides whether to hold a part at all.Open full answer →
27Write the policy for when a GPU is replaced rather than returned to service. What are the triggers and what do they cost?▼hardNewCoreWeaveLambdaMeta4 replies◆ premiumEvery replacement costs a spare, a maintenance window and a return process; every device kept costs the risk of a job it will fail. Four triggers that decide it without a case-by-case argument, the arithmetic that sets the repeat threshold, and why the device's history beats any diagnostic result.Open full answer →
28How much spare capacity does a 16,000-GPU fleet need, and what are you actually reserving it for?▼hardNewMetaMicrosoft4 replies◆ premiumThree separate reserves get merged into one number and then argued about. The repair pipeline from Little's law, the restart pool that has to be instantly available, and the correlated-failure buffer sized by the largest thing that can fail at once, each derived and then added.Open full answer →