A shared pool recovers the GPUs inference holds for its peaks, but a preempted training gang takes minutes to give them back and an SLO breaks in seconds. The utilization of each design, the reclaim-time arithmetic against the traffic ramp, and the floor-plus-borrow split most fleets land on.
One fleet: training that wants every idle GPU, and inference with a p99 SLO. Separate pools, or one pool with preemption? Show the numbers.
A shared pool recovers the GPUs inference holds for its peaks, but a preempted training gang takes minutes to give them back and an SLO breaks in seconds. The utilization of each design, the reclaim-time arithmetic against the traffic ramp, and the floor-plus-borrow split most fleets land on.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on computing what the separate-pool design wastes and what the shared design risks, on knowing that reclaim time is minutes while an SLO breach is seconds, and on the floor-plus-borrow design with the ramp-rate condition that makes it safe.
No comments yet — be the first to share your approach.
