AI Infra Interviews logo
Hardware, Cabling & Cluster Build-Out / 34
hardNewMicrosoftCrusoeCoreWeave

Your next hardware generation doubles power per rack. Plan the hall's upgrade.

Doubling per-rack density is four separate upgrades with different lead times, and the electrical one usually cannot be done while the hall runs. The phasing that keeps capacity available through it, the constraint that decides whether it is possible at all, and the number that says whether to upgrade or move.

Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.

Doubling per-rack density is four separate upgrades with different lead times, and the electrical one usually cannot be done while the hall runs. The phasing that keeps capacity available through it, the constraint that decides whether it is possible at all, and the number that says whether to upgrade or move.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 283 remaining answers · ₹2,000 / $25

The concepts behind this question

Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.

Foundational
🧮 Open Weights & Serving Engines
Capacity Planning for Open-Weights FleetsPlanning a fleet for a sparse open-weights model works differently from planning one for a dense model, because memory follows total parameters and throughput follows active parameters, and those now differ by more than twenty times. The sizing goes in one direction only: from a traffic forecast to tokens per second, to replicas at a measured operating point, to GPUs, to racks and kilowatts. Doing it in the other direction, from an available GPU count, produces a fleet that fits the hardware rather than the demand.
Foundational
🖧 Hardware & Cluster Build-Out
Direct-to-Chip Liquid Cooling and CDUsAbove roughly 40 kW a rack cannot be cooled by air in any practical hall, which is why every dense GPU deployment now runs liquid to the chip. A cold plate sits on each GPU, a coolant distribution unit isolates the clean rack loop from facility water, and the facility side runs warm, typically 30 to 40 degrees supply, because warm water is cheaper to make. The design numbers are flow rate and temperature rise, and both fall out of one equation that every operator should be able to do from memory.
Advanced
🧮 Napkin Math & Capacity🔒 Premium
Power and Datacenter ConstraintsThe binding constraint on new GPU capacity in 2026 is not chips or capital but megawatts: an H100 node draws about 10 kW, a GB200 NVL72 rack about 120 kW, and a 100,000-GPU cluster needs on the order of 150 MW with cooling. This page converts GPU counts to power, power to cooling and facility requirements, and both to cost, so a candidate can size a training hall from a power budget and explain why liquid cooling, PUE and the local grid decide where the next cluster goes.
Foundational
🖧 Hardware & Cluster Build-Out
Rack Power Delivery and BuswaysA GPU rack has gone from 10 kW to over 120 kW in a few generations, and the electrical design changed with it. At 132 kW on a 415 V three-phase feed a rack draws about 184 amps, which is past what a normal power strip carries, so distribution moves to overhead busway and the rack takes redundant high-current taps. On top of the steady draw sits a synchronized transient every training step, because thousands of GPUs finish a collective at the same instant, and that swing is what sizes the upstream equipment.
UP NEXT ON YOUR JOURNEY
FEDITOR'S NOTE

Scored on identifying which upgrades can happen in a live hall, on phasing to avoid a capacity trough, and on comparing upgrade against relocation with a number.

DISCUSSION · 0

No comments yet — be the first to share your approach.