On an HBM part ECC costs almost nothing; on a GDDR part about 6% of capacity and bandwidth. What it buys is the difference between a corrected bit and a silently flipped exponent that turns 1.0 into infinity across a 512-GPU all-reduce. The mechanisms, the signals, and the expected-loss arithmetic.
What does ECC on a GPU cost you, and why do you keep it on across a fleet?
On an HBM part ECC costs almost nothing; on a GDDR part about 6% of capacity and bandwidth. What it buys is the difference between a corrected bit and a silently flipped exponent that turns 1.0 into infinity across a 512-GPU all-reduce. The mechanisms, the signals, and the expected-loss arithmetic.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on distinguishing detection from correction, knowing the HBM-versus-GDDR cost difference, naming the row-remapping and Xid signals an operator watches, and reasoning about silent corruption as a fleet-scale expected cost.
No comments yet — be the first to share your approach.
