The roofline, the sharding arithmetic and the parallelism trade-offs transfer unchanged, and the numbers are in the same units. What changes is who writes the kernels, how shapes must behave, and the interconnect topology you shard against. The v6e and v7 numbers worked through against an H100.
You have trained on GPUs. What transfers to training on TPUs, and what do you have to relearn?
The roofline, the sharding arithmetic and the parallelism trade-offs transfer unchanged, and the numbers are in the same units. What changes is who writes the kernels, how shapes must behave, and the interconnect topology you shard against. The v6e and v7 numbers worked through against an H100.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the candidate applying the same quantitative reasoning to a TPU spec sheet, then naming precisely which habits break: dynamic shapes, hand-written kernels, NVLink-shaped parallelism assumptions.
No comments yet — be the first to share your approach.
