Splitting a model across devices buys memory and spends interconnect
Sharding a model across devices divides the weights and the cache, and adds a collective inside every layer. That communication happens at layer frequency on the decode path, so the interconnect between those devices decides whether the split helps or hurts.
15 MIN · PREMIUM
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
