Put the weights in on-chip SRAM and the HBM wall disappears: tens of terabytes per second per chip, and a batch-1 step limited by the pipeline rather than the memory. The price is capacity: a 70B needs hundreds of chips per replica, and the cost per token depends on keeping every one of them busy.
Why are Groq and Cerebras so fast at batch 1, and what does that speed cost at scale?
Put the weights in on-chip SRAM and the HBM wall disappears: tens of terabytes per second per chip, and a batch-1 step limited by the pipeline rather than the memory. The price is capacity: a 70B needs hundreds of chips per replica, and the cost per token depends on keeping every one of them busy.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on deriving the batch-1 advantage from bandwidth and the scale disadvantage from capacity, using the same roofline the candidate would use for a GPU, and stating the batch regime where each design wins.
No comments yet — be the first to share your approach.
