AI Infra Interviews logo

The KV cache is what actually decides how many users fit

Weights are a fixed cost paid once per replica. The cache is paid per concurrent sequence and grows with every token, so it is the term that decides capacity. The formula, the head-count mistake that inflates it eightfold, and the levers that shrink it.

14 MIN · PLUS

a free account unlocks the core curriculum tier · no card