Temperature zero removes the sampling randomness and leaves the floating-point kind. The batch your request lands in changes the reduction order, the logits move in the last bits, and a near-tie flips a token. The fix has a cost.
A customer sends the same prompt twice at temperature zero and gets different answers. Explain why, and what you can promise them.
Temperature zero removes the sampling randomness and leaves the floating-point kind. The batch your request lands in changes the reduction order, the logits move in the last bits, and a near-tie flips a token. The fix has a cost.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on naming the actual source (batch-dependent kernel reduction order, not sampling), on knowing that a flipped token at step 40 changes everything after it, and on giving a precise promise with its cost.
No comments yet — be the first to share your approach.
