A tail that wide is queueing or prompt-length variance, and the two are distinguishable in one measurement. Four causes ranked by how often they are it, the decomposition that attributes the wait, and the fix that is usually a flag rather than hardware.
Your p99 time to first token is four times p50. Find out why.
A tail that wide is queueing or prompt-length variance, and the two are distinguishable in one measurement. Four causes ranked by how often they are it, the decomposition that attributes the wait, and the fix that is usually a flag rather than hardware.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on decomposing TTFT into queue wait and prefill, on prompt-length variance as the usual cause, and on chunked prefill and admission as the remedies.
No comments yet — be the first to share your approach.
