Unlike a training job, a slow serving replica holds nobody else back, so it survives unnoticed while degrading a slice of users. The five causes and the metric that distinguishes each, why routing can create the symptom with no hardware fault, and the fix that limits damage while you investigate.
One serving replica has a per-token latency 40 percent worse than its peers. Find out why.
Unlike a training job, a slow serving replica holds nobody else back, so it survives unnoticed while degrading a slice of users. The five causes and the metric that distinguishes each, why routing can create the symptom with no hardware fault, and the fix that limits damage while you investigate.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on knowing a slow replica affects a slice of traffic rather than the fleet, on distinguishing a hardware cause from a routing-induced one, and on shedding traffic from it as the immediate action.
No comments yet — be the first to share your approach.
