Four objectives, each measured at a percentile because averages hide the experience you are promising. The error budget in minutes per month, the burn-rate arithmetic that catches a fast failure in an hour and a slow one in a day, and why two of the four need separate targets per traffic class.
Define the service level objectives for an LLM serving fleet, and the alerting that tells you when one is about to be missed.
Four objectives, each measured at a percentile because averages hide the experience you are promising. The error budget in minutes per month, the burn-rate arithmetic that catches a fast failure in an hour and a slow one in a day, and why two of the four need separate targets per traffic class.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on percentile-based latency objectives split by traffic class, on converting availability into minutes of budget, and on multi-window burn-rate alerting rather than threshold alerts.
No comments yet — be the first to share your approach.
