Roll back first and investigate afterward, because the error budget is being consumed while you read logs. The rollback decision rule that removes the argument, the four causes specific to model serving, and the canary design that would have caught it at one percent of the traffic.
Error rate on an inference fleet tripled ten minutes after a deploy. What do you do first, and what should have caught it?
Roll back first and investigate afterward, because the error budget is being consumed while you read logs. The rollback decision rule that removes the argument, the four causes specific to model serving, and the canary design that would have caught it at one percent of the traffic.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on rolling back before diagnosing, on the specific serving failure modes rather than generic deploy bugs, and on a canary with an automatic promotion and rollback rule.
No comments yet — be the first to share your approach.
