AI Infra Interviews logo
Behavioral & Ownership / 21
mediumNewCoreWeaveModalBaseten

Three incidents are open, your pager is going off, and a customer is escalating. What do you do first?

The first move is not technical. It is deciding who runs what, because one person serially debugging three incidents is slower than three people in parallel and much slower than one person coordinating. The ordering rule, and the two things that outrank everything else.

Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.

The first move is not technical. It is deciding who runs what, because one person serially debugging three incidents is slower than three people in parallel and much slower than one person coordinating. The ordering rule, and the two things that outrank everything else.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 283 remaining answers · ₹2,000 / $25

The concepts behind this question

Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.

Foundational
🧭 Ownership & Judgment
Talking About Cost and Capacity with LeadershipInfrastructure engineers are asked to justify large numbers to people who do not share their vocabulary, and the conversations go wrong in predictable ways: a technical objection with no alternative, a forecast with no assumptions, or a cost quoted in a unit the listener cannot act on. What works is a small number of costed options, a stated recommendation, the decision needed by a date, and every figure expressed in whatever the listener actually controls.
Foundational
🧭 Ownership & Judgment
Escalation That WorksEscalation has a reputation as a political act because most of it is done badly: a problem handed upward with no options and an implicit request that someone else choose a side. Done well it is a one-page artifact with two or three costed options, a recommendation, the decision needed, a date, and what you will do by default if no answer arrives. That last line is what converts a message into a decision, and it is the part almost everyone omits.
Foundational
🧭 Ownership & Judgment
Mentoring and Growing EngineersMentoring on an infrastructure team happens mostly under pressure, during incidents and reviews, where the instinct to take the keyboard resolves the problem faster and teaches nothing. The method that works is the mentee driving while the mentor asks questions, with a takeover condition agreed in advance so nobody negotiates it at two in the morning. It costs time, and choosing which situations can absorb that cost is the judgment being assessed.
Advanced
🩺 Fleet Reliability & Observability🔒 Premium
Incident Response for GPU FleetsAn incident on a GPU fleet is a training run that stopped, a serving endpoint burning its error budget, or a fleet-wide symptom nobody has explained yet. The response has a shape: detect, stabilize, diagnose, repair, return through the gate, write it up. The stabilizing move (drain the node, restart from checkpoint, or shift traffic) comes before the diagnosis, because a frontier run loses more per minute than any investigation is worth. This page gives the triage order, the 3am decision tree, the spare-capacity arithmetic behind drain-and-replace, and what a fleet postmortem has to contain.
UP NEXT ON YOUR JOURNEY
FEDITOR'S NOTE

Scored on delegating before debugging, on the explicit ordering rule with data loss and spreading damage at the top, and on communication being a task with an owner.

DISCUSSION · 0

No comments yet — be the first to share your approach.