04A researcher needs 256 GPUs today and the cluster is full. How do you handle it?▼medium★ EssentialNewOpenAIAnthropicGoogle DeepMind4 repliesunlockedSaying no is easy and it costs you the relationship. The move that works is to make the queue visible, give the researcher something they control, and offer a smaller thing today. The three mechanisms that turn this from a recurring argument into a system.Open full answer →
08Your team is drowning in manual work. How do you decide what to automate first?▼mediumNewCoreWeaveGoogleModal4 repliesunlockedMeasure the hours before you rank anything, because the task that feels worst is usually not the one that costs most. The ordering rule that keeps automation from causing the outage it was meant to prevent, and why detection always ships before remediation.Open full answer →
21Three incidents are open, your pager is going off, and a customer is escalating. What do you do first?▼mediumNewCoreWeaveModalBaseten4 replies◆ premiumThe first move is not technical. It is deciding who runs what, because one person serially debugging three incidents is slower than three people in parallel and much slower than one person coordinating. The ordering rule, and the two things that outrank everything else.Open full answer →