11When do you write a kernel in Triton, and when do you have to drop down to CUDA?▼mediumNewOpenAITogether AI4 replies○ sign inTriton hands you a block of data and writes the thread-level code for you, which covers most of what a serving or training stack actually needs. The four things it does not give you, the performance you give up in each case with numbers, the reversal as Triton gains Hopper features, and how to decide in the room.Open full answer →
26Describe something you built that made researchers meaningfully more productive.▼mediumNewAnthropicMetaGoogle DeepMind4 replies◆ premiumMost platform tools are built for the platform team's model of the work rather than the work. What makes a tool get adopted, the measurement that proves it helped, and why the fast path for small jobs beats almost anything else you could build.Open full answer →