Modal LLM Inference & Serving interview questions
LLM Inference & Serving is a core part of the Modal AI Infrastructure Engineer loop. Prefill versus decode, the KV cache, PagedAttention and continuous batching, chunked prefill, speculative decoding, disaggregated serving, quantization, vLLM, SGLang and TensorRT-LLM, multi-LoRA and routing: hosting open-weight models at a latency SLO and a cost you can defend. Below are the llm inference & serving questions to prepare, the ones tagged to Modal first, then the highest-signal questions from our LLM Inference & Serving track, each with an answer written to a senior-engineer bar.
WHAT MODAL LOOKS FOR HERE · Rust and high-performance distributed systems. See the full Modal interview process →
LLM Inference & Serving questions tagged to Modal
More LLM Inference & Serving questions for Modal's loop
The highest-signal llm inference & serving questions candidates rate most useful, modeled on what Modal's AI Infrastructure Engineer loop tests.
Concepts behind Modal's LLM Inference & Serving round
The vocabulary and mental models these questions assume. Start with the foundations free; the deeper, interview-defining ideas are part of premium.
Modal's AI Infrastructure Engineer loop draws llm inference & serving questions such as "Design an autoscaler for GPU inference replicas that reacts to load without thrashing.", "A new replica has to load a 70B model and serve traffic in under a minute. Where do the seconds go, and how do you get there?", "nvidia-smi shows 30 percent utilization on our serving fleet. Is that a problem, and what would you look at instead?". Prefill versus decode, the KV cache, PagedAttention and continuous batching, chunked prefill, speculative decoding, disaggregated serving, quantization, vLLM, SGLang and TensorRT-LLM, multi-LoRA and routing: hosting open-weight models at a latency SLO and a cost you can defend. The full set, ordered easy to hard with expert answers, is below.
Other Modal interview rounds
The other tracks Modal's AI Infrastructure Engineer loop tests.
Prep the whole Modal AI Infrastructure Engineer loop
LLM Inference & Serving is one round. Unlock every answer across Modal's full loop, plus the concept curriculum, for 6 months. One payment, no auto-renewal. Free questions in every track to start.
Independent and not affiliated with Modal. All trademarks belong to their owners.
