Perplexity LLM Inference & Serving interview questions
LLM Inference & Serving is a core part of the Perplexity AI Infrastructure Engineer loop. Prefill versus decode, the KV cache, PagedAttention and continuous batching, chunked prefill, speculative decoding, disaggregated serving, quantization, vLLM, SGLang and TensorRT-LLM, multi-LoRA and routing: hosting open-weight models at a latency SLO and a cost you can defend. Below are the llm inference & serving questions to prepare, the ones tagged to Perplexity first, then the highest-signal questions from our LLM Inference & Serving track, each with an answer written to a senior-engineer bar.
WHAT PERPLEXITY LOOKS FOR HERE · Serving with a retrieval step inside the latency budget. See the full Perplexity interview process →
LLM Inference & Serving questions tagged to Perplexity
More LLM Inference & Serving questions for Perplexity's loop
The highest-signal llm inference & serving questions candidates rate most useful, modeled on what Perplexity's AI Infrastructure Engineer loop tests.
Concepts behind Perplexity's LLM Inference & Serving round
The vocabulary and mental models these questions assume. Start with the foundations free; the deeper, interview-defining ideas are part of premium.
Perplexity's AI Infrastructure Engineer loop draws llm inference & serving questions such as "When does splitting prefill and decode onto separate GPU pools pay for itself, and what does the KV transfer cost?", "How do you route requests across replicas to maximize prefix-cache hits without unbalancing the fleet?", "When does offloading the KV cache to CPU memory or NVMe beat recomputing it?". Prefill versus decode, the KV cache, PagedAttention and continuous batching, chunked prefill, speculative decoding, disaggregated serving, quantization, vLLM, SGLang and TensorRT-LLM, multi-LoRA and routing: hosting open-weight models at a latency SLO and a cost you can defend. The full set, ordered easy to hard with expert answers, is below.
Other Perplexity interview rounds
The other tracks Perplexity's AI Infrastructure Engineer loop tests.
Prep the whole Perplexity AI Infrastructure Engineer loop
LLM Inference & Serving is one round. Unlock every answer across Perplexity's full loop, plus the concept curriculum, for 6 months. One payment, no auto-renewal. Free questions in every track to start.
Independent and not affiliated with Perplexity. All trademarks belong to their owners.
