15vLLM, SGLang or TensorRT-LLM: which engine do you pick for a new deployment, and what would change your mind?▼hardNewBasetenTogether AINVIDIA4 replies○ sign inThree engines, one hardware roofline, and the difference between them is which part of the roofline each one reaches first on your workload. The decision is a table, and the table has reversal conditions.Open full answer →
30A vendor claims 10,000 tokens per second per GPU. Sanity-check it.▼expert★ EssentialNewFireworksTogether AIGroq4 replies◆ premiumTwo bounds, bandwidth and compute, and five questions about what the number counts. For a 70B on an H100 the claim exceeds the bf16 compute peak and is impossible; for an 8B at high batch it is routine; for input tokens it is easy. The chain that tells the cases apart.Open full answer →