Ten thousand requests a second is a fleet of hundreds of nodes, a router that has to know what every replica is caching, and an SLO pair that decides the batch on every one of them. Here is the sizing chain, node by node.
Design an inference platform for 10,000 requests per second on a 70B model. Size it and name the SLOs.
Ten thousand requests a second is a fleet of hundreds of nodes, a router that has to know what every replica is caching, and an SLO pair that decides the batch on every one of them. Here is the sizing chain, node by node.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the sizing chain (tokens per second, prefill FLOPs, decode bandwidth, KV capacity, then nodes with headroom), on a router and scheduler that keep prefix hits and goodput up, and on a cost number at the end with the assumptions visible.
No comments yet — be the first to share your approach.
