Sixteen samples per prompt, a verifier that runs code, no value model, responses to 16k tokens: GRPO moves the cost from the learner to rollouts and rewards. The token arithmetic for one step, the tail that holds a batch for the longest sample, and the three places a fleet idles.
We are running GRPO at scale. What does the infrastructure have to do that plain RLHF did not, and where do the GPUs sit idle?
Sixteen samples per prompt, a verifier that runs code, no value model, responses to 16k tokens: GRPO moves the cost from the learner to rollouts and rewards. The token arithmetic for one step, the tail that holds a batch for the longest sample, and the three places a fleet idles.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on knowing what GRPO removes (the critic) and what it multiplies (samples per prompt), on doing the per-step token budget, and on naming the tail latency of the longest sample as the dominant idle.
No comments yet — be the first to share your approach.
