Home / Chapter 7 · Production and serving
Last edited · 6 min read
Why the GPU is idle
During text generation the GPU spends most of each step waiting for the weights and the KV cache to arrive from memory, and computes for only a fraction of that time. This one effect explains batching, quantisation, speculative decoding, and why output tokens cost more than input tokens.
In plain wordsA lightning-fast chef has to run to a storeroom at the far end of the building for every ingredient. Cooking for one person, he mostly runs. Cooking for a hundred at once, he runs just as much but cooks a hundred times more.
Grow the batch and watch when compute catches up with waiting for memory. Then switch to prefill
One model step
Illustrative model: 8 billion parameters in BF16 (16 GB of weights) on one H100 SXM with 3.35 TB/s of memory bandwidth and about 400 TFLOPS of realistically achievable compute (about 990 TFLOPS dense BF16 on the spec sheet). Reading and computing overlap, so a step takes as long as the longer of the two. The bar ignores the KV cache and overheads, so it is an upper bound; the message adds the cache for 4,000 tokens of context per conversation.
The per-token arithmetic
- A decode step produces one token per conversation but reads every weight from HBM. An 8B model in BF16 is 16 GB, which takes about 4.8 ms to read at 3.35 TB/s. That caps a single conversation at about 210 tokens per second, regardless of the GPU’s compute power.
- Arithmetic intensity is the number of operations per byte read from memory. Multiplying a vector by the weights costs 2 operations (a multiply and an add) per parameter, i.e. per 2 bytes in BF16: about 1 operation per byte for each conversation in the batch. An H100 SXM needs about 295 operations per byte on paper (989 TFLOPS / 3.35 TB/s) and about 120 at a realistic 400 TFLOPS. Counting the weights alone, it takes a batch of roughly 100–300 conversations before compute time catches up with read time.
- Prefill processes the whole prompt in one pass: weights read once serve thousands of tokens, so a single long prompt keeps the GPU fully busy. Prefill sets the time to first token (TTFT); decode sets the time between subsequent tokens.
Where batching stops being free
- A batch shares the cost of reading the weights, but not the KV cache: each conversation reads its own, so attention in decode stays memory-bound at any batch size. In an 8B model with GQA (like Llama 3 8B) the cache takes about 0.125 MB per token, so a conversation with 32k tokens of context is 4 GB. Sixteen such conversations read 64 GB of cache and only 16 GB of weights every step. “The generation loop and KV cache” covers the mechanism.
- Past the crossover point every extra conversation makes the step longer for everyone: total throughput grows more and more slowly while the time between tokens rises. A provider picks a point on that curve. Interactive traffic needs smaller batches and fast tokens; batch traffic can be packed tightly.
- In an MoE model one token reads only the experts it was routed to, but the tokens in a batch spread across different experts, so at larger batch sizes you read almost the whole model (see “Mixture of Experts”).
What it means for the system
- Single-stream speed ≈ memory bandwidth / bytes read per token (active weights plus KV cache). That is why quantising weights to 8 or 4 bits speeds up decode almost proportionally (see “Quantisation”), and why speculative decoding gets several tokens out of one read of the weights (see “Speculative decoding”).
- Cost per token is mostly a matter of batch size. Continuous batching and PagedAttention exist to fit the largest possible batch into limited memory (see “Continuous batching”).
- Output tokens are priced several times higher than input tokens (typically 4–8× at the large providers). Input goes through prefill in parallel, while every output token is a separate decode step that occupies a slot in the batch.
- Prefill and decode have different bottlenecks and get in each other’s way on a shared GPU. Large deployments split them into separate pools (disaggregated serving): DeepSeek-V3 describes prefill on 32-GPU units and decode on 320-GPU units, and vLLM, SGLang and NVIDIA Dynamo support the split. On a single server, chunked prefill softens the conflict.
- For decode you choose a GPU by memory bandwidth and capacity, not TFLOPS. The H200 has the same compute as the H100 but 4.8 TB/s and 141 GB of memory, so it generates faster with large models and long contexts. The Blackwell B200 (shipping as of September 2026) has 180 GB and 8 TB/s, but its compute grew just as much (about 280 operations per byte in BF16), so the conclusions here still hold.
Check yourself
Why is decode memory-bound, and what does that mean for serving?
Every decode step reads all the weights and the KV cache from memory but computes only one token per conversation: in BF16 that is about 1 operation per byte, while an H100 needs about 300 before compute becomes the bottleneck. The speed ceiling for one conversation is memory bandwidth divided by bytes read per token. Weights read once serve the whole batch, so batching raises throughput almost for free until KV cache memory or the latency budget runs out. Quantisation and speculative decoding speed up decode. Prefill is the opposite: it is compute-bound.
Po polsku
Każdy krok decode czyta z pamięci wszystkie wagi i KV cache, a liczy tylko jeden token na rozmowę: w BF16 to ok. 1 operacja na bajt, a H100 potrzebuje ok. 300, żeby wąskim gardłem stało się liczenie. Sufit prędkości jednej rozmowy to przepustowość pamięci podzielona przez bajty na token. Wagi przeczytane raz obsługują cały batch, więc batching podnosi przepustowość prawie za darmo, dopóki nie skończy się pamięć na KV cache albo budżet opóźnienia. Kwantyzacja i spekulacja przyspieszają decode. Prefill jest odwrotny: ogranicza go moc obliczeniowa.
Follow-up questions (4)
- Estimate the maximum speed of a 70B model in FP8 on one H100.
- 70 GB of weights at 3.35 TB/s is about 21 ms per step, so at most about 48 tokens per second per conversation, before you add the KV cache. That leaves about 10 GB of the 80 for the cache, so the batch will be small. In practice such a model is split across 2–4 GPUs (tensor parallelism), which adds up both bandwidth and memory.
- You have an SLO of 50 ms between tokens. How do you size the batch?
- Step time is roughly (weights + the KV cache of the whole batch) / memory bandwidth, as long as compute fits inside it. Grow the batch until the step reaches 50 ms with headroom for p99, or until you run out of cache memory. Beyond that point throughput only grows at the expense of everyone’s latency.
- Why split prefill and decode onto separate machines?
- Prefill is compute-bound and arrives in bursts; decode is memory-bound and runs continuously. On a shared GPU a long prefill stalls other users’ generation, while separate pools can be scaled and sharded independently. The price is transferring the KV cache from the prefill pool to the decode pool.
- Compute utilisation in decode is a few per cent. Is that a problem?
- Not by definition. At small batch sizes the right metric is memory bandwidth utilisation (MBU): bytes of weights and cache per step divided by step time and peak bandwidth. Compute utilisation (MFU) is what you optimise in prefill and training.