Home / Chapter 7 · Production and serving
    Last edited · 6 min read

    Use with AI

    Why the GPU is idle

    During text generation the GPU spends most of each step waiting for the weights and the KV cache to arrive from memory, and computes for only a fraction of that time. This one effect explains batching, quantisation, speculative decoding, and why output tokens cost more than input tokens.

    In plain wordsA lightning-fast chef has to run to a storeroom at the far end of the building for every ingredient. Cooking for one person, he mostly runs. Cooking for a hundred at once, he runs just as much but cooks a hundred times more.

    Grow the batch and watch when compute catches up with waiting for memory. Then switch to prefill

    One model step

    waiting for weights from memorycomputing
    of the step the GPU spends computing
    tokens per second for one conversation
    tokens per second in total

    Illustrative model: 8 billion parameters in BF16 (16 GB of weights) on one H100 SXM with 3.35 TB/s of memory bandwidth and about 400 TFLOPS of realistically achievable compute (about 990 TFLOPS dense BF16 on the spec sheet). Reading and computing overlap, so a step takes as long as the longer of the two. The bar ignores the KV cache and overheads, so it is an upper bound; the message adds the cache for 4,000 tokens of context per conversation.

    The per-token arithmetic

    Where batching stops being free

    What it means for the system

    Check yourself

    Why is decode memory-bound, and what does that mean for serving?

    Every decode step reads all the weights and the KV cache from memory but computes only one token per conversation: in BF16 that is about 1 operation per byte, while an H100 needs about 300 before compute becomes the bottleneck. The speed ceiling for one conversation is memory bandwidth divided by bytes read per token. Weights read once serve the whole batch, so batching raises throughput almost for free until KV cache memory or the latency budget runs out. Quantisation and speculative decoding speed up decode. Prefill is the opposite: it is compute-bound.

    Po polsku

    Każdy krok decode czyta z pamięci wszystkie wagi i KV cache, a liczy tylko jeden token na rozmowę: w BF16 to ok. 1 operacja na bajt, a H100 potrzebuje ok. 300, żeby wąskim gardłem stało się liczenie. Sufit prędkości jednej rozmowy to przepustowość pamięci podzielona przez bajty na token. Wagi przeczytane raz obsługują cały batch, więc batching podnosi przepustowość prawie za darmo, dopóki nie skończy się pamięć na KV cache albo budżet opóźnienia. Kwantyzacja i spekulacja przyspieszają decode. Prefill jest odwrotny: ogranicza go moc obliczeniowa.

    Follow-up questions (4)
    Estimate the maximum speed of a 70B model in FP8 on one H100.
    70 GB of weights at 3.35 TB/s is about 21 ms per step, so at most about 48 tokens per second per conversation, before you add the KV cache. That leaves about 10 GB of the 80 for the cache, so the batch will be small. In practice such a model is split across 2–4 GPUs (tensor parallelism), which adds up both bandwidth and memory.
    You have an SLO of 50 ms between tokens. How do you size the batch?
    Step time is roughly (weights + the KV cache of the whole batch) / memory bandwidth, as long as compute fits inside it. Grow the batch until the step reaches 50 ms with headroom for p99, or until you run out of cache memory. Beyond that point throughput only grows at the expense of everyone’s latency.
    Why split prefill and decode onto separate machines?
    Prefill is compute-bound and arrives in bursts; decode is memory-bound and runs continuously. On a shared GPU a long prefill stalls other users’ generation, while separate pools can be scaled and sharded independently. The price is transferring the KV cache from the prefill pool to the decode pool.
    Compute utilisation in decode is a few per cent. Is that a problem?
    Not by definition. At small batch sizes the right metric is memory bandwidth utilisation (MBU): bytes of weights and cache per step divided by step time and peak bandwidth. Compute utilisation (MFU) is what you optimise in prefill and training.

    Sources

    Report an error · Suggest a fix