Home / Chapter 3 · How a model writes and what it costs
    Last edited · 7 min read

    Use with AI

    The generation loop and KV cache

    Generation is a loop: one forward pass of the model yields one token, which is appended to the input. The KV cache holds the attention keys and values of all earlier tokens, so each step computes only the new token. The price is GPU memory, which limits context length and the number of concurrent conversations.

    In plain wordsYou are writing a letter, and with every new word you glance at notes on what you have already written. You don’t reread the letter, but you do look through the notes every time. The KV cache is those notes. They sit on the desk, that is, in GPU memory, and the longer the letter, the more room they take.

    Step through with and without the cache. Compare how many tokens each step computes

    computed in this stepread from the cachenew token
    0tokens computed in this step
    0in total since the start
    0read from the cache in this step
    –phase

    What the cache holds and why it works

    How much memory it takes

    Pick a model, context length and number of conversations. See how many H100s the cache alone takes up

    on average per token
    per conversation
    in total
    H100s (80 GB) for the cache alone

    Formula: 2 × layers × KV heads × head dimension × bytes × tokens. Instead of K and V, MLA stores one vector of 576 numbers per layer, so there is no factor of 2. In Gemma 3 every sixth layer is full and the others keep at most the last 1024 tokens. Model weights and server overhead not included. The slider stops at 128k tokens, the longest context these models support.

    Prefill and decode

    Increase the context and the number of conversations in the batch. See when reading the cache starts to outweigh reading the weights

    used of the GPU’s 80 GB
    time per step
    tokens/s per conversation
    tokens/s in total
    of reads are the cache

    An illustrative model as in “Why the GPU is idle”: 8B parameters in FP16 (16 GB of weights, 128 KB of cache per token), a single H100 with 80 GB of memory and 3.35 TB/s. Step time = (weights + cache of all conversations) / bandwidth. This is an upper bound on speed, with no overheads.

    The cache on the server

    Check yourself

    What is the KV cache, and how does prefill differ from decode?

    Generation is a loop: one forward pass yields one token. Attention needs the keys and values of all earlier tokens, and thanks to the causal mask they never change, so they are computed once and kept in GPU memory. Prefill processes the whole prompt in parallel, is compute-bound and sets time to first token. Decode goes token by token and is memory-bound, because every step reads the weights and the whole cache. The cache is about 0.33 MB per token for a 70B model with GQA, so it is what caps context length and batch size. GQA, MLA, sliding windows and FP8 shrink it.

    Po polsku

    Generowanie to pętla: jeden przebieg modelu daje jeden token. Attention potrzebuje kluczy i wartości wszystkich wcześniejszych tokenów, a przez maskę przyczynową one się nie zmieniają, więc liczy się je raz i trzyma w pamięci GPU. Prefill liczy cały prompt równolegle, jest ograniczony obliczeniami i wyznacza czas do pierwszego tokena. Decode idzie po tokenie i czeka na pamięć, bo każdy krok czyta wagi i cały cache. Cache to ok. 0,33 MB na token dla 70B z GQA, więc to on ogranicza kontekst i batch. Zmniejszają go GQA, MLA, okno przesuwne i FP8.

    Follow-up questions (4)
    What determines time to first token, and what determines writing speed?
    Time to first token depends on the queue and on how much of the prompt has to be computed outside the cache hit, because that is prefill. Writing speed depends on the size of the weights, the context length, the number of conversations in the batch and memory bandwidth, because that is decode.
    Why does a long context slow down generation if each step computes only one token?
    Because every step reads the whole cache. For a 70B model at 128k tokens that is 43 GB per step, almost a third of the weights’ size, and decode is bound by memory bandwidth.
    vLLM runs out of memory on long contexts. What do you do?
    Work out the budget: weights plus cache per token × maximum context × number of concurrent sequences. Then cap the maximum length and the number of sequences, enable an FP8 cache and prefix sharing, and if that is not enough, split the model across more GPUs. Tensor parallelism splits a GQA cache by KV head, so across at most as many GPUs as there are KV heads (8 in Llama 3.1 70B). An MLA cache is replicated on every GPU, which is why DeepSeek-style models use data-parallel attention instead.
    How do you shrink the KV cache?
    At model design time: GQA or MQA, MLA, sliding windows, SSM layers. At serving time: FP8, prefix sharing, offloading to RAM, evicting tokens. Only sharing and offloading leave the output unchanged.

    Sources

    Report an error · Suggest a fix