Home / Chapter 7 · Production and serving
Last edited · 6 min read
Speculative decoding
A cheap mechanism guesses the next few tokens, and the large model checks all of them in one pass, keeping the ones it would have chosen itself. The text is the same as without speculation, but one read of the weights yields several tokens.
In plain wordsAn intern writes a draft; the boss reads it and corrects the first wrong word, then the intern carries on writing. Reading is much faster than writing, so together they finish sooner, and the text is exactly what the boss would have written alone.
Step through three rounds and watch how many tokens one verification accepts and how many large-model passes you save
Proposals and hits are hand-picked. A small-model pass counts as 1/10 of a large-model pass, and a pass that verifies several tokens costs the same as a normal one.
How verification works
- The draft, i.e. the cheap guessing mechanism, proposes γ tokens (usually a few). In one pass the large model computes the distribution at each of those positions, just as in prefill. It keeps the longest matching prefix, inserts its own token at the first mismatch, and adds one more token if everything matches. So every round yields at least one token.
- Verification is cheap because decode is memory-bound: a pass over five positions reads the same weights as a pass over one and takes almost the same time (see “Why the GPU is idle”).
- When sampling, a proposal x is accepted with probability min(1, p(x)/q(x)), where p is the large model’s distribution and q the small model’s. After a rejection you sample from max(0, p − q), normalised. The result has exactly the large model’s distribution (Leviathan et al.; Chen et al., 2023), and with greedy decoding the text is identical, up to small numerical differences.
What the gain depends on
- With acceptance rate α and γ proposals, a round yields on average (1 − α^(γ+1)) / (1 − α) tokens. For α = 0.8 and γ = 4 that is about 3.4 tokens per large-model pass. The time saving is smaller, because the draft costs something too, and rejected proposals are wasted work.
- Acceptance depends on the text. Code, JSON, repetition and low temperature give a high α; creative text at high temperature a low one. Leviathan et al. measured a 2–3× speed-up on T5-XXL. DeepSeek-V3 reports 85–90% acceptance of the second token from its MTP module and 1.8× more tokens per second.
- The gain shrinks as the batch grows. The weights are read once for the whole batch anyway, so once a step becomes compute-bound the extra positions cost real compute: with large batches and short contexts, speculation can turn into a loss. The KV cache is different, because every conversation reads its own. With long contexts decode stays memory-bound even at large batch sizes, and a draft that is cheap on cache reads still raises throughput (MagicDec: up to 2.5× for Llama 3.1 8B at batch sizes 32–256; EAGLE-3: 1.38× at batch 64 in SGLang). vLLM recommends speculation mainly for low to medium load, but rates EAGLE and MTP as a medium to high gain under heavy load too.
Variants and choosing in practice
- A separate small model from the same family with the same tokeniser: easy to use, but it takes memory and has to be served alongside the large one.
- A draft attached to the large model: extra heads that predict the next positions (Medusa), a light layer working on the large model’s hidden states (EAGLE, now EAGLE-3), or an MTP module the model has from training, as in DeepSeek-V3. These usually guess better than a separate small model. The output stays exact only with standard acceptance and an unchanged large model: Medusa-2, which fine-tunes the large model along with the heads, and Medusa’s “typical acceptance” trade exactness for speed.
- No model at all: n-gram matching copies the continuation from the prompt (prompt lookup). It works for code editing, RAG and summaries, where the answer repeats parts of the input.
- Choose the method and γ from the α measured on your own traffic, and check the gain at the target load, not on a single request. In providers’ APIs speculation usually runs invisibly. The exception is features like Predicted Outputs in the OpenAI API, where you supply the expected text yourself, for example the file before an edit (as of September 2026 only on GPT-4o and GPT-4.1 models and without function calling; rejected predicted tokens are billed as output).
Check yourself
How does speculative decoding work, why doesn’t it change the output, and when does it stop helping?
A cheap drafter, such as a small model, extra heads or n-grams copied from the prompt, proposes several tokens. The large model checks them in one pass, because decode is memory-bound and several positions cost about as much as one. It keeps the matching prefix and inserts its own token at the first mismatch. Accepting with probability min(1, p/q) guarantees exactly the large model’s distribution. The gain depends on the acceptance rate: 2–3× on predictable text at small batch sizes, less as the batch grows, and nothing or a loss once the step is compute-bound (large batch, short contexts).
Po polsku
Tani draft, czyli mały model, dodatkowe głowice albo n-gramy skopiowane z promptu, proponuje kilka tokenów. Duży model sprawdza je w jednym przebiegu, bo decode jest ograniczony pamięcią i kilka pozycji kosztuje prawie tyle co jedna. Przyjmuje zgodny prefiks, a w miejscu pierwszej niezgodności wstawia własny token. Akceptacja z prawdopodobieństwem min(1, p/q) gwarantuje dokładnie rozkład dużego modelu. Zysk zależy od trafności: 2–3× przy przewidywalnym tekście i małym batchu, mniej przy większym, a zero albo strata, gdy krok jest już ograniczony mocą obliczeniową (duży batch, krótkie konteksty).
Follow-up questions (4)
- When does speculation not help, or even hurt?
- When acceptance is low: creative text, high temperature, or a draft trained on different data. Also when the step is already compute-bound, i.e. large batches with short contexts: verifying rejected tokens then takes compute away from other conversations.
- Why is the output distribution exactly the large model’s?
- Token x is accepted with probability min(1, p(x)/q(x)), and after a rejection you sample from max(0, p − q), normalised. The two paths together give exactly p(x) for every token. In practice only small numerical differences remain, from the different shape of the computation.
- Acceptance is 0.6 with γ = 5. Do you increase γ?
- If anything, decrease it. A round yields on average (1 − 0.6⁶) / 0.4 ≈ 2.4 tokens, and with γ = 3 already ≈ 2.2, so the two extra proposals add almost nothing but cost drafting and verification. A better draft gives a bigger gain, for example one fine-tuned on your own traffic.
- Does the draft have to use the same tokeniser as the large model?
- In the classic variant, yes, because you compare tokens and distributions over the same vocabulary. That is why the draft comes from the same model family, or the heads are attached to the large model itself. Newer servers relax this: vLLM can pair a draft and a large model with different vocabularies through a token mapping (use_heterogeneous_vocab, as of September 2026 with greedy drafting only).