Home / Chapter 7 · Production and serving
    Last edited · 6 min read

    Use with AI

    Speculative decoding

    A cheap mechanism guesses the next few tokens, and the large model checks all of them in one pass, keeping the ones it would have chosen itself. The text is the same as without speculation, but one read of the weights yields several tokens.

    In plain wordsAn intern writes a draft; the boss reads it and corrects the first wrong word, then the intern carries on writing. Reading is much faster than writing, so together they finish sooner, and the text is exactly what the boss would have written alone.

    Step through three rounds and watch how many tokens one verification accepts and how many large-model passes you save

    small model’s proposalacceptedrejectedlarge model’s token
    tokens generated
    large-model passes
    passes without speculation
    of the time without speculation

    Proposals and hits are hand-picked. A small-model pass counts as 1/10 of a large-model pass, and a pass that verifies several tokens costs the same as a normal one.

    How verification works

    What the gain depends on

    Variants and choosing in practice

    Check yourself

    How does speculative decoding work, why doesn’t it change the output, and when does it stop helping?

    A cheap drafter, such as a small model, extra heads or n-grams copied from the prompt, proposes several tokens. The large model checks them in one pass, because decode is memory-bound and several positions cost about as much as one. It keeps the matching prefix and inserts its own token at the first mismatch. Accepting with probability min(1, p/q) guarantees exactly the large model’s distribution. The gain depends on the acceptance rate: 2–3× on predictable text at small batch sizes, less as the batch grows, and nothing or a loss once the step is compute-bound (large batch, short contexts).

    Po polsku

    Tani draft, czyli mały model, dodatkowe głowice albo n-gramy skopiowane z promptu, proponuje kilka tokenów. Duży model sprawdza je w jednym przebiegu, bo decode jest ograniczony pamięcią i kilka pozycji kosztuje prawie tyle co jedna. Przyjmuje zgodny prefiks, a w miejscu pierwszej niezgodności wstawia własny token. Akceptacja z prawdopodobieństwem min(1, p/q) gwarantuje dokładnie rozkład dużego modelu. Zysk zależy od trafności: 2–3× przy przewidywalnym tekście i małym batchu, mniej przy większym, a zero albo strata, gdy krok jest już ograniczony mocą obliczeniową (duży batch, krótkie konteksty).

    Follow-up questions (4)
    When does speculation not help, or even hurt?
    When acceptance is low: creative text, high temperature, or a draft trained on different data. Also when the step is already compute-bound, i.e. large batches with short contexts: verifying rejected tokens then takes compute away from other conversations.
    Why is the output distribution exactly the large model’s?
    Token x is accepted with probability min(1, p(x)/q(x)), and after a rejection you sample from max(0, p − q), normalised. The two paths together give exactly p(x) for every token. In practice only small numerical differences remain, from the different shape of the computation.
    Acceptance is 0.6 with γ = 5. Do you increase γ?
    If anything, decrease it. A round yields on average (1 − 0.6⁶) / 0.4 ≈ 2.4 tokens, and with γ = 3 already ≈ 2.2, so the two extra proposals add almost nothing but cost drafting and verification. A better draft gives a bigger gain, for example one fine-tuned on your own traffic.
    Does the draft have to use the same tokeniser as the large model?
    In the classic variant, yes, because you compare tokens and distributions over the same vocabulary. That is why the draft comes from the same model family, or the heads are attached to the large model itself. Newer servers relax this: vLLM can pair a draft and a large model with different vocabularies through a token mapping (use_heterogeneous_vocab, as of September 2026 with greedy drafting only).

    Sources

    Report an error · Suggest a fix