Home / Chapter 1 · How a model reads and predicts
Last edited · 6 min read
The next token
A model doesn’t return text, only a score (logit) for every token in the vocabulary. Separate code, the sampler, turns the scores into probabilities and draws one token. Every token of an answer is produced this way, one after another.
In plain wordsYour phone keyboard suggests three words. A model does the same for a hundred thousand pieces at once, then rolls a die weighted by those odds. Temperature decides how heavily the die is loaded.
Change the temperature and top-p, then sample. Compare a question the model knows, one it gets confidently wrong and one where it is guessing
Real next-token probabilities from Qwen3-0.6B-Base, a small open model without chat training: its eight most likely tokens, rescaled to 100%. Struck-through bars have been cut by top-p.
From logits to a token
- Softmax with temperature:
p = exp(logit / T) / Σ exp(logit / T). T below 1 sharpens the distribution, above 1 flattens it, T → 0 always picks the top token (greedy), and a very high T pushes the distribution towards uniform. Temperature doesn’t change the order of the tokens, only the gaps between them. - Cutting the tail: top-k keeps the k best tokens, top-p (nucleus) the smallest set covering e.g. 90% of the probability, and min-p drops tokens weaker than a fraction of the best one. Top-p and min-p adapt to the model’s confidence: when it is sure, one candidate remains; when it is unsure, dozens. In Hugging Face and vLLM, temperature is applied before truncation, so it affects what survives.
- The sampled token is appended to the text and can’t be taken back. Every following token has to fit it, so one unlucky token from the tail can drag a whole made-up sentence behind it.
- The sampler can also zero out tokens that don’t fit a JSON schema. That is how structured output works (see “Enforcing output format”).
Pitfalls
- Greedy doesn’t mean best. Always picking the most likely token leads to repetition and loops (Holtzman et al., 2019). DeepSeek recommends a temperature of 0.5–0.7 for R1 precisely to prevent endless repetition. Repetition, frequency and presence penalties lower the logits of tokens already used: they curb loops, but also punish legitimate repeats such as identifiers in code or JSON keys.
- Temperature 0 doesn’t give full reproducibility. GPU computations produce slightly different numbers depending on how many requests landed in the batch, and when logits are nearly tied, that changes the chosen token and the entire rest of the answer. The
seedparameter, where an API has one, doesn’t guarantee determinism. When self-hosting, batch-invariant kernels make the output deterministic at some cost in speed (vLLM:VLLM_BATCH_INVARIANT=1). - The distribution doesn’t measure truth. A flat distribution can signal ignorance, but also several equally good phrasings. A model can be confident in an error, and the sampler always picks something, so the sentence sounds just as confident. A lower temperature doesn’t cure fabrication; it only makes it repeatable (see “Hallucinations”).
What you set in practice
- In non-reasoning models: extraction, classification and code at a low temperature, 0–0.3; writing and brainstorming at 0.7–1. Change temperature or top-p, not both at once, and check the effect on an eval set, not on a single sample.
logprobsin an API show this distribution, usually before temperature and top-p (the default in vLLM). They are useful for a confidence threshold in classification: you have the model answer with a single label token and read the probability of each label. The Claude API doesn’t return them, and OpenAI does only with reasoning off.- Several samples at T > 0 plus a majority vote (self-consistency) raise accuracy on tasks with one correct answer, and disagreement between the samples signals uncertainty. It costs as many times more tokens as there are samples.
- With reasoning models the sampler is often locked (as of September 2026). Claude Opus 4.7 and later, and Claude Sonnet 5, reject non-default
temperature,top_pandtop_kwith a 400 error, and OpenAI with reasoning enabled doesn’t accepttemperature,top_porlogprobs. You steer with the reasoning level and the prompt instead. Where a reasoning model does accept them, keep the provider’s recommended values: Google strongly recommends temperature 1.0 for all Gemini 3 models and warns that lower values can cause looping. A low temperature is a tool for non-reasoning models.
Check yourself
How does a model choose the next token, and how do you set temperature and top-p?
The model returns logits for the whole vocabulary. A temperature softmax, exp(logit/T), turns them into a distribution: T below 1 sharpens it, above 1 flattens it, and T = 0 always picks the top token. Top-p keeps the smallest set of tokens covering, say, 90% of the probability and cuts the tail where odd tokens come from. On non-reasoning models, use a low temperature for extraction and code and a higher one for creative text, change one parameter at a time and check on evals. T = 0 is not fully reproducible, and reasoning models often lock the sampler or, like Gemini 3, work best at the default 1.0.
Po polsku
Model zwraca logity dla całego słownika. Softmax z temperaturą, exp(logit/T), zamienia je w rozkład: T poniżej 1 wyostrza, powyżej 1 spłaszcza, a T = 0 to zawsze najlepszy token. Top-p zostawia najmniejszy zbiór tokenów pokrywający np. 90% prawdopodobieństwa i odcina ogon, z którego biorą się dziwne tokeny. W modelach bez rozumowania ekstrakcja i kod dostają niską temperaturę, teksty twórcze wyższą; zmieniaj jeden parametr naraz i sprawdzaj na ewaluacjach. T = 0 nie daje pełnej powtarzalności, a modele rozumujące często blokują sampler albo, jak Gemini 3, działają najlepiej przy domyślnym 1,0.
Follow-up questions (4)
- Why doesn’t temperature 0 give full reproducibility?
- The result depends on which other requests landed in the same batch: a different batch size means a different order of floating-point operations and slightly different logits. With two nearly tied tokens that is enough to pick the other one, and from there the whole answer diverges.
- Top-k, top-p, min-p: what’s the difference?
- Top-k keeps a fixed number of candidates regardless of how confident the model is. Top-p keeps as many as it takes to cover a given share of the probability, so it adapts to the distribution. Min-p drops tokens weaker than a fraction of the best one, e.g. 0.1 × its probability. Its authors report that it copes better with high temperatures, but a 2025 reanalysis found their evidence doesn’t support that.
- Does a lower temperature reduce hallucinations?
- Only the ones that come from drawing a token from the tail. If the model “believes” a wrong fact, T = 0 will pick it every time (see “Hallucinations”).
- How do you use logprobs for classification?
- Have the model answer with a single label token and read the probability of each label. You get a distribution instead of a single answer, and you set a threshold below which the case goes to a human. Calibrate the threshold on your own data, because after post-training the model’s confidence is often inflated.