## Training

*Where a model’s knowledge comes from*

*Last edited: 28 September 2026*

A model is built in three stages. Pretraining provides language and knowledge through next-token prediction, SFT teaches the assistant role, and RL polishes behaviour and reasoning. After training the weights are frozen, so anything the model doesn’t know has to be given to it in context.

**In plain words:** First, years of reading everything in sight. Then a vocational course on example conversations, after which it knows how an assistant behaves. Finally, graded practice: it tries on its own and an examiner rewards good results, so it learns whatever earns points.

*Interactive widget on the page: Click a stage and compare what the model does with the same prompt.*

### Pretraining

- The task: predict the next token across a huge corpus of text and code (Llama 3: about 15 trillion tokens). The loss is cross-entropy, i.e. −log of the probability the model gave the correct token. Backpropagation computes the gradient, and an optimiser, usually AdamW, nudges every weight slightly towards a smaller error. Every position in the text is a separate example, so one document teaches thousands of predictions at once.
- The cost is about 6 × parameters × tokens operations. Chinchilla (2022) found that for a fixed budget the optimum is about 20 tokens per parameter. Today models are trained much longer (Qwen3, 2025: 36 trillion tokens for every size, about 60,000 per parameter for the 0.6B model; Llama 3 8B: almost 2000 per parameter), because a smaller model trained for longer is cheaper at inference over its whole lifetime.
- Before training the data is filtered: Llama 3 removed duplicates at the URL, document (MinHash) and line level, dropped junk with heuristics and picked high-quality text with classifiers. Public benchmarks sit on the same web and leak into the corpus. On GSM1k, fresh problems in the style of GSM8K, some models scored up to 8% lower, so a public score can overstate what you will see on your own task (see “Evals”).
- Towards the end, the model is trained on progressively longer sequences to extend its context, and on a small portion of the best data, e.g. code and maths, with a decaying learning rate (annealing). This stage is also called mid-training.
- The result, the base model, is trained to continue documents, not to answer. Given “How do I bake bread?” it may add more questions from a forum. Pretraining data now includes plenty of question–answer and synthetic instruction text, so modern base models often answer anyway, but unreliably and without the assistant role.

### Post-training: SFT and RL

- SFT uses the same loss, but on conversations written in the chat template (see “How a model sees a chat”) and computed only on the assistant’s answer tokens. Hundreds of thousands to millions of examples, written by people and generated by stronger models, teach format, role and tool calling.
- RLHF: people compare pairs of answers, a reward model is trained on those comparisons, and the language model is optimised against its score (classically with PPO), with a KL penalty for drifting away from the SFT model. In InstructGPT (2022) a 1.3-billion-parameter model trained this way beat the 175-billion-parameter GPT-3 in human ratings. DPO learns directly from better/worse pairs, without a separate reward model. Today the comparisons are often made by a model judge following written principles (RLAIF, Constitutional AI); in Tülu 3, GPT-4o rated the answers.
- RLVR: the reward is checked by a program, such as tests or a comparison with the correct result. The model generates many attempts, and the better ones are reinforced. GRPO compares attempts on the same prompt with each other, without a separate critic. This is how reasoning models came about: a long chain of thought pays off because it raises the chance of a reward (see “Reasoning models”).
- RL in environments: the same idea for agents. Multi-turn tasks with tools (a terminal, a repository with tests, simulated APIs) are rewarded for the end state; Kimi K2 (2025) had a joint RL stage in real and synthetic environments. SFT teaches the tool-call format, and this is where the model practises long, multi-step runs.
- Distillation: a smaller model learns from the answers or probability distributions of a larger one, which is why small models beat what their size suggests (see “Distillation”).

### What this means in practice

- After training the weights are frozen. The model doesn’t learn from the conversation, and “memory” in apps is notes appended to the context (see “Context engineering and memory”). About the world after its cutoff date it knows only what it gets in context.
- RL optimises the reward, not the intent (reward hacking). A model rewarded for passing tests can learn to game them instead of fixing the code, and one rewarded by human ratings learns to agree with people (sycophancy). In RLHF a KL penalty and diverse rewards limit this but don’t eliminate it. RLVR for reasoning often drops the KL term (DAPO, 2025), because a long chain of thought has to move far from the starting model, so there the quality of the verifier is what holds reward hacking back.
- Your own training starts from an existing model. Fine-tuning is good at teaching style, format and narrow tasks, and poor at teaching new knowledge. Knowledge that changes or needs a source goes into the context (see “Fine-tuning and LoRA” and “RAG”).

### Check yourself

**Question:** How is an LLM made: what is the difference between pretraining, SFT, RLHF and RLVR?

**Short answer:** Pretraining teaches next-token prediction on trillions of tokens of text with a cross-entropy loss. That gives language and knowledge, but the base model only continues documents. SFT on conversations in the chat template teaches the assistant role and the tool format. RLHF optimises against a reward model learned from comparisons made by people or a model judge, with a KL penalty for drifting away from the SFT model. RLVR optimises against a reward checked by a program, such as tests: that is how reasoning models are made, and, in multi-turn tool environments, models for agents. The risk of RL is reward hacking. After training the weights are frozen, so current knowledge has to be supplied in context.

### Follow-up questions

- **RLHF versus RLVR?** RLHF takes its reward from a model trained on the preferences of people or a model judge: good for style and helpfulness, but the reward model can be fooled. RLVR takes its reward from an automatic check, such as tests or the task’s result: harder to fool and easy to scale, but only where a program can verify the result.
- **PPO, DPO, GRPO?** PPO is classic RL with a reward model and a separate critic that estimates the value of a state. DPO learns directly from “better/worse answer” pairs, with no reward model and no generation in the loop. GRPO (DeepSeek) generates several answers to the same prompt and compares them with each other, so it needs no critic.
- **Why a KL penalty in RL?** It keeps the model close to the SFT model. Without it, optimisation quickly finds answers that the reward model rates highly and people rate poorly, and the model loses general skills. RLVR for reasoning often drops it (DAPO), because the model has to move far from its starting point; there a good verifier is what stops the model gaming the reward.
- **When fine-tuning and when RAG?** Fine-tuning for style, format and narrow skills. RAG for knowledge that changes and needs a source (see “Fine-tuning and LoRA” and “RAG”).

### Sources

- [Ouyang et al.: Training language models to follow instructions with human feedback (InstructGPT, 2022)](https://arxiv.org/abs/2203.02155)
- [Lambert et al.: Tülu 3, an open post-training recipe with RLVR (2024)](https://arxiv.org/abs/2411.15124)
- [Llama Team: The Llama 3 Herd of Models (2024), data and decontamination](https://arxiv.org/abs/2407.21783)
- [Kimi Team: Kimi K2: Open Agentic Intelligence (2025)](https://arxiv.org/abs/2507.20534)
- [Andrej Karpathy: Deep Dive into LLMs like ChatGPT](https://www.youtube.com/watch?v=7xTGNNLPyMI)

Interactive page: https://howaiworks.dev/training/
