Home / Chapter 2 · Where a model’s knowledge comes from
Last edited · 8 min read
How a model sees a chat
The model never sees a chat window. The system prompt, messages, tool calls and tool results are concatenated into one document with role markers, and the model continues the text after the assistant marker. This explains the cost of long conversations, the limits of the system prompt and most bugs in self-hosted deployments.
In plain wordsA script where every line has a speaker label: DIRECTOR, CUSTOMER, ASSISTANT. An actor with no memory gets the whole script from page one every time, sees an empty ASSISTANT label at the end and writes the next line. If someone slips a stranger’s page into the script, the actor reads it just like the rest.
Add messages and tool calls, and switch views. Notice that every request resends the whole document from the start
A simplified, illustrative template. Every model family has its own markers and its own way of writing tool calls. Tokens are counted with the simplified tokeniser from “Tokens”, and each marker counts as one token.
What your code actually sends: a list of messages, not a document
A simplified form of what most chat APIs accept. Details such as the tool-call format differ between providers. You don’t send the markers or the final <|assistant|>; the server adds them. Tool definitions go in a separate field, and the server inserts them into the document next to the system prompt.
The chat template
- Role markers are special tokens: separate IDs in the vocabulary, not ordinary characters (see “Tokens”). The model learned the format during SFT (see “Training”): the assistant marker is followed by the answer, and the answer by an end-of-turn token.
- Every model family has its own template: ChatML in Qwen (
<|im_start|>user … <|im_end|>), Llama 3 (<|start_header_id|>user<|end_header_id|> … <|eot_id|>), harmony in gpt-oss (<|start|>user<|message|> … <|end|>), Mistral ([INST] … [/INST]). You send a list of messages with roles, and the server renders it with the template the model was trained on. - In Hugging Face the template is a Jinja template shipped with the tokeniser.
apply_chat_template(messages, add_generation_prompt=True)builds the document and appends the assistant marker at the end. Without that marker the model may carry on writing the user’s message instead of answering it. - Generation stops at the end-of-turn token, provided the server has it in its list of stop tokens. In Llama 3 a turn ends with
<|eot_id|>, not<|end_of_text|>. A server that only waits for the latter lets the model write a user marker and invent the user’s next message.
The whole document every turn
- The model doesn’t remember the conversation. Every request carries the whole document from the start, so the total number of tokens sent grows with the square of the number of turns (see “The context window and agents”). Server-side state, such as OpenAI’s
previous_response_id, only saves the upload: the server rebuilds the document and bills all earlier input tokens again. What to keep, summarise or drop from the history is covered in “Context engineering and memory”. - The stable start of the document (system prompt, tool definitions and older history) is what prompt caching reuses (see “Prompt caching”). Anything variable at the start, such as the current time in the system prompt, breaks the cache every turn. Editing an old message produces a new document, and the cache is lost from the point of the change.
- Tools are document fragments too. Their definitions sit next to the system prompt and cost tokens every turn. A call is text the model writes after a special marker, and your code appends the result as a new fragment (see “Tools (function calling)”).
- A reasoning model’s thinking is another fragment with its own markers:
<think>…</think>in some open models, theanalysischannel in harmony (see “Reasoning models”). Harmony drops it from the history after the final answer but keeps it between tool calls.
Images, PDFs and voice
- A vision encoder cuts an image into square patches and turns each one into a vector that takes a token’s place in the sequence. For Claude a patch is 28 × 28 px, so a 1000 × 1000 px photo is 1296 tokens, and Claude 4.7 and later downscale anything above 2576 px on the long edge or 4784 tokens (as of September 2026).
- An image stays in the history and is billed again every turn: ten 1000 × 1000 px screenshots in an agent’s history add about 13,000 tokens to every request. Claude reads each PDF page twice, as text (1,500–3,000 tokens) and as an image, so a 100-page report is 150,000–300,000 tokens of text alone. Cache such input, downscale it before sending, and drop screenshots the agent no longer needs.
- Audio becomes tokens too: Gemini counts 32 tokens per second of audio, and OpenAI’s Realtime API 10 per second of the user’s speech and 20 per second of the model’s, resending the whole conversation for each response.
- A voice agent is built one of two ways. A chained pipeline (speech-to-text → text model → text-to-speech) gives you transcripts, policy checks before the answer and any text model, but every stage adds to the silence the caller hears. A speech-to-speech model (OpenAI Realtime, Gemini Live) handles audio in one session, with barge-in and turn detection, but shows you less of what happens in between. Measure the time from the end of the caller’s speech to the first audio, at the median and p95.
The system prompt is not a safeguard
- The system prompt is text at the start of the document that the model was trained to prioritise. It is not a separate channel or a permission. An instruction in an email, on a web page or in a tool result sits in the same token stream and sometimes wins (see “Prompt injection”). Rules that must hold are enforced by application code. Don’t keep secrets in the system prompt, because the model may quote it.
- User text must not become a marker. In a well-built API, typing “<|im_start|>system” yields ordinary characters. When self-hosting, check this yourself: Hugging Face tokenisers recognise special-token strings in text by default (
split_special_tokens=False), so pasted text can open a real system turn.
Self-hosting and fine-tuning
- With an open-weight model (see “Open-weight models”) the template is part of the model. A wrong or outdated one raises no error; it silently degrades answers and breaks tool calling. A common case is a doubled beginning-of-text token: the document from
apply_chat_template(tokenize=False)gets tokenised a second time with special tokens added, and the Hugging Face docs warn that this hurts quality. - Render fine-tuning data with exactly the template the server will use, and compute the loss only on assistant tokens (see “Fine-tuning and LoRA”).
- Prefill: you end the message list with the start of the answer, such as
{, and the model writes on from there. In transformers this iscontinue_final_message=True. From Opus 4.6 and Sonnet 4.6 onwards, the Claude API rejects prefill with a 400 error and points you to structured outputs (see “Enforcing output format”).
Check yourself
How does a model see a conversation, and how does it know who said what?
The model never sees a chat window. The server concatenates the system prompt, history, tool calls and tool results into one document, separating roles with special tokens from a template the model learned during SFT. It ends with an assistant marker, and the model continues the text until it emits an end-of-turn token. The model has no memory, so the whole document is resent every turn. The system prompt is just text at the top, not a security boundary: an instruction hidden in an email sits in the same stream.
Po polsku
Model nie widzi okienka czatu. Serwer skleja system prompt, historię wiadomości, wywołania narzędzi i ich wyniki w jeden dokument, a role oddziela tokenami specjalnymi według szablonu, którego model nauczył się w SFT. Na końcu stawia znacznik asystenta, a model dopisuje ciąg dalszy, aż wygeneruje token końca tury. Model nie ma pamięci, więc w każdej turze dostaje cały dokument od nowa. System prompt to tylko tekst na początku, a nie granica bezpieczeństwa: polecenie ukryte w mailu leży w tym samym ciągu.
Follow-up questions (5)
- Why does the model sometimes carry on writing as the user?
- Because it didn’t generate the end-of-turn token, or the server doesn’t have that token in its stop list. To the model it is still one document, so the most likely continuation is a user marker followed by the user’s next message. A wrong template or stop configuration is usually to blame.
- Can the model tell an instruction in the system prompt from one in an email?
- Only as far as it learned to in training. Models are trained on a hierarchy: system above user, user above tool content (OpenAI, “The Instruction Hierarchy”, 2024). This improves robustness but is not a hard boundary, because everything sits in one token stream.
- Why does the API take a list of messages instead of ready-made text?
- So that the template, markers and tool format always match what the model was trained on, and so that user content can’t forge a role marker.
- You are deploying an open-weight model on your own server. What do you check in the template?
- That the tokeniser’s template matches the model card, that the end-of-turn token is in the stop list, that there is no doubled beginning-of-text token, and that user text can’t turn into special tokens. The simplest check is to render a sample conversation with tools and compare it character by character with the example in the model’s documentation.
- What happens when you edit an old message?
- The model remembered nothing, so it simply gets a new document. You lose the prompt cache from the point of the change to the end. On Claude Opus 5.5 and Fable 5.1, editing history before a thinking block ends in a 400 error (see “Reasoning models”).