Home / Chapter 2 · Where a model’s knowledge comes from
    Last edited · 8 min read

    Use with AI

    How a model sees a chat

    The model never sees a chat window. The system prompt, messages, tool calls and tool results are concatenated into one document with role markers, and the model continues the text after the assistant marker. This explains the cost of long conversations, the limits of the system prompt and most bugs in self-hosted deployments.

    In plain wordsA script where every line has a speaker label: DIRECTOR, CUSTOMER, ASSISTANT. An actor with no memory gets the whole script from page one every time, sees an empty ASSISTANT label at the end and writes the next line. If someone slips a stranger’s page into the script, the actor reads it just like the rest.

    Add messages and tool calls, and switch views. Notice that every request resends the whole document from the start

    tokens in the document now
    API requests since the conversation began
    input tokens sent in total

    A simplified, illustrative template. Every model family has its own markers and its own way of writing tool calls. Tokens are counted with the simplified tokeniser from “Tokens”, and each marker counts as one token.

    What your code actually sends: a list of messages, not a document

    A simplified form of what most chat APIs accept. Details such as the tool-call format differ between providers. You don’t send the markers or the final <|assistant|>; the server adds them. Tool definitions go in a separate field, and the server inserts them into the document next to the system prompt.

    The chat template

    The whole document every turn

    Images, PDFs and voice

    The system prompt is not a safeguard

    Self-hosting and fine-tuning

    Check yourself

    How does a model see a conversation, and how does it know who said what?

    The model never sees a chat window. The server concatenates the system prompt, history, tool calls and tool results into one document, separating roles with special tokens from a template the model learned during SFT. It ends with an assistant marker, and the model continues the text until it emits an end-of-turn token. The model has no memory, so the whole document is resent every turn. The system prompt is just text at the top, not a security boundary: an instruction hidden in an email sits in the same stream.

    Po polsku

    Model nie widzi okienka czatu. Serwer skleja system prompt, historię wiadomości, wywołania narzędzi i ich wyniki w jeden dokument, a role oddziela tokenami specjalnymi według szablonu, którego model nauczył się w SFT. Na końcu stawia znacznik asystenta, a model dopisuje ciąg dalszy, aż wygeneruje token końca tury. Model nie ma pamięci, więc w każdej turze dostaje cały dokument od nowa. System prompt to tylko tekst na początku, a nie granica bezpieczeństwa: polecenie ukryte w mailu leży w tym samym ciągu.

    Follow-up questions (5)
    Why does the model sometimes carry on writing as the user?
    Because it didn’t generate the end-of-turn token, or the server doesn’t have that token in its stop list. To the model it is still one document, so the most likely continuation is a user marker followed by the user’s next message. A wrong template or stop configuration is usually to blame.
    Can the model tell an instruction in the system prompt from one in an email?
    Only as far as it learned to in training. Models are trained on a hierarchy: system above user, user above tool content (OpenAI, “The Instruction Hierarchy”, 2024). This improves robustness but is not a hard boundary, because everything sits in one token stream.
    Why does the API take a list of messages instead of ready-made text?
    So that the template, markers and tool format always match what the model was trained on, and so that user content can’t forge a role marker.
    You are deploying an open-weight model on your own server. What do you check in the template?
    That the tokeniser’s template matches the model card, that the end-of-turn token is in the stop list, that there is no doubled beginning-of-text token, and that user text can’t turn into special tokens. The simplest check is to render a sample conversation with tools and compare it character by character with the example in the model’s documentation.
    What happens when you edit an old message?
    The model remembered nothing, so it simply gets a new document. You lose the prompt cache from the point of the change to the end. On Claude Opus 5.5 and Fable 5.1, editing history before a thinking block ends in a 400 error (see “Reasoning models”).

    Sources

    Report an error · Suggest a fix