Home / Chapter 1 · How a model reads and predicts
    Last edited · 6 min read

    Use with AI

    Tokens

    A model never sees letters or words, only a sequence of integers. Each integer is the ID of a piece of text from a fixed vocabulary. Price, context limits and speed are all counted in tokens, and the way text gets split explains several classic model mistakes.

    In plain wordsA model reads like someone who knows syllables and whole common words but has never seen individual letters. Familiar words it takes in whole; rare ones it assembles from pieces.

    Pick a tokeniser and an example, or type your own text. See where the token boundaries fall and how many characters one token covers

    The same sentence in English and Polish

    The dot · marks a space. A piece like \xc5 is a single byte: the tokeniser split one letter across two tokens. Tiktokenizer (linked in the sources) counts tokens for any text.

    How the vocabulary is built

    Effects you see in answers

    Cost and limits

    Check yourself

    How does BPE tokenisation work, and how does it affect a model’s cost and behaviour?

    A tokeniser splits text into pieces from a fixed vocabulary and maps them to integers. BPE builds the vocabulary from 256 bytes by repeatedly merging the most frequent adjacent pair until it reaches the target size, usually 100–260k. Because it works on bytes, any text can be encoded. Common words become one token; rare and non-English words split into pieces, so Polish takes about 1.3–1.6 times as many tokens as English, depending on the tokeniser. The model never sees individual letters or digits, hence mistakes in letter counting and arithmetic. A larger vocabulary shortens sequences but grows the embedding table and the output layer.

    Po polsku

    Tokenizer tnie tekst na kawałki ze stałego słownika i zamienia je na liczby. BPE buduje słownik od 256 bajtów, scalając najczęstsze sąsiednie pary, aż dojdzie do zadanego rozmiaru, zwykle 100–260 tys. Dzięki bajtom każdy tekst da się zakodować. Częste słowa są jednym tokenem, rzadkie i nieangielskie rozpadają się na kawałki, więc polski tekst zajmuje ok. 1,3–1,6 raza więcej tokenów niż angielski, zależnie od tokenizera. Model nie widzi pojedynczych liter ani cyfr, stąd błędy przy liczeniu liter i arytmetyce. Większy słownik skraca sekwencje, ale powiększa embeddingi i warstwę wyjściową.

    Follow-up questions (4)
    Why not tokenise by character or by whole word?
    Characters make sequences several times longer, and the cost of attention and the KV cache grows with length. Whole words need a gigantic vocabulary, and a typo, a new word or a proper name gets no ID. Byte-level BPE combines short sequences with full coverage.
    What does a larger vocabulary change?
    Shorter sequences, so cheaper context and faster generation, especially in languages other than English. The price is a bigger embedding table and output layer (about 1 billion parameters each in Llama 3 70B) and rare tokens that saw few examples in training.
    Can you swap the tokeniser of a trained model?
    In practice, no. The embedding table and the output layer are tied to specific IDs, so a new vocabulary needs at least retraining those layers, and usually continued pretraining. Adding a few special tokens during fine-tuning is possible, but their embeddings have to be trained (see “Fine-tuning and LoRA”).
    How would you estimate the cost of a new feature before launch?
    On a sample of real data: input and tool definitions with the model’s own tokeniser or the provider’s token-counting endpoint, and output from the usage field of a few real calls. Reasoning models bill thinking tokens as output even when you don’t see them, and no tokeniser can count those in advance (see “Reasoning models”). Price a repeated prefix at the cache rate (see “Prompt caching”), then multiply by the number of calls. Per-million-token prices can’t be compared directly across providers, because the same text is a different number of tokens for each of them.

    Sources

    Report an error · Suggest a fix