Home / Chapter 3 · How a model writes and what it costs
    Last edited · 6 min read

    Use with AI

    The context window and agents

    The model remembers nothing between requests; it knows only what it received in the current one. The context window is the token limit of a single request, shared by input and output. An agent resends the whole growing history every turn, so it pays for that history many times over, and answer quality drops as it grows.

    In plain wordsThe model is a consultant with amnesia. Before every meeting they get a folder and know only what is in it. The folder can only be so thick, and the agent adds pages to it at every step.

    Move the agent-turn slider. See what fills the window and how fast the cumulative input cost grows

    input tokens in this turn
    input tokens sent since the start
    input cost without caching
    with prompt caching
    cumulative input cost without cachingwith prompt caching

    Illustrative numbers. A 200k-token window, 32k of it reserved for the answer and thinking. Each turn adds 6k tokens of tool results and 300 tokens of answer. Rates as in “Prompt caching”: $3 per million input tokens, cache writes at 125%, reads at 10%. Output cost not included.

    What counts towards the limit

    Input limit and output limit

    Cost grows with every turn

    Quality drops with length

    Check yourself

    What counts towards the context window, and why does a long agent session get expensive and progressively worse?

    The model is stateless and only sees the current request. The window caps tokens per request and is shared by the system prompt, tool definitions, history with tool results and everything the model generates, thinking included. The output limit is separate and much smaller. An agent resends the whole history every turn, so cumulative input tokens grow quadratically with turns. Prompt caching lowers the price, not the shape of the curve. Quality degrades with length long before the window is full, so keep context lean and test at realistic lengths.

    Po polsku

    Model jest bezstanowy i widzi tylko bieżący request. Okno to limit tokenów na request, wspólny dla system promptu, definicji narzędzi, historii z wynikami narzędzi i tego, co model wygeneruje, łącznie z myśleniem. Limit samego wyjścia jest osobny i dużo mniejszy. Agent w każdej turze wysyła całą historię, więc łączne tokeny wejściowe rosną z kwadratem liczby tur. Prompt caching obniża cenę, ale nie kształt krzywej. Jakość spada z długością na długo przed końcem okna, więc kontekst warto trzymać krótki i testować na realnych długościach.

    Follow-up questions (4)
    What is context rot?
    Quality degrading with input length before the window runs out. It is stronger when the context holds similar but irrelevant passages and when the answer doesn’t repeat words from the question, so a needle-in-a-haystack test doesn’t reveal it.
    The model has a 1M-token window. Why not put the whole knowledge base in the prompt?
    Because every request pays for the whole million, at some providers at a higher rate above a threshold, time to first token grows, and quality drops with length. For a small, stable set with prompt caching it is reasonable; for a large or changing one, retrieval wins (see “RAG”).
    An agent overflows the window after 40 turns. What do you do?
    First measure what fills the window, usually old tool results and tool definitions. Then trim results at the source, clear old ones, move side tasks to sub-agents, and only as a last resort compact the history (see “Context engineering and memory”).
    Does an API that keeps history server-side lower the cost?
    No. Features such as previous_response_id in OpenAI’s Responses API save on transfer, but the whole history still goes to the model and counts as input. Only prompt caching or a shorter context lowers the cost.

    Sources

    Report an error · Suggest a fix