Home / Chapter 5 · Agents
    Last edited · 6 min read

    Use with AI

    Enforcing output format

    Structured output in strict mode doesn’t ask the model for a format; it enforces it: at every step the sampler zeroes the probability of tokens that would break the schema. You get a guarantee of syntax, not of correct content.

    In plain wordsA form with checkboxes instead of a blank sheet. You can’t write anything outside the allowed options, but you can still tick the wrong box. And when there is no box for the right answer, you tick the nearest one.

    Step through generation. See which tokens the mask cuts, and stop at token 4

    Schema: {"label": "spam" | "ham"}

    Illustrative probabilities. Struck-through tokens break the schema and get zero; the odds of the rest are rescaled to 100%.

    How masking works

    Modes in the API

    What the guarantee doesn’t cover

    Check yourself

    How do you guarantee that a model returns JSON matching a schema, and what does that guarantee not cover?

    In strict mode the schema is compiled into a grammar, and at every step the sampler zeroes out tokens that would break it. The JSON matches the schema unless the length limit cuts it off or the model refuses, which the stop reason shows. At Anthropic, enum values can also come back in a different letter case, so compare them case-insensitively. The guarantee covers syntax, not content: values are validated in code. The schema gets an escape hatch, because a tight enum pushes the model into the nearest option, and the reasoning field goes before the decision, because the model writes left to right. Strict tool calling works the same way.

    Po polsku

    W trybie ścisłym schemat kompiluje się do gramatyki, a w każdym kroku sampler zeruje prawdopodobieństwo tokenów, które by ją złamały. JSON pasuje do schematu, chyba że odpowiedź utnie limit długości albo model odmówi, co widać po powodzie zatrzymania. U Anthropic wartości enum mogą też wrócić z inną wielkością liter, więc porównuje się je bez jej rozróżniania. Gwarancja dotyczy składni, nie treści: wartości waliduje kod. Schemat dostaje wyjście awaryjne, bo ciasny enum wpycha model w najbliższą opcję, a pole z uzasadnieniem stoi przed decyzją, bo model pisze od lewej do prawej. Tool calling w trybie strict działa tak samo.

    Follow-up questions (4)
    Can enforcing a format lower quality?
    Yes, when the schema asks for the decision before the reasoning or has no “other” option. What helps: a reasoning field before the decision, a reasoning model’s thinking, which the mask doesn’t cover, or two steps: a free-form answer first, then extraction.
    JSON mode, structured output or tool calling: when do you use which?
    Structured output when the answer feeds your code. Tool calling when the model has to choose an action from several, with strict mode for the arguments. JSON mode only where there is no strict mode.
    How do you do this on your own model?
    vLLM and SGLang have built-in constrained decoding engines (including XGrammar and llguidance): the schema turns into an automaton, and the automaton into a token mask at every step. The per-token overhead is small; the cost is compiling each new schema. With a reasoning model, start the server with a reasoning parser (e.g. --reasoning-parser in vLLM); otherwise the grammar applies from the first token and cuts off the thinking.
    The JSON parses, but the data is wrong. What next?
    Validation in code (constrained types, business rules) and a retry with a specific error message, with a cap on attempts. A recurring error is a signal to change the schema or the prompt, and the case goes into the eval set.

    Sources

    Report an error · Suggest a fix