Home / Chapter 5 · Agents
Last edited · 6 min read
Enforcing output format
Structured output in strict mode doesn’t ask the model for a format; it enforces it: at every step the sampler zeroes the probability of tokens that would break the schema. You get a guarantee of syntax, not of correct content.
In plain wordsA form with checkboxes instead of a blank sheet. You can’t write anything outside the allowed options, but you can still tick the wrong box. And when there is no box for the right answer, you tick the nearest one.
Step through generation. See which tokens the mask cuts, and stop at token 4
Schema: {"label": "spam" | "ham"}
Illustrative probabilities. Struck-through tokens break the schema and get zero; the odds of the rest are rescaled to 100%.
How masking works
- The schema is compiled into an automaton. For regular expressions and non-recursive schemas a finite automaton is enough; recursive schemas and arbitrary grammars need a pushdown automaton. The automaton’s state says which characters may come next.
- Tokens don’t line up with JSON syntax: a single token might be
{"or":". So the engine (Outlines, XGrammar) works out, for each automaton state, which tokens in the vocabulary are allowed. The rest get a logit of minus infinity, and the sampler draws from what remains (see “The next token”). - The per-token overhead is close to zero, because most checks are precomputed. You pay on the first use of a schema: compilation adds latency, after which the grammar is cached (at Anthropic, for 24 hours since its last use).
Modes in the API
- JSON mode guarantees valid JSON, but not conformance to a schema. Strict mode guarantees the schema:
json_schemawithstrict: truein OpenAI,output_config.formatin Anthropic. - Tool calling is the same JSON in a different role: the model writes a tool name and arguments (see “Tools (function calling)”). Without strict mode the format is only learned, so the arguments usually, but not always, match the schema.
strict: trueon a tool definition turns on the same masking. In OpenAI’s Responses API tools are strict by default when the schema allows it and otherwise silently fall back to best effort (the response showsstrict: false), so set the flag explicitly. - Strict modes support a subset of JSON Schema. OpenAI requires every field to be listed in
required(an optional field is a type that allowsnull) and every object to haveadditionalProperties: false, with limits of 5,000 properties and 10 levels of nesting. Anthropic doesn’t support, among other things, recursive schemas or string length limits, and allows at most 20 strict tools per request, 24 optional parameters and 16 union-typed parameters across all strict schemas. Beyond that the API returns “Schema is too complex for compilation”. - The mask covers only the answer. A reasoning model’s thinking stays unconstrained, so the model can reason freely first and then fill in the schema (see “Reasoning models”).
What the guarantee doesn’t cover
- Content. The model can still write a wrong invoice number or a made-up date in a valid format. Business rules (ranges, totals, whether IDs exist) are checked by code, and the error goes back to the model with a specific message, with a cap on retries.
- A response cut off by the length limit (
stop_reason: "max_tokens",finish_reason: "length") or a refusal: OpenAI returns it in a separaterefusalfield, Anthropic asstop_reason: "refusal", and the content may not match the schema. Check the stop reason before parsing. - Letter case in enums at Anthropic. String
enumandconstvalues can come back with different capitalisation (“Conversation Topic 3” for “Conversation topic 3”), with a normal stop reason and no error, in JSON outputs and strict tool use alike. Compare them case-insensitively and avoid values that differ only in case. - What the model wanted to say. When the most likely answer lies outside the schema, the mask pushes the model into the nearest allowed one, like phishing turned into spam in the widget. Add an escape hatch: an “other” or “unknown” value and a comment field.
- Reasoning quality. The model writes left to right, so putting the decision before the justification makes it decide before it has worked anything out. Put the justification field before the decision field and list both in
required: Anthropic writes required properties before optional ones (OpenAI keeps the schema’s order and requires every field anyway). In the “Let Me Speak Freely?” study (2024), format restrictions lowered scores on reasoning tasks, partly for this very reason: on one task, GPT-3.5 in JSON mode put the answer before the reason every time. The result is contested. A re-run with matched prompts (.txt, “Say What You Mean”) found no drop, while other work measures a cost from the mask itself, mainly with small models and tight schemas (Reddy et al., 2026). When your evals show a drop, let the model answer freely and extract the structure with a second, cheap call.
Check yourself
How do you guarantee that a model returns JSON matching a schema, and what does that guarantee not cover?
In strict mode the schema is compiled into a grammar, and at every step the sampler zeroes out tokens that would break it. The JSON matches the schema unless the length limit cuts it off or the model refuses, which the stop reason shows. At Anthropic, enum values can also come back in a different letter case, so compare them case-insensitively. The guarantee covers syntax, not content: values are validated in code. The schema gets an escape hatch, because a tight enum pushes the model into the nearest option, and the reasoning field goes before the decision, because the model writes left to right. Strict tool calling works the same way.
Po polsku
W trybie ścisłym schemat kompiluje się do gramatyki, a w każdym kroku sampler zeruje prawdopodobieństwo tokenów, które by ją złamały. JSON pasuje do schematu, chyba że odpowiedź utnie limit długości albo model odmówi, co widać po powodzie zatrzymania. U Anthropic wartości enum mogą też wrócić z inną wielkością liter, więc porównuje się je bez jej rozróżniania. Gwarancja dotyczy składni, nie treści: wartości waliduje kod. Schemat dostaje wyjście awaryjne, bo ciasny enum wpycha model w najbliższą opcję, a pole z uzasadnieniem stoi przed decyzją, bo model pisze od lewej do prawej. Tool calling w trybie strict działa tak samo.
Follow-up questions (4)
- Can enforcing a format lower quality?
- Yes, when the schema asks for the decision before the reasoning or has no “other” option. What helps: a reasoning field before the decision, a reasoning model’s thinking, which the mask doesn’t cover, or two steps: a free-form answer first, then extraction.
- JSON mode, structured output or tool calling: when do you use which?
- Structured output when the answer feeds your code. Tool calling when the model has to choose an action from several, with strict mode for the arguments. JSON mode only where there is no strict mode.
- How do you do this on your own model?
- vLLM and SGLang have built-in constrained decoding engines (including XGrammar and llguidance): the schema turns into an automaton, and the automaton into a token mask at every step. The per-token overhead is small; the cost is compiling each new schema. With a reasoning model, start the server with a reasoning parser (e.g. --reasoning-parser in vLLM); otherwise the grammar applies from the first token and cuts off the thinking.
- The JSON parses, but the data is wrong. What next?
- Validation in code (constrained types, business rules) and a retry with a specific error message, with a cap on attempts. A recurring error is a signal to change the schema or the prompt, and the case goes into the eval set.
Sources
- OpenAI docs: Structured Outputs
- Claude docs: Structured outputs
- Willard, Louf: Efficient Guided Generation for Large Language Models (arXiv)
- Tam et al.: Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models (arXiv)
- .txt: Say What You Mean, a response to “Let Me Speak Freely?”