AI Frontiers, part 45: Structured output — schemas, grammars, and constrained decoding
Part 45from the AI Frontiers series · 65 parts in all
Every AI integration eventually needs the model to return something a program can parse.
The first version of that integration is a sentence in the prompt: "Return JSON with keys
title, summary, and tags." It works. It also fails, at a rate of maybe one in fifty calls,
in ways that are individually boring and collectively expensive — a trailing comma, a
markdown fence around the object, a key renamed to summary_text for no reason, or
an apologetic sentence preceding the brace because the model decided to be helpful.
By late 2024 this was solved in the way infrastructure problems get solved: as a decoding problem rather than a prompting problem. Instead of asking for structure and validating afterward, you make non-conforming tokens impossible at generation time. OpenAI's Structured Outputs shipped that as a product feature in August 2024, guaranteeing conformance to a supplied JSON Schema; open-source libraries had been doing the same thing at the logits layer for a year and a half. This entry is about how constrained decoding works, why it is a real improvement and not a panacea, and how to build the interface around it so that valid output and correct output stay as close together as the tooling allows.
Why "please return JSON" is not a contract
A language model generates one token at a time from a probability distribution over its vocabulary. Nothing in that process knows about braces. When the model emits an opening brace, nothing forces it to emit a closing one; when it emits a key, nothing constrains the value type. Instruction-following ability means the distribution is strongly biased toward well-formed JSON, which is not the same as being constrained to it. Over a million calls, a one-percent malformed rate is ten thousand parse failures, and the retry logic that handles them doubles your latency and cost on a nontrivial slice of traffic.
There is a second, subtler failure: schema shape drifts. The model invents fields, nests an object where a string was expected, or returns a single element where an array was expected. These are the failures that make people give up on lightweight parsing and reach for a formal guarantee.
The three implementation families
Constrained decoding has one core idea — at each generation step, mask the logits of every token that cannot lead to a valid string — and three families of implementation, differing in what "valid" can express.
Regular languages. If the target is a regex or a finite-state machine, you can precompute which tokens are admissible from each parser state and mask the rest. This is the fastest and simplest case, and it covers a surprising amount: enumerations, dates, identifiers, phone numbers, and small flat JSON objects. Outlines was the library that made this concrete and public in 2023, compiling JSON schemas and regexes into index-based finite-state machines over the tokenizer's vocabulary and applying them with near-zero per-token overhead (Willard and Louf).
Context-free grammars. Nested structures — arbitrary JSON depth, balanced parentheses, SQL — need a pushdown automaton rather than a finite one. Grammar-constrained decoding with context-free grammars extended the approach to exactly this case (Geng et al.). The complication is that deciding admissibility requires the parser stack state, not just a node in an FSM, so the precomputation is heavier and incremental parsing has to be efficient. The classic pre-LLM result here is PICARD, which constrained sequence-to-SQL decoding to valid SQL fragments and measurably improved execution accuracy (Scholak et al.) — proof that the technique was not just about tidier output but about better answers.
Programmatic specification. The third family abandons "specify a grammar" for "write the generation as a program." LMQL let you interleave Python-like control flow with model calls, expressing constraints as part of the query rather than after it (Beurer-Kellner et al.), and Microsoft's Guidance took a similar templating approach. This is the most expressive option and the hardest to keep portable; it tends to be chosen by teams that are already treating prompt construction as software.
In practice, all three converge on the same mechanism at inference time: a mask over the logits, applied token by token, with an incremental parser tracking state. Understanding that mechanism matters because it explains the technique's limitations exactly.
Constraints do not add knowledge
This is the sentence I would put at the top of every structured-output design document: a
grammar can make an invalid output impossible, but it cannot make a valid output correct. If
the passage does not state a dollar amount, a schema requiring amount: number
gets you a confidently fabricated number that parses perfectly. The quantity of malformed
output goes to zero; the quantity of wrong output does not move, and in the case of
required fields it may go up, because you have removed the model's ability to say "not
specified."
The mitigation is structural, not grammatical. Make the fields you cannot guarantee nullable, and make the model's null a first-class answer rather than an error. Add a verbatim-quote or span field next to each extracted value so a downstream check can verify the value appears in the source — a cheap grounding assertion that catches fabrication without a second model call. Prefer enumerated types over free strings wherever the true answer space is closed, because selection is easier to get right than generation and easier to evaluate. And when the schema has more than a handful of required fields, consider splitting one extraction call into several narrower ones: each added constraint narrows the model's options, and at some point the model is spending capacity satisfying the grammar rather than reading the document.
Schemas are interfaces, and interfaces need versions
The moment a schema is in production, three things are true simultaneously: downstream
code depends on it, evaluation fixtures encode it, and the model's behavior is conditioned on
it. That makes schema changes migrations, with all the usual consequences. What I have found
workable: treat the schema as a checked-in artifact under version control with a changelog;
make new fields optional with a defined default so additions are backward compatible; never
repurpose an existing field name with a different meaning; and run the eval suite against both
the old and new schema during the transition. Union types earn their cost here, because
string | object is often the honest description of a field that used to be a
string and now sometimes carries structure.
One spec-level caution: not every JSON Schema is expressible as a grammar. Features that
require unbounded lookahead or complicated numeric validation — very large maxLength
values, intricate multipleOf constraints, complex oneOf combinations —
are either restricted or approximated by every implementation I have used. OpenAI's
implementation, for instance, requires additionalProperties: false and supports a
documented subset. Read the subset list before designing the schema, not after the call
starts failing.
Where structured output actually lives
Three production patterns absorb most of the value. Extraction — turning an invoice, a support email, or a ticket thread into typed fields — where the win is deleting hand-written parser code and its test suite. Tool and function calls, covered in part 30, where the schema is the API contract between the model and the runtime and constrained decoding removes the class of failures where the model invents an argument that does not exist. And agent control flow, where each step of a loop is a structured decision — which tool, what arguments, whether to stop — and malformed steps abort trajectories. In all three, the measurable win is not "cleaner output"; it is the deletion of retry logic, repair heuristics, and regex post-processing, which is where the real latency lives.
The practical checklist
Ship a schema, not a hope: constrain decoding when the provider supports it, and validate anyway, because conformance guarantees are per-call and your pipeline is not. Keep a nullable escape hatch for every field whose truth is not guaranteed present in the source. Pair extracted values with evidence spans. Version the schema and run the old and new versions side by side. Log the raw model output for a sample of traffic — even valid output, especially in the first weeks — because debugging a closed-loop extraction pipeline without the raw string is genuinely hard. And measure two numbers separately: schema-validity rate, which should be essentially 100%, and field-accuracy rate against human labels, which is the only one that tells you whether the system works. Teams that watch only the first number ship a beautiful pipeline that extracts the wrong things very reliably.
The deeper lesson of 2024's structured-output work is a general one about applied AI. Whenever a reliability problem can be moved from the model's judgment into the surrounding software, it should be. Decoding constraints, validation, retries, and interface contracts are all places where ordinary engineering beats prompt engineering, and the best AI systems of the next few years will be the ones with the least of their correctness entrusted to the model's goodwill.
What to do when the provider cannot constrain
Not every model you use will support grammar-guided decoding, and even when it does, the fastest path to a working integration is often a repair loop that costs nothing to build. The pattern that has held up for me is three layers, in increasing cost.
Extract. Find the outermost balanced object or array in the response and parse that, tolerating the markdown fence and the polite preamble. This is ten lines of code and it handles the majority of the malformed-output tail for models that are trying. Its limit is that it cannot fix a truncated generation — if the output hit the token cap mid-object, there is nothing to salvage.
Repair. On a validation failure, make one more call that includes the broken output and the exact validation error, and ask for a corrected version that conforms to the schema. This works surprisingly well because repair is a much easier task than generation, and it is cheap when run on a small model. Cap it at one or two attempts; the retry storm is the cost trap, and each attempt should be logged as a distinct span so that the retry rate is visible in your traces rather than hidden inside latency.
Escalate. If repair fails, treat the request as failed and either fall back to a human or to a less structured response. Silently substituting a default value is the choice that causes real damage, because it produces a system that always answers and is therefore never caught.
All three layers become unnecessary if you control the decoding. That is the argument for preferring providers and open models that support grammars, and it is why constrained decoding moved from a research technique to a checklist item on procurement in under two years.
Streaming, latency, and the tokenization trap
Two implementation details surprise people the first time they build this.
The first is that constrained decoding has a runtime cost of its own. The mask must be computed before each sampling step, which means parsing the partial output and consulting the automaton; done naively, this can add measurable overhead per token. The good implementations precompute token-to-state transitions once per grammar and reduce the per-step cost to a table lookup (Willard and Louf), which is why the difference between a careful library and a hand-rolled mask loop shows up as a real throughput gap rather than a rounding error.
The second is subtler and worth knowing before you debug it. Tokenizers are not
prefix-stable: a vocabulary token that decodes to "name" may be a single token in
one context and split across three in another, and a mask that forces a particular token
boundary can push the model into a completion it would never have chosen freely. The
engineering folklore names this token healing, and the symptom is a model that produces
technically valid output that is noticeably worse than its unconstrained counterpart on the
same task. Comparing constrained and unconstrained quality on your own evaluation set is the
only way to know whether you paid a quality tax for the reliability guarantee — and if you did
not measure it, you will not notice, because both outputs parse.
Works Cited
Beurer-Kellner, Luca, Marc Fischer, and Martin Vechev. "Prompting Is Programming: A Query Language for Large Language Models." arXiv, 2022, arxiv.org/abs/2212.06094. Accessed 7 Nov. 2024.
Geng, Saibo, et al. "Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning." arXiv, 2023, arxiv.org/abs/2305.13971. Accessed 7 Nov. 2024.
OpenAI. "Introducing Structured Outputs in the API." OpenAI, 6 Aug. 2024, openai.com/index/introducing-structured-outputs-in-the-api/. Accessed 7 Nov. 2024.
Scholak, Torsten, Dzmitry Bahdanau, and Christopher Pal. "PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models." arXiv, 2021, arxiv.org/abs/2109.05093. Accessed 7 Nov. 2024.
Shin, Richard, Christopher Lin, and Xinyun Chen. "Constrained Language Models Yield Few-Shot Semantic Parsers." arXiv, 2021, arxiv.org/abs/2104.08768. Accessed 7 Nov. 2024.
Willard, Brandon T., and Rémi Louf. "Efficient Guided Generation for Large Language Models." arXiv, 2023, arxiv.org/abs/2307.09702. Accessed 7 Nov. 2024.
Wright, Austin, Henry Andrews, and Ben Hutton. "JSON Schema: A Media Type for Describing JSON Documents." JSON Schema Specification, 2022, json-schema.org/specification. Accessed 7 Nov. 2024.