Structured Outputs: Getting JSON You Can Actually Parse
AI Agents

Structured Outputs: Getting JSON You Can Actually Parse

Asking a model for JSON and hoping is a bug waiting to happen. Schema-enforced outputs, and the validation you still need around them.

Sooner or later you need a model's output in a shape your code can consume. Asking politely for JSON works most of the time, and "most of the time" is the problem — the failures arrive in production, at volume.

Three levels of reliability

Prompt and pray. "Respond only with JSON." Works often. Fails with markdown fences, a preamble, a trailing explanation, or a subtly different shape. Requires defensive parsing.

JSON mode. The provider guarantees syntactically valid JSON. This removes parse errors but not schema errors — you still get valid JSON with missing fields or invented keys.

Schema-enforced output. You supply a JSON Schema and the decoder is constrained so only conforming tokens can be produced. Structure is guaranteed, not requested. When available, use it.

Structure is not correctness

This distinction matters and is routinely missed. A schema guarantees the shape. It says nothing about the content.

A schema requiring {"email": "string"} will happily accept "not-an-email". One requiring a confidence score between 0 and 1 will return a number in range that is nonetheless meaningless.

Validate semantics separately: check the email parses, the referenced ID exists, the enum value is one your system handles. Treat model output as untrusted input from a well-behaved but unreliable client.

Designing schemas models handle well

  • Flat beats nested. Deeply nested structures produce more errors. Two flat calls often beat one nested one.
  • Enums over free strings wherever the value comes from a known set. This eliminates an entire class of normalisation work.
  • Describe every field. Schema descriptions are read by the model and function as instructions.
  • Require reasoning first. A reasoning field before the answer field measurably improves the answer, because the model generates its justification before committing.
  • Avoid optional-everything. If every field is optional you will get sparse objects. Require what you actually need.

The ordering trick

JSON object key order follows generation order, and generation is sequential. A schema that produces the conclusion before the analysis gives you a conclusion generated without the analysis.

Put reasoning fields first and the verdict last. This is the same principle as chain-of-thought, expressed through schema design, and it is free.

Handling failure

Even with enforcement, things go wrong — truncation from hitting the token limit, refusals, edge cases. Build for it:

  1. Validate against the schema on receipt. Never trust the guarantee blindly.
  2. On validation failure, retry once with the error message included as feedback. Most failures resolve on the retry.
  3. Set a max token limit high enough for the largest plausible output — truncated JSON is invalid JSON.
  4. Log failures with the raw output. Patterns in the failures usually indicate a schema that needs simplifying.

Streaming structured output

Structured output and streaming coexist awkwardly. Partial JSON is invalid JSON, so you cannot parse incrementally without a streaming-tolerant parser, and those add complexity for modest benefit.

For most cases, stream only when the payload is large enough that waiting is genuinely painful. For a small object, wait for completion and parse once — the latency difference rarely justifies the machinery.

Versioning schemas

Schemas change, and outputs get stored. A field you make required today breaks every consumer reading records written yesterday.

Treat schemas the way you would any API contract: add optional fields freely, avoid changing the meaning of existing ones, and if you must break compatibility, version explicitly and migrate. Storing the schema version alongside each record costs almost nothing and saves a genuinely miserable afternoon later.

Refusals break your schema assumptions

A model that declines to answer still has to emit something. Under schema enforcement it will produce a conforming object with empty or placeholder values, which your code will happily accept as a real result.

Guard against it by including an explicit escape route in the schema — a nullable result plus a status field the model can set. Giving the model a valid way to say "I could not do this" is far better than forcing it to fabricate a well-formed answer.

When to use tools instead

If you want one structured result, use structured output. If you want the model to choose among several actions, use tool calling — the mechanisms are similar but tool calling models the choice explicitly, and gives you a natural place to put per-action schemas.

Reaching for a single enormous schema with a discriminated union usually means you wanted tools.

Common questions

Does JSON mode guarantee my schema is followed?

No. JSON mode guarantees syntactically valid JSON, not that it matches your schema. Only schema-enforced output constrains structure, and even then you must validate the content.

Why does field order in my schema matter?

Fields are generated sequentially, so a verdict emitted before the reasoning was produced without it. Put reasoning fields first and conclusions last.

What should I do when validation fails?

Retry once, including the validation error as feedback. Most failures resolve on retry. Log the raw output — repeated failures usually mean the schema is too complex.

Similar articles

Tool Calling: The Feature That Turns a Model Into an Agent
AI Agents
AI Agents·9 min read

Tool Calling: The Feature That Turns a Model Into an Agent

Tool calling is how a model asks your code to do something. Understanding the handshake — and designing good tools — is most of what makes an agent reliable.

Read
Agent Audit Logging: What to Record and What to Redact
AI Agents
AI Agents·9 min read

Agent Audit Logging: What to Record and What to Redact

When an agent does something surprising, the log is the only account of what happened. What a usable agent audit record contains, and what it must not.

Read
Agent Caching Strategies: Prompt, Tool and Result Layers
AI Agents
AI Agents·9 min read

Agent Caching Strategies: Prompt, Tool and Result Layers

Agents resend the whole transcript every turn and re-read the same files repeatedly. Three caching layers fix that, and each has different keying rules.

Read